VCyber Twin
Detection Reality Index · Vol.3
The Between-Audit Gap · July 2026
What supervisors ask for — and what a default build actually measures

Between two mandated engagements, what evidence do you have that detection still works?

Financial-sector supervisors have spent several years moving the same direction: from evidence that is point-in-time to evidence that is current. For most institutions the honest answer to the question above is a coverage number — rules deployed, techniques mapped, controls documented. That is a count of what should fire. It is not a measurement of what did.
3 of 24
ATT&CK techniques that produced an alert on a default open-source SOC build. Twenty-one were silent.
0 of 51,937
Rules that fired on an ordinary port scan with the full ET open ruleset loaded
3.3s
Measured MTTD for network recon on QRadar 7.3.3 — after the telemetry gap was closed
The shift supervisors already made
Saudi Arabia's SAMA framework expects control effectiveness to be measured with periodic KPIs and KRIs at its "managed and measurable" maturity level, and its financial entities ethical red-teaming programme requires intelligence-led exercises at least every three years — testing detection and response, not just prevention, against live production.
In the EU, DORA has required a digital operational resilience testing programme running at least annually since January 2025 — vulnerability assessment, scenario-based testing, penetration testing — alongside incident reporting timelines tight enough that the response has to be rehearsed rather than improvised.
Neither framework names any particular tool, and this report does not claim otherwise. What both make unavoidable is the question that sits underneath them — the one at the top of this page.
The distinction this report is about

A coverage number counts what is supposed to fire. Replaying the technique tells you what actually fired, which rule caught it, and how long it took. Those are not the same number — and only one of them survives a follow-up question from an examiner.

Vol.1 — a default-configured open-source SOC build (Wazuh), 24 techniques
Discovery, lateral movement, collection, exfiltration, impact — across the kill chain. Each launch timestamped, each technique executed for real, then the SIEM checked for whether anything fired. Three produced an alert. Twenty-one were silent.
That is one reference configuration, not a claim about SOCs in general. What made it worth pursuing was the question of why a technique stays silent — because the answers turn out to be different in kind, and only one of them is a detection-engineering problem at all.
Scope, stated up front. Everything in this report was measured on our own bench — an isolated build with a real SIEM, real sensors, and real techniques replayed against it. It describes the configurations we ran. It is not an industry benchmark, and it is not a claim about your estate.
VCyber Twin
Detection Reality Index · Vol.3
What we measured · Why techniques stay silent
Vol.2.1 — the same method against IBM QRadar 7.3.3
TechniqueInitialRoot cause foundAfter the fix
T1486
Data Encrypted for Impact
ransomware file-encryption
SILENT auditd was never installed on the victim host — file-encryption activity generated no telemetry at all. DETECTED
syscall events reach the SIEM
T1046
Network Service Discovery
port scanning
SILENT No NIDS on that segment — a network scan produces no host syslog, so the SIEM was blind by construction. DETECTED
measured MTTD 3.3 seconds
Neither miss was a QRadar failure. Both were gaps in the telemetry reaching it. A detection stack cannot alert on an event it never receives, and no amount of rule tuning changes that. Finding exactly where that boundary sits is the work.
The number worth arguing with
While closing the second gap we installed Suricata on the victim host and loaded the full Emerging Threats open ruleset — 51,937 rules. Then we ran the port scan again. Zero of those 51,937 rules fired. The scan was only detected after we added one correctly-scoped threshold rule of our own.
This is not a criticism of Suricata or of the ET ruleset: Suricata dropped its port-scan preprocessor deliberately, and ET's scan signatures largely match tool fingerprints rather than raw scanning behaviour. Both are working as designed.
That is precisely what makes the number useful. A SOC looking at "51,937 detection rules loaded" reasonably concludes the surface is covered. Replaying one ordinary reconnaissance technique showed that reconnaissance remained invisible.
The sentence this whole report exists to support

Coverage has to be measured. It cannot be counted.

What we will not claim
A report that only contains good news is a brochure. Three limits belong on the record.
MTTD was not measurable cleanly for every technique.
On QRadar Community Edition, events become searchable roughly 60–90 seconds after they occur — a platform ingest characteristic, not detection logic. For T1486 we therefore report detection as restored rather than publishing a precise latency. A licensed production deployment measures this far more cleanly. The 3.3 second figure for T1046 is a measured value on this bench, carrying the same ingest and clock-skew caveats.
One bench is one configuration.
These numbers describe the build we ran, not an industry benchmark, and certainly not your estate.
This does not replace an intelligence-led red team, and it is not run against production.
It is a non-production measurement layer that sits between the exercises you are already required to run. Anyone telling a supervisor that a non-production replay satisfies a red-teaming obligation is misreading both.
VCyber Twin
Detection Reality Index · Vol.3
What we will not claim · Method · How to start
Why this matters between audits
An examiner asking "how do you know your detection works" can be answered three ways.
"We have 51,937 rules deployed."
A count.
"Our last red team was fourteen months ago, and here is the report."
A snapshot, decaying since the day it was signed.
"Here is what we replayed last month: which techniques fired, which rule caught each one, the measured time to detect, which stayed silent, why, and what we changed afterwards."
A measurement — current, and repeatable on demand. The third answer is the one that survives a follow-up question. Producing it does not require touching production, granting anyone access, or waiting three years for the next mandated cycle.
The method, so you can check it
Every figure above comes from a run we can reproduce, and the raw events are retained. The technique executes for real, launch time is recorded, the SIEM is queried for the resulting event, and the delta is the measured detection time. Where a technique is silent, we go and find the reason rather than recording a score — that is how both misses above turned into fixes, and then into detections. Claims about detection should be checkable. This is how we would want anyone to check ours.
A matched slice, run on our bench

We run a matched slice on request. You choose the technique that actually concerns you. We replay it on our bench against a build matched to yours and send back a one-page readout: what fired, which rule, measured time to detect, what stayed silent and why.

It runs in our lab. Nothing touches your environment, nobody gets access, no data leaves your side, and there is nothing to sign to receive it.

We run these one at a time, so how soon we can start depends on what is already on the bench. If it is useful, there is a paid engagement on the other side of it; if it is not, you have spent one email.

dongnx@atkvn.com

ATK New Technology · atkvn.com · Detection Reality Index Vol.1 · dongnx@atkvn.com  or  use the form
Figures are measured on ATK's own bench and characterise that build — they are not a measurement of any third party's production environment.

If you want this on a build of yours

This number came from our bench, and it characterises that build, not your live estate. If you want the same thing run against a SIEM build you care about, the whole ladder is written out with prices, including what we will not do.

How to work with us  ·  dongnx@atkvn.com

Step one on that ladder is free and needs one hostname you are responsible for. It measures cryptography rather than detection, but it costs you nothing to see how we write.