Detection Reality Index · Vol.1 · reference build

24 techniques went in.
3 produced an alert.

We replayed twenty-four real ATT&CK techniques against a default-configured Wazuh build and measured which ones actually fired. Three did. The entire quiet chain — discovery, collection, exfiltration, command-and-control — produced nothing at all.

Published in full, free, ungated. Method below, every technique listed, and an explicit statement of what this number does not mean.

24techniques replayed
3produced an alert
21silent
15–25sdetection window queried

Why this is not a coverage number

Most coverage claims mean one thing: "we have a rule for technique X." That is a statement about a ruleset. It can be true while the technique runs end to end and nobody hears about it.

We measured a different statement: "when technique X actually ran, did a rule fire, and how long did it take?" Those two questions have different answers, and the distance between them is where real intrusions live. A ruleset can be counted from a config file. Firing can only be observed by running something.

How we measured

Isolated lab, three machines: an attacker host, a victim host running the SOC agent, and the SIEM manager on an out-of-the-box configuration: the stock Wazuh ruleset plus a small number of community rules, which is what a build looks like before anyone has tuned it. We name the platform because a measurement you cannot reproduce is not a measurement, and you cannot reproduce this one without knowing what we ran it against.

For each of the 24 techniques:

Techniques were spread across the kill chain rather than clustered where detection is easy.

What we publish, and what we do not

Open, because you cannot check a number you cannot inspect: which techniques we ran, what we queried, the time window, the rule ids that fired, and which SIEM builds we can characterise.

Not published: how the bench itself is constructed — the replica harness, the orchestration, and the internal scoring. That is the part we would have to rebuild from scratch if we gave it away, and it is not needed to audit the result. Everything required to disagree with this report is above.

Results

Detected — 3 of 24

TechniqueWhat fired
T1136.001 — Create accountDefault rule 5902 (MITRE-mapped)
T1110.001 — SSH brute forceDefault rule 5503 (PAM) plus a custom rule
T1486 — Ransomware encryption / mass renameReal-time file integrity monitoring: default rule 553 plus a custom rule

All three are loud, high-volume, filesystem- or auth-level events. That is the pattern, and it is the finding: the build detects noise well.

Silent — 21 of 24, zero alerts

Discovery T1046 · T1082 · T1016 · T1049 · T1033 · T1518 · T1087 · T1057
Credential access reading /etc/shadow (T1003.008) · credentials in files (T1552.001)
Collection and staging T1005 · T1074
Exfiltration T1048
Command and control T1071.001
Defence evasion obfuscation T1027 · file deletion T1070.004
Ingress tooling T1105
File-based persistence cron (T1053.003) · sudoers (T1548.003) · SSH keys (T1098.004) · permission change (T1222.002)

Read the middle of that list again. Discovery, collection, exfiltration and C2 form a complete chain — land, look around, gather, take it out, keep talking to it. On this build that chain runs from end to end and the alert console stays empty.

Three uncomfortable truths

1. Having a rule is not detecting. The build lit up on the three noisiest actions and missed every quiet one. A rule count would have described this deployment as broadly covered.

2. Configuration decides — and the default chooses for you. Real-time file integrity monitoring caught ransomware. Scheduled file integrity monitoring missed every file-based persistence technique we ran. Same product, same ruleset, opposite outcome. Nobody made that trade-off on purpose; the default made it.

3. A host SIEM is blind to the network. Exfiltration and C2 produced zero alerts because there was no network sensor in the build. From the alert console there is no way to tell that data is leaving.

What this means if you run a SOC

An attacker running the quiet chain is invisible to a default host SOC until they do something loud. If the loud step never comes — and in a data-theft intrusion it often does not — the build never speaks.

A coverage percentage cannot tell you this, because it is computed from the rules you own rather than from what happened when something ran. Only replay tells you, and replay produces a number you can hand to someone else.

Closing the gap

This is fixable, and mostly not by buying anything:

Limitations — stated, not buried

This is one reference configuration, and it is not a verdict on Wazuh. It characterises the build we configured. It is not a measurement of any organisation's live environment, and it does not generalise to "all SOCs".

Some misses are missing telemetry, not missing rules. The default agent did not have command auditing enabled, so a portion of the silent techniques were never observable to the ruleset in the first place. That distinction is recorded per technique, and it matters: telemetry gaps are cheap to fix, rule gaps are not.

This report is vendor-authored. We built the bench and we sell an instrument that produces numbers like this one. That is exactly why the method is above and the technique list is complete: so the number can be argued with rather than believed.

If you think we are wrong

Good — the method is published so that this is possible. Run the same techniques against your own build and see whether you get a different answer; if you do, we would like to know, and we will say so in the next volume. A measurement that cannot come out badly is not a measurement.

Getting the number for a build you care about

The number above describes our reference build. It is not your number, and we would not pretend otherwise. If it is useful to see the same measurement run against a build configured to match one of the stacks you operate, that is what we do: it runs on our bench, nothing is installed in your environment, and the report says plainly that it characterises that build rather than a live estate.

Ask us a question About ATK

A matched slice, run on our bench

We run a matched slice on request. You choose the technique that actually concerns you. We replay it on our bench against a build matched to yours and send back a one-page readout: what fired, which rule, measured time to detect, what stayed silent and why.

It runs in our lab. Nothing touches your environment, nobody gets access, no data leaves your side, and there is nothing to sign to receive it.

We run these one at a time, so how soon we can start depends on what is already on the bench. If it is useful, there is a paid engagement on the other side of it; if it is not, you have spent one email.

dongnx@atkvn.com

The same question, measured elsewhere: Vol.2 — IBM QRadar · Vol.3 — what happens to a detection posture between two audits. Published in full, same terms: open method, stated limits, nothing gated.

If you want this on a build of yours

This number came from our bench, and it characterises that build, not your live estate. If you want the same thing run against a SIEM build you care about, the whole ladder is written out with prices, including what we will not do.

How to work with us  ·  dongnx@atkvn.com

Step one on that ladder is free and needs one hostname you are responsible for. It measures cryptography rather than detection, but it costs you nothing to see how we write.