How to check in one minute
First find out on which side of the manager the alert disappears. Put your rule ID in place of 100080:
grep -cE '"id": ?"100080"' /var/ossec/logs/alerts/alerts.json
- More than 0: the rule fires. The alert is lost after the manager, on the way into the index. See mapper_parsing_exception.
- 0: go through the five causes below. The first two checks cost one command each: restart the manager (
/var/ossec/bin/wazuh-control restart) if the rule file changed since the last restart, and look for dropped events:
grep -E "Input queue is full|dropping events" /var/ossec/logs/ossec.log
grep events_dropped /var/ossec/var/run/wazuh-analysisd.state
If logtest shows the parent's rule ID rather than yours, that is a different problem: the rule was dropped at load.
Why it happens
1. The manager has not loaded your edit yet
In our 4.14.7 runs, wazuh-logtest used the edited local_rules.xml immediately. We changed one rule's level and parent four times without restarting, and logtest followed every change. Rule changes reach the running manager when it restarts, so a rule can pass logtest before it exists in production.
2. Something earlier in the sequence takes the event
We tested a child of 5715 at level 9 matching Accepted password, and fed two inputs through wazuh-logtest:
| input | rule that took the last line |
|---|---|
one Accepted password line | 100080, our rule |
8 Failed password lines from the same IP in 8 seconds, then the same Accepted password line | 40112, level 12 |
Same rule, same final line, different history. In production the line arrives after whatever came before it. Logtest keeps state within one session, so it shows this too if you feed it the real sequence:
/var/ossec/bin/wazuh-logtest < sequence.log 2>&1 | grep -E "^[[:space:]]+(id|level|description): " | tail -3
The 2>&1 matters: logtest on 4.14 prints its results to stderr, and a script that reads only stdout sees nothing. More on this case: rule shadowed by a sibling.
3. Frequency rules count from zero after a restart
We loaded a correlation rule with frequency="4" timeframe="300" and same_srcip on sshd failures (on 4.14.7 the sshd "Failed password" rule is 5760), and sent real log lines through logcollector, not logtest:
| case | what we sent | correlation rule |
|---|---|---|
| control | 4 failures, no restart | fired |
| restart in the middle | 2 failures, restart, 2 failures, all within 300 s | did not fire |
| after that | 2 more failures, 4 since the restart | fired |
All six lines of the middle case reached the manager, so it was the counter, not lost events. Since rule changes need a restart, every rule deploy also resets any correlation that was counting.
4. The event never reaches the manager
Logtest takes the line you paste; it never sees the agent or the network. We flooded a 4.14.7 agent and manager three ways:
| setup | reached the manager | alert about the loss |
|---|---|---|
| agent defaults, client buffer on, 900,000 lines at 10,000/s | 45,493 (5%) | 202 once, 203 44 times, 204 never |
| client buffer disabled, same flood | 57,047 (6.3%) | none; one warning in the agent's own ossec.log |
manager overloaded (<limits><eps> at 100), 120,000 lines | 22,391 (18.7%) | none; warnings in the manager's ossec.log and events_dropped in the state file |
On the agents, the case with no alert at all shows up only here:
grep "message queue is full" /var/ossec/logs/ossec.log
5. For agent events, <hostname> is the agent name
By default logtest treats the line as a local event, so <hostname> is compared with the host name in the syslog header. For an event that comes from an agent, analysisd replaces that value with the agent name taken from the event's location (cleanevent.c in 4.14.7), and the rule compares against the agent name. The alert still prints the syslog host under predecoder.hostname, because that field is parsed again from full_log when the JSON is written. So the alert can show the value your rule was looking for, and the rule still did not match.
We checked it with wazuh-logtest on a 4.14.7 manager, one su line whose syslog host is DN4, from an agent named cae:
wazuh-logtest -l | <hostname>DN4</hostname> | <hostname>^cae$</hostname> |
|---|---|---|
stdin (default, local event) | fires | does not fire |
[002] (cae) 192.168.8.28->journald (agent event) | does not fire | fires |
[002] (caesar) 192.168.8.29->journald | does not fire | does not fire |
To test your own line the way analysisd sees it, pass the agent-style location. Use your agent's ID, name and IP:
/var/ossec/bin/wazuh-logtest -l '[002] (cae) 192.168.8.28->journald'
Also check: the rule level
A rule below log_alert_level can match and still write no alert. The stock 4.14.7 config we read sets it to 3. The stock Microsoft Graph parent 99500 is an example: sign-ins match it at level 0 and never become alerts.
grep log_alert_level /var/ossec/etc/ossec.conf
Fix
- Not loaded yet: restart the manager after every rule change, then test again.
- Earlier events take the line: test with the real sequence, not one line. If a sibling wins, change the level or the parent of your rule; the options we measured are on the shadowing page.
- Frequency rules and restarts: expect a correlation in progress to be lost at each restart, and plan rule deploys with that in mind.
- Agent events and
<hostname>: match the agent name, anchored, for example<hostname>^cae$</hostname>. Without^and$,caealso matches an agent calledcaesar. If the agent is renamed, the rule goes quiet. - Lost before the manager: keep the agent's client buffer on. It is the only one of the three loss paths that raised an alert in our runs, and it is on by default. Alert on
203, not204. Watchossec.logandevents_droppedfor the other two.
Limits of what we measured
Everything here ran on Wazuh 4.14.7 in throwaway single-node containers. The sequence test ran through wazuh-logtest; the correlation test ran through logcollector. The flood test used one agent reading a file; we did not measure agents receiving syslog over the network, a manager receiving syslog directly, other kinds of manager overload such as CPU or disk, or clusters. We saw logtest follow rule edits without a restart; we did not separately measure the running manager ignoring an edit until restart. The <hostname> case was measured with wazuh-logtest and an agent-style location, not with a journald event sent through a real agent. We have not measured other cases where logtest and live input decode the same event differently, for example Windows eventchannel events.