Detection and rules
What the agent finds in the logs a server already writes: the fourteen source types and their thresholds, how a source the installer did not find is added, what leaves the machine and what never does, and the rule files every detector is made of, including the ones you write yourself.
What the agent finds in your logs
Each source has its own detector and its own threshold, because five failed SSH logins in ten minutes
and fifty suspicious web requests in two hours are the same statement about an attacker and not the
same number. Rotation, copytruncate and a log that disappears for a while are handled,
and an unreadable file produces one warning rather than one per read.
The fourteen source types
| Type | Typical path | What it counts as a hit |
|---|---|---|
sshd | journald, else /var/log/auth.log or /var/log/secure | Failed passwords, invalid users, failed keys for a user that does not exist, exceeded attempt limits, and the pre-authentication failures that mark scanners: no identification string, a bad protocol version, garbage in the banner exchange. A refused key for a valid user is deliberately not a hit, because an admin with five keys writes four of those per successful login. A successful login clears that address's failures. |
web | /var/log/nginx/access.log, or a per-site glob | Login POSTs, counted regardless of status code, and paths no legitimate client asks for. Everything else in an access log is ignored, because this is the source with real customers behind it. |
web-error | /var/log/nginx/error.log, /var/log/apache2/error.log | Requests the web server itself refused, failed HTTP basic auth, and ModSecurity lines of critical severity when they land here instead of an audit log. Rate-limit lines are deliberately not counted: a limit fires on a real browser with twenty tabs. TLS handshake and PHP messages say something about the server, not the client. |
postfix | /var/log/mail.log, /var/log/maillog | Three separate event sources out of one file: SASL authentication failures, NOQUEUE rejects that are really 5xx (a 450 is greylisting and does not count), and amavis blocks. Delivered mail is never a hit. |
dovecot | the same mail log | IMAP and POP3 login failures, one per line whatever the "N attempts" says. The remote address is read from the rip= field, never lip=, which is the server's own. An aborted login with no auth attempt is not a hit: that is a TLS scan, and no password was tried. |
exim | /var/log/exim4/mainlog, /var/log/exim/main.log | Authentication failures and rejected senders. Unverified against a live host: the patterns come from the documented log format. |
ftp | /var/log/syslog, /var/log/messages | Failed authentications of pure-ftpd, proftpd and vsftpd. All three log through syslog, and each writes the peer differently, so three patterns are used. pure-ftpd is measured: 494 of 494 lines on a production host had one single form. |
named | the same syslog | Only an incoming query that bind turned down. The big trap is the reverse direction: connection refused resolving is this host's own resolver failing to reach someone else's name server, and 90 percent of the measured lines were that. A detector without the exclusion reports other people's name servers as attackers. |
panel | /var/log/ispconfig/auth.log | Failed logins on the ISPConfig panel. The file is 0 bytes on every host measured, which is not a fault: the panel simply writes nothing there. doctor points it out after a week. |
modsec | /var/log/apache2/modsec_audit.log | One event per transaction whose verdict carries critical severity. This is the only stateful detector, because a transaction spans several lines. Nothing from the request travels in the event: not the rule ID, not the matched data, not the URI. |
csf | /var/log/lfd.log | What CSF actually blocked, never what it merely detected. No second threshold on top: lfd counted before the agent saw the line, so one block is one report. |
fail2ban | the unit's journal, else /var/log/fail2ban.log | Ban actions only. A Restore Ban is never reported, because that is what a fail2ban restart replays from its own database and not a new attack. A filter Found line is not reported either: the jail has not decided yet. |
imunify360 | polls imunify360-agent | Incidents from the CLI, mapped to categories by rule name prefix. Experimental, and unverified against a licensed installation; the decoder skips what it cannot parse rather than turning an unexpected answer into a dead source. |
scan | the kernel journal (journalctl -k), else /var/log/kern.log or /var/log/messages | The lines of the port scan rule the sync run installs with this source: a SYN to a port this host does not listen on, prefixed rip-scan:. Ten probes inside ten minutes from one address are a scan. Nothing in a kernel log line is written by the peer. See below. |
web-app is not a source type and cannot be written into sources. It is a
second detector that runs on the lines a web source already reads, and it has its own
threshold, which is why it appears under its own name in the threshold table above. An agent whose
configuration names it refuses to start: sources[0]: type "web-app" is not supported in this version.
Port scans, from the kernel log
Every other source reads a log a service writes. The scan source reads the kernel,
and it is the one source that also changes the firewall: with it in the config, the sync run adds a
rule behind the lists that logs a SYN to a port this host does not listen on, with the prefix
rip-scan: and at most ten lines a minute after a first burst of 20, so a scan of all 65535 ports cannot flood the
journal. With the ipset backend that is a chain of its own, rip-blacklist-scan; with
nftables it is one rule in the chain. The listening ports are read from /proc on every
sync, so a service you start is not a scan target after the next run, and there is nothing to
declare. Only a socket something outside could reach counts as listening. A port bound to
127.0.0.1 or ::1 alone stays in the rule, because a packet that arrives
there from the network is a probe of a closed port and nothing else.
The list from the last sync is not the only check: the rule also asks the kernel when the packet arrives, and a SYN to a port on which a socket is listening at that moment is a connection and is never logged. That covers a service started since the last sync and a daemon that opens a port for a single transfer. The passive port range of a running FTP server (pure-ftpd, proftpd, vsftpd) is read from its configuration and left out of the rule; if the server runs without a configured range, the kernel's ephemeral port range is left out instead. A kernel without the socket match keeps the rule without that check.
The rule stands behind the whitelist and behind the lists. An address on your whitelist returned
from the chain before it and produces no line, and an address the lists already drop never reaches
it either. In mode: log and mode: off it does what everything else does in
those modes.
The source reads the kernel journal, journalctl -k, and falls back to
/var/log/kern.log or /var/log/messages where there is no journal; a
path you set wins. The shipped rule scan counts ten probes inside ten
minutes from one address, reports them under category 14 (Port Scan) as port probes and
bans the address like any other find. A client that retries one closed port sends at most six
packets, which is why the threshold is ten. install writes the source on a fresh host
as the last entry of sources; a host installed earlier keeps its configuration, and
adding it is one line, followed by a restart of the watch service and a sync for the rule.
# config.yaml: the source, next to the ones the installer wrote
sources:
- type: sshd
- type: scan
# /etc/reportedip-agent/rules.d/50-scan.yaml: ban a scanner on this
# host, but leave the report to hosts that see more of it. Every field
# not named here stays as shipped.
format: 1
rules:
- id: scan
action: {ban: true, report: false}
# The same file with enabled: false switches the rule off; the firewall
# rule and its log lines stay. A threshold of your own goes into
# config.yaml and wins over the file:
thresholds:
scan: {hits: 20, window_minutes: 10}
Adding a source the installer did not find
A source the installer did not find is simply absent from the config, and adding it is one entry with
a path or a glob. Because a file with a sources key replaces the default list entirely,
add your entry next to the existing ones rather than in a second block.
sources:
- type: sshd
- type: web
path: /var/log/nginx/access.log
# A second web server on a different path
- type: web
path: /var/log/caddy/access.log
# Every site of a panel host, rescanned every ten minutes.
# At most five wildcard segments.
- type: web
glob: /var/www/clients/*/web*/log/access.log
exclude:
# This site reports to the API on its own, through Hive or a
# honeypot. Without the exclude the host sends every address
# twice and pays twice out of the daily quota.
- /var/log/ispconfig/httpd/honeypot.example.com/*
# Exim on a host the installer saw as a Postfix machine
- type: exim
path: /var/log/exim4/mainlog
Two paths that resolve to the same file are not a problem and need no exclude: ISPConfig publishes
every access log twice, under /var/www and under /var/log/ispconfig, and
the agent reads such a file once. Before you trust a new entry, point
reportedip-agent test at the file, then restart the watch service, then check
status: the sources section shows hits per source. doctor shows a
files= count per source, and a source with a count of zero resolves to nothing at all.
What leaves the machine
A report carries the address, the threat category IDs and a generated sentence. No log line, no request body, no User-Agent, no URL and no user name is ever transmitted, so a report cannot leak your customers, your paths or your credentials even by accident.
One precision, because an absolute promise would be wrong here. Two sources put a name from the log
into that generated sentence: with fail2ban it is the name of the jail that banned, with
imunify360 the name of the rule that fired. Both are filtered to letters, digits, dot,
hyphen and underscore and cut to 32 characters before they leave the machine, because a jail name
comes out of a log line and is therefore attacker-adjacent input. No other source sends anything out
of a log line at all.
The hostname is sent, for one purpose only: so that you recognise your own machine in the server list
instead of comparing random identifiers. It is filtered on the client to A-Za-z0-9._-,
with everything else removed rather than replaced, cut to 191 characters, and it never goes out
without the install ID next to it.
The identity is still the install ID. The hostname is a label and nothing more: renaming the machine changes nothing about the licence, nothing about which host holds it and nothing the server decides. You can also give a host a label of your own in your account, which is shown ahead of the hostname and is never derived from the machine.
Four addresses are never reported and never blocked, and this is a boundary in the code rather than a setting: private and reserved ranges, the loopback, every address of this server's own interfaces and its default gateway, and everything in your whitelist. The install-time SSH client address sits in the auto-whitelist, which counts as part of that last group.
Rules
A rule says which lines of one source are hits, where the address stands in such a line, what a hit is called, and how many of them inside how many minutes become a ban or a report. Every detector of the agent is a rule, and every rule is a file. Three layers are read and merged by rule id:
| Layer | Where it lives | What it is |
|---|---|---|
| 1 | Inside the binary | The shipped rules, one file per source type. They run from the binary. After the next sync a copy for reading lies under /var/lib/reportedip-agent/rules.d/standard/; an edit there is overwritten by the sync after it. install creates that directory and packs/ next to it, empty, so the three layers are visible before the first sync fills two of them. |
| 2 | A signed rule pack from reportedip.com | The sync run fetches it, verifies its signature against a key built into the binary before a single byte of it is read as a rule, and keeps it under the state directory. A host that cannot reach the service keeps the pack it has. A pack that does not verify is refused, and the one on disk stays. |
| 3 | /etc/reportedip-agent/rules.d/*.yaml | Your files: overrides of shipped rules, and rules of your own. |
The pack wins over the binary, your files win over the pack. An override names a rule by its
id and only the fields to change; every field it does not name keeps the value of the
layer below. match and examples are replaced as a whole,
action field by field. enabled: false switches a rule off. The file name
plays no part: a file called 10-sshd.yaml under rules.d does not shadow the
shipped file of that name, it is read like any other and merged by the ids it contains, so a release
that renames a shipped file cannot turn your override into a duplicate.
The file format
One file holds any number of rules. This is a complete one, taken from the agent's own test suite:
format: 1
rules:
- id: my-sshd
source: sshd
match:
regex: '(Failed password|Invalid user) .* from '
anchors: ["Failed password", "Invalid user", "Accepted "]
reset: 'Accepted (password|publickey) for .* from '
addr: {after: " from ", occurrence: last}
categories: [22, 18]
noun: failed logins
threshold: {hits: 5, window_minutes: 10}
examples:
match:
- {line: "Sep 24 14:47:38 host sshd[1]: Failed password for root from 45.33.32.156 port 51422 ssh2", addr: 45.33.32.156}
inject:
- {line: "Sep 24 14:47:38 host sshd[1]: Invalid user x from 8.8.8.8 port 22 from 45.33.32.156 port 51422", addr: 45.33.32.156}
nomatch:
- "Sep 24 14:47:38 host sshd[1]: Connection closed by 45.33.32.156 port 1"
| Field | Meaning | What the loader checks |
|---|---|---|
format | The first line of every rule file. This agent reads format 1. | A file with a higher number was written for a newer agent and is skipped; so is a file without the line. |
id | The name of the rule, as rules list and the journal show it. | a-z, 0-9 and - after the first character, at most 32 bytes. An id that exists in a lower layer makes the rule an override. |
source | The source type whose lines the rule reads, one of the fourteen: sshd, fail2ban, web, web-error, postfix, dovecot, exim, ftp, named, panel, modsec, csf, imunify360, scan. | Required for a new rule. |
event | The event source the hits count under, the name an entry under thresholds in config.yaml refers to, and the name the comment of a report carries. Defaults to the id. | Same charset as an id. Rules that share an event share its threshold, and a hit counts once whichever of them caught it. A rule of yours cannot take over the event of a shipped rule: override that rule by its id instead. |
enabled | true or false; a missing key is true. | enabled: false in an override switches a shipped rule off without touching its file. |
match.regex | The pattern that decides whether a line is a hit, and nothing else. RE2 syntax, as Go reads it. | At most 512 bytes. No flags such as (?i), (?s) or (?m), no named groups: the address does not come from the pattern. |
match.anchors | Literal words every hit carries. The pattern runs only on a line that carries one of them; every other line never reaches it. | Required. Each at least 6 bytes and a literal part of regex or reset. When reset is set, at least one anchor has to occur in it, or the reset never fires. |
match.ignore | Patterns that drop a line before regex sees it. | RE2, the same limits as regex. |
match.reset | The pattern of a line that clears the counter of an address, a successful login. | RE2, the same limits. The address is read the same way as for a hit. |
match.builtin | Instead of a pattern, the name of a compiled detector. The shipped rules for web, web-app, web-error, modsec, exim, panel, named, fail2ban, csf and imunify360 are of this kind and carry policy only. | Not together with regex, and without anchors, ignore, reset, addr or examples. |
addr | Where the address stands in a matching line. Exactly one form: after with occurrence: last or first; between: [left, right]; in_brackets: N, the Nth [...] counting from 1; before_byte: ":", one byte; field: N, the Nth whitespace-separated field counting from 0. An optional upto: "<" cuts the line at the first occurrence of that marker before the form is applied, so a field the peer writes behind it cannot be reached. | Required for a regex rule. after needs at least 3 bytes and an occurrence; upto is at most 16 bytes. |
categories | The threat category ids a report carries. There are 63, numbered 1 to 63. | Required when the rule reports. Each at least 1. |
noun | What a hit is called in the comment of a report, such as failed logins. | Letters, digits and spaces, at most 40. |
threshold | hits inside window_minutes, the pair a fail2ban jail called maxretry and findtime. | hits 1 to 10000, window_minutes 1 to 1440. Without the block, min_hits and window_minutes of config.yaml apply. An entry under thresholds in config.yaml wins over both. |
action | ban and report, each true or false; both are true when missing. | Both false is refused: a rule that does nothing is switched off with enabled: false instead. |
examples | The tests of the rule. match: lines and the address each yields. inject: lines that carry a second, decoy address in a field the peer writes, and the address the rule must still yield. nomatch: lines the rule must not match. | Required for a regex rule, at least one of each kind, and at least one nomatch line has to carry an address. They run every time the rule is loaded; a rule whose examples do not hold does not load. |
Why the address is not a capture group
fail2ban had this twice, CVE-2013-2178 and CVE-2009-0362, and both times it was the same mistake:
the address came out of the pattern, and an attacker had put an address of their choosing into a
field the pattern reached. RE2 finds the leftmost match, so in
Invalid user x from 8.8.8.8 port 22 from 45.33.32.156 port 51422 a group behind the
first from would ban 8.8.8.8, which is the user name the attacker typed.
addr: {after: " from ", occurrence: last} takes the last one, the one sshd wrote itself.
That is why a rule names a position instead of a group, and why every rule carries an
inject line that proves it.
What protects you from your own rule
The examples run every time the rule is loaded, so a pattern that has stopped matching its own
lines does not load. A pattern rule your layer adds or overrides bans and does not report for the first 24 hours of its
content (a rule with match.builtin never does): a wrong kernel entry expires and shows in ban list, a wrong report cannot be
taken back. rules status marks such a rule as young. A daemon restart starts the day
again.
Two guards watch a rule while it runs. When three addresses from your whitelist reach its threshold inside ten minutes, the rule reads the wrong field, because real attackers never stand on the whitelist; it stops reporting until its file changes, and keeps banning, which the whitelist covers. When a young rule pushes more than 50 distinct addresses over its threshold within one minute, it matches every visitor and is suspended whole, bans and reports, until its content changes. An established rule is never suspended for a burst: a distributed brute force with hundreds of sources a minute is what an sshd rule exists for. A rule from the pack is watched by the burst guard like a young one, because a pack reaches every host at once.
An entry under thresholds in config.yaml keeps winning over the file, so a host you
tuned loses nothing. A file that is not valid YAML, or names a newer format, is skipped alone and
the others load. A rule that fails one of the checks above costs your whole layer: the shipped
rules and the pack run on, status and doctor say so, and
rules check names the file, the line and the id.
The rules commands
reportedip-agent rules list # every rule: id, source, event, kind, state, threshold, action
reportedip-agent rules show sshd # one rule as merged, and whether it is overridden
reportedip-agent rules check # load rules.d the way the daemon does, and say what is wrong
reportedip-agent rules check /root/my.yaml # try a file before it is put in place
reportedip-agent rules export sshd # print the shipped file of a rule, as the template
reportedip-agent rules status # state, hits since the daemon started, young and suspended rules
reportedip-agent rules disable sshd # write enabled: false for one id; enable takes it back
reportedip-agent test --rule my-sshd /var/log/auth.log # run one rule alone over a real log
list, show, check, export and
status read and print. check without a file loads your directory exactly
as the daemon would; with files it loads those on top of the shipped set instead, so a file can be
tried before it lies in rules.d. export prints the shipped file a rule
lives in, comments included, as the starting point of an override. enable and
disable are the two that write: one generated file,
90-agent-toggles.yaml, root only, rewritten on every call. test --rule
runs that one rule and no compiled detector beside it, takes the source type from the rule, and
reports and bans nothing, like test always does.
Writing a rule of your own
- Start from a shipped file.
reportedip-agent rules export sshd > /root/60-sshd.yamlprints the file the rule lives in, comments and examples included. - Cut it down, or give it a new id. For an override keep the
idand only the fields to change; the file below is complete. For a rule of your own give it a newidwithsource,match,addr,categoriesandexamples, likemy-sshdabove. A new id is a new event source, and the shipped rule keeps running beside it. - Check it.
reportedip-agent rules check /root/60-sshd.yamlloads the file on top of the shipped set exactly as the daemon would and prints every problem with file, line and id, or one line that says ok. - Put it in place.
install -m 644 -o root -g root /root/60-sshd.yaml /etc/reportedip-agent/rules.d/. The directory and every file in it have to belong to root and be writable by nobody else, and a symlink that points out of the directory is refused; otherwise the whole layer is not used. - Wait a minute. The daemon picks the change up on its own, without a restart, and
the counters start afresh.
reportedip-agent rules statusshows the rule, marked young for its first day.reportedip-agent test --rule my-sshd /var/log/auth.logshows meanwhile what it would have caught.
# /etc/reportedip-agent/rules.d/60-sshd.yaml: the shipped sshd rule,
# stricter on this host. Everything not named here stays as shipped.
format: 1
rules:
- id: sshd
threshold: {hits: 3, window_minutes: 10}
Last updated: · Maintained by the ReportedIP team