Monitoring the Linux Agent
How to watch the agent from Zabbix, Nagios, Icinga, Checkmk or Prometheus: what a check should cover,
the JSON that status --json prints, the permissions it needs, a ready Zabbix template,
and short setups for the other systems. Available from version 0.3.40.
What a check should cover
The agent protects a host in two ways that fail differently. The lists in the kernel keep blocking even when the agent is gone, so a stopped agent does not look broken from outside: the firewall still drops the known addresses, and nothing new is detected, banned or reported. A useful check therefore looks at more than the process.
| Question | Where the answer is | Healthy |
|---|---|---|
| Does the watch daemon run? | daemon.running, or systemctl is-active reportedip-agent | true, active |
| Does the sync timer run? | sync.last_attempt_age_s | below 5400 (90 minutes; a host fetches every 15 or every 60 minutes, depending on the plan) |
| Is the chain in the packet path? | chain_ok | true, on a host that blocks |
| Anything else wrong? | exit code, status, problems | 0, ok, empty |
The daemon writes a heartbeat every 30 seconds. status calls it dead when that heartbeat is
older than two minutes, so a stopped daemon shows up in the next check that runs at least two minutes
after the stop. Nothing else in status changes when the daemon stops, and that is why this
line exists.
status --json
reportedip-agent status --json runs the same checks as status and prints the
result as one JSON document. The exit code follows the same rules, with one difference: the text mode asks the API live and is degraded when that fails, --json asks nothing and so never is for that reason. It sends no request to the API:
the account part is the answer the last sync cached, and account.cached_age_s says how old
it is. A check every few minutes therefore costs nothing and cannot hang on the network.
{
"schema": 1,
"status": "degraded",
"code": 1,
"problems": [
{ "key": "daemon", "text": "the watch daemon is not running, last heartbeat 6m0s ago" }
],
"version": "v0.3.40",
"generated_at": "2026-09-29T20:40:00Z",
"daemon": { "running": false, "heartbeat_age_s": 361 },
"update": { "latest": "0.3.40", "behind": false, "last_check_age_s": 5120, "last_error": "" },
"backend": "nftables",
"mode": "drop",
"chain_ok": true,
"disk": { "free_mb": 18234, "min_mb": 200 },
"sync": { "running": false, "last_attempt_age_s": 412, "last_ok_age_s": 412 },
"lists": { "ssh": { "ipv4": 21873, "ipv6": 1204, "last_ok_age_s": 412, "fails": 0, "last_error": "" } },
"bans": { "enabled": true, "kernel": 3, "records": 3, "missing": 0 },
"sources": { "/var/log/auth.log": { "state": "ok", "last_line_age_s": 12, "hits": 58 } },
"queue": { "entries": 0, "max": 5000, "oldest_age_s": -1, "sent_total": 311, "sender_paused": false },
"account": { "role": "reportedip_professional", "reports_today": 214, "report_limit": 1000,
"feed": true, "license": "licensed", "reputation_listed": false,
"group": "", "cached_age_s": 14320 },
"rules": { "total": 20, "operator": 0, "disabled": 0, "problems": 0 },
"health": {}
}
| Field | Meaning |
|---|---|
schema | The version of this format. Fields are only ever added under the same number; a field is never renamed, retyped or removed without raising it. |
status, code | ok and 0, degraded and 1, error and 2. The same as the exit code. |
problems | One entry per reason for degraded, each with a short key (daemon, chain, lists, sources, queue, bans, disk, update, license, reputation, backend, state) and a sentence. Empty when the host is healthy. |
*_age_s | Ages in whole seconds. -1 means never or unknown, for example before the first sync. |
backend | ipset, nftables, or none on a host that only reports. Such a host has no chain and no sync, so chain_ok is false there by design. |
lists, sources | One object per list and per log source, keyed by name. |
health | The open health conditions, the same ones the state mail is about. |
error | Only with code 2: the configuration or the state directory cannot be used. The document is printed anyway, so the check gets a value instead of nothing. |
The document holds counts, ages and states only: no address, no key and no log line, so it can be stored by a monitoring server without further thought.
Permissions
status needs root. The configuration holds your API key and is readable by root only, the
state directory is 0700, and the firewall sets can only be read with root rights. A
monitoring agent runs as its own user, so it needs a sudo rule for exactly this one command and nothing
else:
# /etc/sudoers.d/reportedip-monitoring, mode 0440
# Replace zabbix with the user your monitoring agent runs as (nagios, icinga, ...).
zabbix ALL=(root) NOPASSWD: /usr/local/bin/reportedip-agent status --json
# Check the file before it is used, then test as that user:
visudo -cf /etc/sudoers.d/reportedip-monitoring
sudo -u zabbix sudo -n /usr/local/bin/reportedip-agent status --json
Zabbix
For Zabbix 7.0 and newer, with Zabbix agent or Zabbix agent 2. One UserParameter fetches the document every five minutes, and every other item takes its value from that one call, so the agent runs once per interval and not once per item.
Install the check
# Zabbix agent 2: /etc/zabbix/zabbix_agent2.d/reportedip.conf
# Zabbix agent: /etc/zabbix/zabbix_agentd.d/reportedip.conf
UserParameter=reportedip.status,sudo -n /usr/local/bin/reportedip-agent status --json 2>/dev/null || true
# then the sudo rule from the section above, and:
systemctl restart zabbix-agent2 # or zabbix-agent
zabbix_agent2 -t reportedip.status # prints the document
The || true keeps a degraded host from turning into an unsupported item: the exit code is in
the document as code, and the item has to receive the document in every case.
Import the template
Two variants of the same template, depending on how your agents talk to the server. Import one of them under Data collection, Templates, Import and link it to every host that runs the agent.
- reportedip_agent_zabbix7.yaml: items of type Zabbix agent (passive)
- reportedip_agent_zabbix7_active.yaml: items of type Zabbix agent (active)
The template sets a timeout of 20 seconds on the status item. An agent older than 7.0 ignores that and
uses its own Timeout, which is 3 seconds by default; raise it to 20 in the agent
configuration on such hosts.
What the template watches
| Trigger | Severity | Fires when |
|---|---|---|
| watch daemon is not running | High | daemon.running is false, or no reportedip-agent watch process runs. The process count needs no UserParameter, so it still works when the status check does not. |
| firewall chain is not in place | High | chain_ok is false on a host with a backend. |
| configuration error, the agent cannot run | High | code is 2. |
no sync for {$RIP.SYNC.MAXAGE} | Average | The last feed request is older than the macro (90 minutes), or there was none, on a host with a backend. It depends on the licence trigger, because without a licence no request goes out. |
no data for {$RIP.NODATA} | Average | The status item got nothing for 15 minutes. |
report queue over {$RIP.QUEUE.PCT}% | Warning | The queue is fuller than the macro (80). |
| host has no server licence, feed paused | Warning | account.license is unlicensed. |
| reporting address is listed | Warning | The address this host reports from stands in the community database. |
| agent degraded | Warning | Every other reason for code 1. It depends on the triggers above, so one problem raises one alert and not two. The item Problems names the reason. |
| newer agent version available | Info | A newer release exists. The self update installs it within six hours, so this only stays open when the update path is broken or auto_update is off. |
Two discovery rules add one item set per list (IPv4 and IPv6 entries, failures in a row) and per log source (state, hits, age of the last line). Every threshold is a macro on the template and can be changed per host.
When the items stay empty
Value of type "string" is not suitable for value type "Numeric" or an item that never gets a value
almost always means the command could not run as the Zabbix user: the sudo rule is missing, has the wrong
mode (it must be 0440, otherwise sudo ignores the file), or names a different path. Run the
test line from the permissions section as the Zabbix user. An active agent fetches its item list only every
RefreshActiveChecks seconds, so the first value can take a few minutes after the template is
linked.
Nagios and Icinga
The exit codes of status are the plugin codes already: 0 OK, 1 WARNING, 2 CRITICAL. A plugin
should print one line, though, so a short wrapper turns the document into one line with performance
data. It needs jq.
#!/bin/sh
# /usr/local/lib/nagios/plugins/check_reportedip
out=$(sudo -n /usr/local/bin/reportedip-agent status --json 2>/dev/null)
rc=$?
[ -n "$out" ] || { echo "REPORTEDIP UNKNOWN: no status document, check the sudo rule"; exit 3; }
line=$(printf '%s' "$out" | jq -r '"REPORTEDIP " + (.status | ascii_upcase) + ": "
+ (if (.problems | length) > 0 then ([.problems[].text] | join("; ")) else "daemon running, chain ok" end)
+ " | queue=\(.queue.entries) bans=\(.bans.kernel) sync_age=\(.sync.last_attempt_age_s)s"') || {
echo "REPORTEDIP UNKNOWN: no status document"; exit 3; }
echo "$line"
[ "$rc" -le 2 ] && exit "$rc" || exit 3
# NRPE: /etc/nagios/nrpe.d/reportedip.cfg
command[check_reportedip]=/usr/local/lib/nagios/plugins/check_reportedip
# Icinga 2 with the agent: a CheckCommand that runs the same script
object CheckCommand "reportedip" {
command = [ "/usr/local/lib/nagios/plugins/check_reportedip" ]
}
The sudo rule from the permissions section applies, with the user NRPE or the Icinga agent runs as.
Checkmk
A local check runs as root inside the Checkmk agent, so it needs no sudo rule. Put the script into the local directory of the agent and discover the services of the host once.
#!/bin/sh
# /usr/lib/check_mk_agent/local/reportedip, mode 0755
out=$(/usr/local/bin/reportedip-agent status --json 2>/dev/null)
rc=$?
[ -n "$out" ] || { echo "3 ReportedIP_Agent - no status document"; exit 0; }
[ "$rc" -gt 2 ] && rc=3
printf '%s' "$out" | jq -r --argjson rc "$rc" '"\($rc) ReportedIP_Agent queue=\(.queue.entries)|bans=\(.bans.kernel)|sync_age=\(.sync.last_attempt_age_s) "
+ (if (.problems | length) > 0 then ([.problems[].text] | join(", ")) else "daemon running, chain ok" end)' \
|| echo "3 ReportedIP_Agent - no status document"
Prometheus
The textfile collector of the node exporter reads metrics from files, so a timer that writes one file
every five minutes is enough. The directory is whatever --collector.textfile.directory
points to on your hosts.
#!/bin/sh
# /usr/local/sbin/reportedip-prom, run as root every 5 minutes (cron or a systemd timer)
dir=/var/lib/prometheus/node-exporter
out=$(/usr/local/bin/reportedip-agent status --json 2>/dev/null)
[ -n "$out" ] || exit 1 # keep the old file; its age raises the alert
printf '%s' "$out" | jq -r '
"reportedip_status_code \(.code)",
"reportedip_daemon_running \(if .daemon.running then 1 else 0 end)",
"reportedip_heartbeat_age_seconds \(.daemon.heartbeat_age_s)",
"reportedip_sync_attempt_age_seconds \(.sync.last_attempt_age_s)",
"reportedip_chain_ok \(if .chain_ok then 1 else 0 end)",
"reportedip_bans_active \(.bans.kernel)",
"reportedip_queue_entries \(.queue.entries)",
(.lists | to_entries[] | "reportedip_list_entries{list=\"\(.key)\",family=\"ipv4\"} \(.value.ipv4)",
"reportedip_list_entries{list=\"\(.key)\",family=\"ipv6\"} \(.value.ipv6)")
' > "$dir/reportedip.prom.tmp" && mv "$dir/reportedip.prom.tmp" "$dir/reportedip.prom"
The rename at the end matters: the collector must never read a half written file. Alert on
reportedip_status_code > 0, on reportedip_daemon_running == 0, and on the age
of the file itself through node_textfile_mtime_seconds, which catches a timer that stopped.
Without a monitoring system
The agent mails you itself when something is wrong and again when it is fixed, once per condition, if
notify.email is set; see Configuration.
That mail comes from the sync run, so it covers a broken feed, chain, source or licence, but it cannot
report that the whole host is down. A plain check from outside, a ping or an SSH port check, covers that
part. For a quick look by hand, reportedip-agent status and its exit code are enough.
Last updated: · Maintained by the ReportedIP team