Skip to main contentSkip to footer

Monitoring the Linux Agent

How to watch the agent from Zabbix, Nagios, Icinga, Checkmk or Prometheus: what a check should cover, the JSON that status --json prints, the permissions it needs, a ready Zabbix template, and short setups for the other systems. Available from version 0.3.40.

What a check should cover

The agent protects a host in two ways that fail differently. The lists in the kernel keep blocking even when the agent is gone, so a stopped agent does not look broken from outside: the firewall still drops the known addresses, and nothing new is detected, banned or reported. A useful check therefore looks at more than the process.

QuestionWhere the answer isHealthy
Does the watch daemon run?daemon.running, or systemctl is-active reportedip-agenttrue, active
Does the sync timer run?sync.last_attempt_age_sbelow 5400 (90 minutes; a host fetches every 15 or every 60 minutes, depending on the plan)
Is the chain in the packet path?chain_oktrue, on a host that blocks
Anything else wrong?exit code, status, problems0, ok, empty

The daemon writes a heartbeat every 30 seconds. status calls it dead when that heartbeat is older than two minutes, so a stopped daemon shows up in the next check that runs at least two minutes after the stop. Nothing else in status changes when the daemon stops, and that is why this line exists.

status --json

reportedip-agent status --json runs the same checks as status and prints the result as one JSON document. The exit code follows the same rules, with one difference: the text mode asks the API live and is degraded when that fails, --json asks nothing and so never is for that reason. It sends no request to the API: the account part is the answer the last sync cached, and account.cached_age_s says how old it is. A check every few minutes therefore costs nothing and cannot hang on the network.

json
{
  "schema": 1,
  "status": "degraded",
  "code": 1,
  "problems": [
    { "key": "daemon", "text": "the watch daemon is not running, last heartbeat 6m0s ago" }
  ],
  "version": "v0.3.40",
  "generated_at": "2026-09-29T20:40:00Z",
  "daemon":  { "running": false, "heartbeat_age_s": 361 },
  "update":  { "latest": "0.3.40", "behind": false, "last_check_age_s": 5120, "last_error": "" },
  "backend": "nftables",
  "mode": "drop",
  "chain_ok": true,
  "disk":    { "free_mb": 18234, "min_mb": 200 },
  "sync":    { "running": false, "last_attempt_age_s": 412, "last_ok_age_s": 412 },
  "lists":   { "ssh": { "ipv4": 21873, "ipv6": 1204, "last_ok_age_s": 412, "fails": 0, "last_error": "" } },
  "bans":    { "enabled": true, "kernel": 3, "records": 3, "missing": 0 },
  "sources": { "/var/log/auth.log": { "state": "ok", "last_line_age_s": 12, "hits": 58 } },
  "queue":   { "entries": 0, "max": 5000, "oldest_age_s": -1, "sent_total": 311, "sender_paused": false },
  "account": { "role": "reportedip_professional", "reports_today": 214, "report_limit": 1000,
               "feed": true, "license": "licensed", "reputation_listed": false,
               "group": "", "cached_age_s": 14320 },
  "rules":   { "total": 20, "operator": 0, "disabled": 0, "problems": 0 },
  "health":  {}
}
FieldMeaning
schemaThe version of this format. Fields are only ever added under the same number; a field is never renamed, retyped or removed without raising it.
status, codeok and 0, degraded and 1, error and 2. The same as the exit code.
problemsOne entry per reason for degraded, each with a short key (daemon, chain, lists, sources, queue, bans, disk, update, license, reputation, backend, state) and a sentence. Empty when the host is healthy.
*_age_sAges in whole seconds. -1 means never or unknown, for example before the first sync.
backendipset, nftables, or none on a host that only reports. Such a host has no chain and no sync, so chain_ok is false there by design.
lists, sourcesOne object per list and per log source, keyed by name.
healthThe open health conditions, the same ones the state mail is about.
errorOnly with code 2: the configuration or the state directory cannot be used. The document is printed anyway, so the check gets a value instead of nothing.

The document holds counts, ages and states only: no address, no key and no log line, so it can be stored by a monitoring server without further thought.

Permissions

status needs root. The configuration holds your API key and is readable by root only, the state directory is 0700, and the firewall sets can only be read with root rights. A monitoring agent runs as its own user, so it needs a sudo rule for exactly this one command and nothing else:

bash
# /etc/sudoers.d/reportedip-monitoring, mode 0440
# Replace zabbix with the user your monitoring agent runs as (nagios, icinga, ...).
zabbix ALL=(root) NOPASSWD: /usr/local/bin/reportedip-agent status --json

# Check the file before it is used, then test as that user:
visudo -cf /etc/sudoers.d/reportedip-monitoring
sudo -u zabbix sudo -n /usr/local/bin/reportedip-agent status --json
Test as the monitoring user, not as root. A test as root always works and proves nothing about the rule. Only the last line above, run as root, shows what the monitoring agent will get.

Zabbix

For Zabbix 7.0 and newer, with Zabbix agent or Zabbix agent 2. One UserParameter fetches the document every five minutes, and every other item takes its value from that one call, so the agent runs once per interval and not once per item.

Install the check

bash
# Zabbix agent 2: /etc/zabbix/zabbix_agent2.d/reportedip.conf
# Zabbix agent:   /etc/zabbix/zabbix_agentd.d/reportedip.conf
UserParameter=reportedip.status,sudo -n /usr/local/bin/reportedip-agent status --json 2>/dev/null || true

# then the sudo rule from the section above, and:
systemctl restart zabbix-agent2    # or zabbix-agent
zabbix_agent2 -t reportedip.status # prints the document

The || true keeps a degraded host from turning into an unsupported item: the exit code is in the document as code, and the item has to receive the document in every case.

Import the template

Two variants of the same template, depending on how your agents talk to the server. Import one of them under Data collection, Templates, Import and link it to every host that runs the agent.

The template sets a timeout of 20 seconds on the status item. An agent older than 7.0 ignores that and uses its own Timeout, which is 3 seconds by default; raise it to 20 in the agent configuration on such hosts.

What the template watches

TriggerSeverityFires when
watch daemon is not runningHighdaemon.running is false, or no reportedip-agent watch process runs. The process count needs no UserParameter, so it still works when the status check does not.
firewall chain is not in placeHighchain_ok is false on a host with a backend.
configuration error, the agent cannot runHighcode is 2.
no sync for {$RIP.SYNC.MAXAGE}AverageThe last feed request is older than the macro (90 minutes), or there was none, on a host with a backend. It depends on the licence trigger, because without a licence no request goes out.
no data for {$RIP.NODATA}AverageThe status item got nothing for 15 minutes.
report queue over {$RIP.QUEUE.PCT}%WarningThe queue is fuller than the macro (80).
host has no server licence, feed pausedWarningaccount.license is unlicensed.
reporting address is listedWarningThe address this host reports from stands in the community database.
agent degradedWarningEvery other reason for code 1. It depends on the triggers above, so one problem raises one alert and not two. The item Problems names the reason.
newer agent version availableInfoA newer release exists. The self update installs it within six hours, so this only stays open when the update path is broken or auto_update is off.

Two discovery rules add one item set per list (IPv4 and IPv6 entries, failures in a row) and per log source (state, hits, age of the last line). Every threshold is a macro on the template and can be changed per host.

When the items stay empty

Value of type "string" is not suitable for value type "Numeric" or an item that never gets a value almost always means the command could not run as the Zabbix user: the sudo rule is missing, has the wrong mode (it must be 0440, otherwise sudo ignores the file), or names a different path. Run the test line from the permissions section as the Zabbix user. An active agent fetches its item list only every RefreshActiveChecks seconds, so the first value can take a few minutes after the template is linked.

Nagios and Icinga

The exit codes of status are the plugin codes already: 0 OK, 1 WARNING, 2 CRITICAL. A plugin should print one line, though, so a short wrapper turns the document into one line with performance data. It needs jq.

bash
#!/bin/sh
# /usr/local/lib/nagios/plugins/check_reportedip
out=$(sudo -n /usr/local/bin/reportedip-agent status --json 2>/dev/null)
rc=$?
[ -n "$out" ] || { echo "REPORTEDIP UNKNOWN: no status document, check the sudo rule"; exit 3; }
line=$(printf '%s' "$out" | jq -r '"REPORTEDIP " + (.status | ascii_upcase) + ": "
  + (if (.problems | length) > 0 then ([.problems[].text] | join("; ")) else "daemon running, chain ok" end)
  + " | queue=\(.queue.entries) bans=\(.bans.kernel) sync_age=\(.sync.last_attempt_age_s)s"') || {
  echo "REPORTEDIP UNKNOWN: no status document"; exit 3; }
echo "$line"
[ "$rc" -le 2 ] && exit "$rc" || exit 3
bash
# NRPE: /etc/nagios/nrpe.d/reportedip.cfg
command[check_reportedip]=/usr/local/lib/nagios/plugins/check_reportedip

# Icinga 2 with the agent: a CheckCommand that runs the same script
object CheckCommand "reportedip" {
  command = [ "/usr/local/lib/nagios/plugins/check_reportedip" ]
}

The sudo rule from the permissions section applies, with the user NRPE or the Icinga agent runs as.

Checkmk

A local check runs as root inside the Checkmk agent, so it needs no sudo rule. Put the script into the local directory of the agent and discover the services of the host once.

bash
#!/bin/sh
# /usr/lib/check_mk_agent/local/reportedip, mode 0755
out=$(/usr/local/bin/reportedip-agent status --json 2>/dev/null)
rc=$?
[ -n "$out" ] || { echo "3 ReportedIP_Agent - no status document"; exit 0; }
[ "$rc" -gt 2 ] && rc=3
printf '%s' "$out" | jq -r --argjson rc "$rc" '"\($rc) ReportedIP_Agent queue=\(.queue.entries)|bans=\(.bans.kernel)|sync_age=\(.sync.last_attempt_age_s) "
  + (if (.problems | length) > 0 then ([.problems[].text] | join(", ")) else "daemon running, chain ok" end)' \
  || echo "3 ReportedIP_Agent - no status document"

Prometheus

The textfile collector of the node exporter reads metrics from files, so a timer that writes one file every five minutes is enough. The directory is whatever --collector.textfile.directory points to on your hosts.

bash
#!/bin/sh
# /usr/local/sbin/reportedip-prom, run as root every 5 minutes (cron or a systemd timer)
dir=/var/lib/prometheus/node-exporter
out=$(/usr/local/bin/reportedip-agent status --json 2>/dev/null)
[ -n "$out" ] || exit 1   # keep the old file; its age raises the alert
printf '%s' "$out" | jq -r '
  "reportedip_status_code \(.code)",
  "reportedip_daemon_running \(if .daemon.running then 1 else 0 end)",
  "reportedip_heartbeat_age_seconds \(.daemon.heartbeat_age_s)",
  "reportedip_sync_attempt_age_seconds \(.sync.last_attempt_age_s)",
  "reportedip_chain_ok \(if .chain_ok then 1 else 0 end)",
  "reportedip_bans_active \(.bans.kernel)",
  "reportedip_queue_entries \(.queue.entries)",
  (.lists | to_entries[] | "reportedip_list_entries{list=\"\(.key)\",family=\"ipv4\"} \(.value.ipv4)",
                           "reportedip_list_entries{list=\"\(.key)\",family=\"ipv6\"} \(.value.ipv6)")
' > "$dir/reportedip.prom.tmp" && mv "$dir/reportedip.prom.tmp" "$dir/reportedip.prom"

The rename at the end matters: the collector must never read a half written file. Alert on reportedip_status_code > 0, on reportedip_daemon_running == 0, and on the age of the file itself through node_textfile_mtime_seconds, which catches a timer that stopped.

Without a monitoring system

The agent mails you itself when something is wrong and again when it is fixed, once per condition, if notify.email is set; see Configuration. That mail comes from the sync run, so it covers a broken feed, chain, source or licence, but it cannot report that the whole host is down. A plain check from outside, a ping or an SSH port check, covers that part. For a quick look by hand, reportedip-agent status and its exit code are enough.

Last updated: · Maintained by the ReportedIP team

Security Focused
GDPR Compliant
Made in Germany
Back to Docs