Skip to main contentSkip to footer

Replacing fail2ban with the Linux Agent

How a host moves from fail2ban to the agent without a gap: first check what already owns the firewall objects, run both for a while, then switch fail2ban off for good, and how to go back in seconds if you have to.

Replacing fail2ban

The agent does not sit next to fail2ban, it takes over from it. fail2ban remains available as one source among many, so a host that still runs it keeps getting its bans reported, and a host without fail2ban loses nothing: every attack type that used to arrive through a jail now has a detector reading the original log instead.

Measured over eleven days on a production host: the agent found 187 of the 189 addresses fail2ban had banned on the same machine, the two it missed were test addresses fed in by hand, it reported 78 further addresses that had stayed under fail2ban's thresholds, and it matched the FTP bans three out of three. Escalation for repeat offenders comes from the agent's own ban history, and reportedip-agent test <file> takes the place of fail2ban-regex.

First check whether something already owns these names

This one cost an afternoon on a production fleet, so it comes before everything else. A hand-written firewall script of the kind this site used to document can use exactly the names the agent uses. Found in the field: the sets rip-whitelist, rip-ssh, rip-mail, rip-web, rip-ftp and rip-edge, plus a chain rip-blacklist hooked in with -A INPUT -j rip-blacklist, rebuilt every hour by a cron job, and nearly character for character the chain the agent builds itself.

On such a host the usual order is wrong in both directions. The agent's first sync rebuilds the chain and replaces the existing DROP rules. On a host installed with --mode log it replaces them with LOG rules, which leaves the host unprotected until the old cron job runs again, up to an hour later. And when that job does run, it rebuilds the chain its own way and removes the agent's rip-local rule along with it. Two programs, one chain, and each one undoing the other every hour.

So check before the first sync, and if the names collide, do not install this host with --mode log. Install the agent, which touches no rule, take the old whitelist over, disable the old cron job, and only then run the first sync with mode: drop, the default. Keep an observation phase with --mode log for hosts where no names collide.

bash
# 1. Does something already own these names?
ipset list -n | grep "^rip-"
iptables -S INPUT | grep rip-
iptables -S rip-blacklist 2>/dev/null | head
crontab -l | grep -iE "ipset|blacklist|reportedip"
ls -l /etc/cron.d/ /etc/cron.hourly/ 2>/dev/null

# 2. If yes, take the old whitelist over FIRST. The feed can be
#    downloaded again; that list cannot.
ipset list rip-whitelist    | sed -n "/^Members/,\$p" | tail -n +2 >  /tmp/rip-wl
ipset list rip-whitelist-v6 | sed -n "/^Members/,\$p" | tail -n +2 >> /tmp/rip-wl
wc -l /tmp/rip-wl
while read -r a; do
  [ -n "$a" ] && reportedip-agent whitelist add "$a" "imported from the old script"
done < /tmp/rip-wl
reportedip-agent whitelist list | wc -l

# 3. Disable the old job, then hand the chain over in one step.
crontab -l | grep -v blacklist | crontab -      # or: rm /etc/cron.d/<job>
# Only if the agent was installed with --mode log:
sed -i "s/^mode: log/mode: drop/" /etc/reportedip-agent/config.yaml
reportedip-agent sync
reportedip-agent status
Step 2 is not optional. On the hosts checked, the old script's whitelist held 208 IPv4 and 61 IPv6 prefixes: search engine crawlers, Cloudflare and the operator's own machines. Without importing it, the agent would have reported Googlebot, Cloudflare and the fleet itself on its first day. A whitelist is the one piece of state a migration cannot reconstruct from anywhere else, so carry it over before you switch anything, and read it back with reportedip-agent whitelist list before the first sync.

Running both for a while

Running both is safe and it is the sensible way to migrate. The agent never writes to /etc/fail2ban and fail2ban knows nothing about rip- sets, so the two use separate chains and separate state. Two rules that both drop the same address cost one packet comparison.

The one thing not to do is report the same event twice. If you keep the fail2ban action that posts bans over HTTP and configure fail2ban as an agent source, every ban leaves the host twice and is paid for twice out of your daily quota. Pick the agent or the action.

Install with REPORTEDIP_FAIL2BAN_SOURCE=0 when you are replacing fail2ban. The installer detects a running fail2ban and writes it as a source, which is right for a host that keeps it and wrong for a migration: once fail2ban is gone, so is the log that source points at. The variable leaves the source out, and without it the install prints one line saying the source is there and what to do with it after the removal. On a host already installed, take the fail2ban entry out of sources by hand, then stop and start the watch service.

bash
# What did fail2ban ban, and what does the agent make of the
# same logs? Same file, same window, two verdicts.
fail2ban-client status sshd
reportedip-agent test /var/log/auth.log

# Per jail, then the agent over the file behind it
for j in $(fail2ban-client status | sed -n "s/.*Jail list:\t*//p" | tr -d " " | tr "," " "); do
  echo "== $j"; fail2ban-client status "$j" | grep -E "Total banned|Currently banned"
done
reportedip-agent status | sed -n "/^sources:/,/^queue:/p"

Switching fail2ban off for good

Do this once you have compared the two for a few days and the agent's source list covers every jail you had.

bash
# 1. Stop it and keep it stopped.
systemctl stop fail2ban
systemctl disable fail2ban

# 2. Remove the fail2ban source from the agent config, then stop
#    and start the daemon. A restart is not enough for a removed
#    source: the daemon keeps tailing the old file.
$EDITOR /etc/reportedip-agent/config.yaml
systemctl stop reportedip-agent.service
systemctl start reportedip-agent.service

# 3. Check what fail2ban left in the kernel, in both address
#    families. Stopping it does not always clean up: a jail that
#    once used an iptables action leaves old-style f2b- chains behind.
nft list tables | grep f2b    || echo "no f2b table"
iptables -S  | grep "f2b-"    || echo "no f2b chain (v4)"
ip6tables -S | grep "f2b-"    || echo "no f2b chain (v6)"
ipset list -n | grep f2b      || echo "no f2b set"

# 4. Remove the leftovers from the RUNNING ruleset. The jump rule
#    carries -p and --dports, so it is deleted by its full text; a
#    plain -D INPUT -j f2b-sshd does not match it, and the chain then
#    refuses to go ("Device or resource busy"). Measured on a fleet:
#    the v6 leftover was the one that stayed.
nft delete table inet f2b-table 2>/dev/null
for T in iptables ip6tables; do
  $T -S INPUT | grep "f2b-" | sed "s/^-A INPUT //" | while read -r rule; do eval "$T -D INPUT $rule"; done
  for c in $($T -S | awk "/^-N f2b-/{print \$2}"); do $T -F "$c" && $T -X "$c"; done
done

# 5. And from the SAVED one, or they come back at the next boot.
#    Save without the rip- lines as well: the agent rebuilds its chain
#    at every sync, and a saved rule that names a rip- set which does
#    not exist yet at boot makes iptables-restore fail as a whole.
iptables-save  | grep -v -e "f2b-" -e "rip-" > /etc/iptables/rules.v4
ip6tables-save | grep -v -e "f2b-" -e "rip-" > /etc/iptables/rules.v6

# 6. Confirm the agent is happy without it.
reportedip-agent status | grep -i fail2ban
reportedip-agent status; echo "exit $?"
Step 5 is the one people skip. If you only delete the f2b- chains from the running ruleset, netfilter-persistent restores them from /etc/iptables/rules.v4 and from /etc/iptables/rules.v6 at every boot. They come back empty, so they drop nothing and look harmless, and they will still be there in a year confusing whoever reads the ruleset next. Check both files: on one host the leftovers were only in the v6 file, and whoever greps rules.v4 alone walks straight past them. Clean the running ruleset first, then save it, and verify after a reboot.

Rollback is one command and takes seconds, because the agent never touched /etc/fail2ban: systemctl enable --now fail2ban brings every jail and its own table back with the agent unaffected. Measured on a live host, all eleven jails were back in eight seconds.

Last updated: · Maintained by the ReportedIP team

Security Focused
GDPR Compliant
Made in Germany
Back to Docs