Replacing fail2ban with the Linux Agent
How a host moves from fail2ban to the agent without a gap: first check what already owns the firewall objects, run both for a while, then switch fail2ban off for good, and how to go back in seconds if you have to.
Replacing fail2ban
The agent does not sit next to fail2ban, it takes over from it. fail2ban remains available as one source among many, so a host that still runs it keeps getting its bans reported, and a host without fail2ban loses nothing: every attack type that used to arrive through a jail now has a detector reading the original log instead.
Measured over eleven days on a production host: the agent found 187 of the 189 addresses fail2ban had
banned on the same machine, the two it missed were test addresses fed in by hand, it reported 78
further addresses that had stayed under fail2ban's thresholds, and it matched the FTP bans three out
of three. Escalation for repeat offenders comes from the agent's own ban history, and
reportedip-agent test <file> takes the place of fail2ban-regex.
First check whether something already owns these names
This one cost an afternoon on a production fleet, so it comes before everything else. A hand-written
firewall script of the kind this site used to document can use exactly the names the agent
uses. Found in the field: the sets rip-whitelist, rip-ssh,
rip-mail, rip-web, rip-ftp and rip-edge, plus a
chain rip-blacklist hooked in with -A INPUT -j rip-blacklist, rebuilt every
hour by a cron job, and nearly character for character the chain the agent builds itself.
On such a host the usual order is wrong in both directions. The agent's first sync rebuilds the chain
and replaces the existing DROP rules. On a host installed with --mode log it
replaces them with LOG rules, which leaves the host unprotected until the old cron job
runs again, up to an hour later. And when that job does run, it
rebuilds the chain its own way and removes the agent's rip-local rule along with it. Two
programs, one chain, and each one undoing the other every hour.
So check before the first sync, and if the names collide, do not install this host with
--mode log. Install the agent, which touches no rule, take the old whitelist over,
disable the old cron job, and only then run the first sync with mode: drop, the
default. Keep an observation phase with --mode log for hosts where no names collide.
# 1. Does something already own these names?
ipset list -n | grep "^rip-"
iptables -S INPUT | grep rip-
iptables -S rip-blacklist 2>/dev/null | head
crontab -l | grep -iE "ipset|blacklist|reportedip"
ls -l /etc/cron.d/ /etc/cron.hourly/ 2>/dev/null
# 2. If yes, take the old whitelist over FIRST. The feed can be
# downloaded again; that list cannot.
ipset list rip-whitelist | sed -n "/^Members/,\$p" | tail -n +2 > /tmp/rip-wl
ipset list rip-whitelist-v6 | sed -n "/^Members/,\$p" | tail -n +2 >> /tmp/rip-wl
wc -l /tmp/rip-wl
while read -r a; do
[ -n "$a" ] && reportedip-agent whitelist add "$a" "imported from the old script"
done < /tmp/rip-wl
reportedip-agent whitelist list | wc -l
# 3. Disable the old job, then hand the chain over in one step.
crontab -l | grep -v blacklist | crontab - # or: rm /etc/cron.d/<job>
# Only if the agent was installed with --mode log:
sed -i "s/^mode: log/mode: drop/" /etc/reportedip-agent/config.yaml
reportedip-agent sync
reportedip-agent status
reportedip-agent whitelist list before the first sync.
Running both for a while
Running both is safe and it is the sensible way to migrate. The agent never writes to
/etc/fail2ban and fail2ban knows nothing about rip- sets, so the two use
separate chains and separate state. Two rules that both drop the same address cost one packet
comparison.
The one thing not to do is report the same event twice. If you keep the fail2ban action that posts bans over HTTP and configure fail2ban as an agent source, every ban leaves the host twice and is paid for twice out of your daily quota. Pick the agent or the action.
Install with REPORTEDIP_FAIL2BAN_SOURCE=0 when you are replacing fail2ban.
The installer detects a running fail2ban and writes it as a source, which is right for a host that
keeps it and wrong for a migration: once fail2ban is gone, so is the log that source points at. The
variable leaves the source out, and without it the install prints one line saying the source is there
and what to do with it after the removal. On a host already installed, take the
fail2ban entry out of sources by hand, then stop and start the watch
service.
# What did fail2ban ban, and what does the agent make of the
# same logs? Same file, same window, two verdicts.
fail2ban-client status sshd
reportedip-agent test /var/log/auth.log
# Per jail, then the agent over the file behind it
for j in $(fail2ban-client status | sed -n "s/.*Jail list:\t*//p" | tr -d " " | tr "," " "); do
echo "== $j"; fail2ban-client status "$j" | grep -E "Total banned|Currently banned"
done
reportedip-agent status | sed -n "/^sources:/,/^queue:/p"
Switching fail2ban off for good
Do this once you have compared the two for a few days and the agent's source list covers every jail you had.
# 1. Stop it and keep it stopped.
systemctl stop fail2ban
systemctl disable fail2ban
# 2. Remove the fail2ban source from the agent config, then stop
# and start the daemon. A restart is not enough for a removed
# source: the daemon keeps tailing the old file.
$EDITOR /etc/reportedip-agent/config.yaml
systemctl stop reportedip-agent.service
systemctl start reportedip-agent.service
# 3. Check what fail2ban left in the kernel, in both address
# families. Stopping it does not always clean up: a jail that
# once used an iptables action leaves old-style f2b- chains behind.
nft list tables | grep f2b || echo "no f2b table"
iptables -S | grep "f2b-" || echo "no f2b chain (v4)"
ip6tables -S | grep "f2b-" || echo "no f2b chain (v6)"
ipset list -n | grep f2b || echo "no f2b set"
# 4. Remove the leftovers from the RUNNING ruleset. The jump rule
# carries -p and --dports, so it is deleted by its full text; a
# plain -D INPUT -j f2b-sshd does not match it, and the chain then
# refuses to go ("Device or resource busy"). Measured on a fleet:
# the v6 leftover was the one that stayed.
nft delete table inet f2b-table 2>/dev/null
for T in iptables ip6tables; do
$T -S INPUT | grep "f2b-" | sed "s/^-A INPUT //" | while read -r rule; do eval "$T -D INPUT $rule"; done
for c in $($T -S | awk "/^-N f2b-/{print \$2}"); do $T -F "$c" && $T -X "$c"; done
done
# 5. And from the SAVED one, or they come back at the next boot.
# Save without the rip- lines as well: the agent rebuilds its chain
# at every sync, and a saved rule that names a rip- set which does
# not exist yet at boot makes iptables-restore fail as a whole.
iptables-save | grep -v -e "f2b-" -e "rip-" > /etc/iptables/rules.v4
ip6tables-save | grep -v -e "f2b-" -e "rip-" > /etc/iptables/rules.v6
# 6. Confirm the agent is happy without it.
reportedip-agent status | grep -i fail2ban
reportedip-agent status; echo "exit $?"
f2b- chains
from the running ruleset, netfilter-persistent restores them from
/etc/iptables/rules.v4 and from /etc/iptables/rules.v6 at
every boot. They come back empty, so they drop nothing and look harmless, and they will still be
there in a year confusing whoever reads the ruleset next. Check both files: on one host the
leftovers were only in the v6 file, and whoever greps rules.v4 alone walks straight
past them. Clean the running ruleset first, then save it, and verify after a reboot.
Rollback is one command and takes seconds, because the agent never touched
/etc/fail2ban: systemctl enable --now fail2ban brings every jail and its own
table back with the agent unaffected. Measured on a live host, all eleven jails were back in eight
seconds.
Last updated: · Maintained by the ReportedIP team