vyos: move PPPoE off the config plane onto a gated systemd unit
Some checks failed
Some checks failed
PPPoE HA could not work as written, and the reason is structural rather than a
bug: `set interfaces pppoe pppoe0 disable` and `delete` are handled identically
by interfaces_pppoe.py -- both UNLINK /etc/ppp/peers/pppoe0. That path is pppd's
own options file, so the resting state destroyed exactly what the promotion path
needed, and `ppp@pppoe0` restart-looped against it (observed: 47 restarts, zero
sessions at the access concentrator). It also made op-mode `connect interface
pppoe0` unusable, and put every failover behind a priority-322 commit where one
unrelated invalid node fails the whole thing -- which has already taken the
10 gig down once.
pppoe0 is now configured identically and ENABLED on both routers, so the peers
file always exists, and dialling is gated by a drop-in on the unit:
ConditionPathExists=/run/vrrp-wan/may-dial
ConditionPathExists=/etc/ppp/peers/pppoe0
/run is tmpfs, so the gate is shut at boot and neither box can dial before VRRP
has decided. That matters more than it first appears: with the node enabled,
interfaces_pppoe.py restarts ppp on EVERY commit touching the pppoe subtree when
the daemon is not running -- so the backup actively tries to dial whenever
anything commits. The gate is the only thing making that a no-op, which is why
vrrp-wan-reconcile now refuses to bless a box whose drop-in is missing: /etc is
per-image, and a VyOS upgrade would otherwise silently remove the protection.
may-dial is a LEASE, not a flag. ConditionPathExists is evaluated at start only
-- it can prevent a dial, never revoke one -- so a reconciler that stops running
while its box is demoted would keep the one ISP session for ever. The reconciler
renews the lease; a new 5s vrrp-wan-guard revokes it, and only ever revokes. It
fired correctly first time: "GUARD: lease stale (81s > 75s)".
Also: remove-then-stop on release (the file's absence blocks a NEW start that a
concurrent commit would trigger); a flap damper, because two routers that both
believe they hold the VIP will both dial and each dial kills the other's session
-- against a real ISP that is how an account gets rate-limited; and a guard on
`cfg` returning empty under commit-lock contention, which had already produced
one spurious "releasing" on a box that needed nothing.
GRACE 90 -> 180. accel-ppp's dead-peer budget is lcp-echo-interval(30) x
failure(3) = 90s, so the old value sat exactly on the boundary: a hard failover
into an AC that does not replace the stale session would fail its own check,
shed the VIPs, and leave both routers in FAULT.
The sim could not have tested any of this. Both routers now get the identical
WAN -- the secondary had none "because two PPPoE clients sharing one credential
is a different failure mode than anything production has", which is backwards:
that IS production. It also left the pair incomparable, ten NAT rules against
none. Safety now comes from resting state, not asymmetry.
Three more things the sim was hiding:
- the drift check's secondary regex omitted interfaces pppoe/bonding, nat
source and protocols failover, so it reported "in sync" for a box with no
WAN at all;
- the VRRP health-check and transition-script hooks existed on both live VMs
and in NEITHER generator -- the mechanism under test was pure undetected
drift;
- labsim-vyos's only default route was the libvirt-NAT scaffold, so every
"the LAN still has internet" verdict on it was answered by eth2 rather than
the WAN. --drop-scaffold applied; the earlier DHCP-failover proof is being
re-run because of it.
vrrp-wan-install ends the other half of that: the sim's previous proof came from
scripts hand-`sed`-ed in place, so the tested behaviour was not the committed
behaviour. `--check` now makes that a hard failure.
First green run: master holds both WANs, backup released, and the AC reports
exactly ONE session. sim-net-apply.sh check: all four in sync.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01DMVzWZgiKW2wquf5z8S1yH
This commit is contained in:
50
migration/ppp-vrrp-gate.conf
Normal file
50
migration/ppp-vrrp-gate.conf
Normal file
@@ -0,0 +1,50 @@
|
||||
# Installed to /etc/systemd/system/ppp@pppoe0.service.d/10-vrrp-wan-gate.conf
|
||||
#
|
||||
# This drop-in is the ONLY thing preventing both routers from dialling the one
|
||||
# ISP credential at the same time. Do not remove it without reading this.
|
||||
#
|
||||
# pppoe0 is configured identically and ENABLED on both routers, because the
|
||||
# alternative -- `set interfaces pppoe pppoe0 disable` -- unlinks
|
||||
# /etc/ppp/peers/pppoe0 (interfaces_pppoe.py treats `disable` and `delete`
|
||||
# identically), and pppd's options file IS that path. A promotion then had to
|
||||
# re-render it via a full config commit at priority 322, where one unrelated
|
||||
# invalid node fails the whole commit and takes the 10 gig down with it. It also
|
||||
# made op-mode `connect interface pppoe0` unusable, since that refuses when the
|
||||
# peers file is absent.
|
||||
#
|
||||
# With the node enabled, interfaces_pppoe.py's apply() does this on EVERY commit
|
||||
# that touches the pppoe subtree:
|
||||
#
|
||||
# if not is_systemd_service_running('ppp@pppoe0.service') or shutdown_required:
|
||||
# call('systemctl restart ppp@pppoe0.service')
|
||||
#
|
||||
# -- i.e. the backup actively tries to dial whenever anything commits. A
|
||||
# `pulumi up`, a `sim-net-apply.sh apply`, or the boot-time config load are all
|
||||
# that commit. This gate is what makes that a no-op.
|
||||
#
|
||||
# /run is tmpfs, so the gate is shut at boot on both boxes and neither can dial
|
||||
# before VRRP has decided. ppp@.service is already After=vyos-router.service, so
|
||||
# no extra ordering is needed.
|
||||
[Unit]
|
||||
# Both must hold; multiple ConditionPathExists are ANDed.
|
||||
# may-dial -- vrrp-wan-reconcile has blessed this box (a renewed lease)
|
||||
# /etc/ppp/peers -- refuse to start pppd against a missing options file, which
|
||||
# is what produced a restart loop of 47 and counting on
|
||||
# 2026-09-05. A failed Condition is NOT a failure: the job
|
||||
# succeeds, the unit stays inactive, and `systemctl start`
|
||||
# exits 0 -- so callers must check is-active, never rc.
|
||||
ConditionPathExists=/run/vrrp-wan/may-dial
|
||||
ConditionPathExists=/etc/ppp/peers/pppoe0
|
||||
|
||||
# Belt to that brace. The stock unit is Restart=on-failure/RestartSec=5s against
|
||||
# systemd's default StartLimitIntervalSec=10s/Burst=5 -- two restarts per window,
|
||||
# so the limiter can never trip and a doomed pppd retries for ever.
|
||||
StartLimitIntervalSec=600
|
||||
StartLimitBurst=6
|
||||
|
||||
[Service]
|
||||
RestartSec=15
|
||||
# A hung pppd must be resolved inside the failover budget. The stock 90s means a
|
||||
# demoted router could still hold the session while the new master is dialling.
|
||||
# 20s still allows a clean LCP Terminate + PADT in the normal case.
|
||||
TimeoutStopSec=20
|
||||
Reference in New Issue
Block a user