Controlled failover in production: takeover 52s, failback 36s, ~88s of interruption across two deliberate transitions. The cloned-MAC lease TRANSFERS. That was named in this file as "the largest untested item in any failover" -- whether the 10 gig ISP would re-issue 87.192.101.48 to f0:9f:c2:12:9b:4f arriving on a different switch port. It did, same address, inside the takeover window. Moved out of "what the sim cannot prove" rather than left there contradicting the evidence. VyOS dialling Vodafone also worked from both routers, which had never been done -- PPPoE was only ever proven on the USG. Vodafone did not refuse either re-dial, so its session-control behaves like `replace`, not the hostile `deny`. GRACE stays at 300 anyway: one drill on one evening is not the ISP's policy under all conditions. Worth knowing for anything added later: Vodafone hands out a DIFFERENT IPv4 on every dial (83.106.5.72 -> 90.251.153.180 -> 90.251.142.103). Nothing may be pinned to the pppoe0 address. Checked the HE IPv6 tunnel specifically, since it carries a hardcoded source-address -- it is pinned to 87.192.101.48, which is the 10 gig and stable across failover, so tun0 was untouched and IPv6 stayed up at 15.5ms.
14 KiB
PPPoE high availability
Proven in labsim. Deployed to production 2026-09-06 — mechanism on both
routers, pppoe0 disable removed from vyos002, vyos002 out of FAULT and in
BACKUP. The Pulumi override (step 8) is still staged, not merged.
What it does
One consumer ISP account, two routers. The 10 gig lease is bound to a cloned MAC
(f0:9f:c2:12:9b:4f, the retired USG's) and the Vodafone line to a single
credential, so neither may be live on both boxes. The WAN follows VRRP
mastership — but the two halves use different control planes, and that is the
whole design:
| plane | why | |
|---|---|---|
bond0.53 (10 gig) |
VyOS config (disable) |
only config can move a MAC |
pppoe0 (Vodafone) |
systemd unit gate | see below |
Why PPPoE cannot live on the config plane
interfaces_pppoe.py treats disable and delete identically: both unlink
/etc/ppp/peers/pppoe0, call PPPoEIf.remove() (withdrawing the FRR default
route) and stop the unit. That path is pppd's own options file
(ExecStart=/usr/sbin/pppd call %I), so the resting state destroyed exactly what
the promotion path needed. ppp@pppoe0 then restart-looped against the missing
file — 47 restarts observed, zero sessions at the access concentrator — and
never tripped systemd's limiter, because RestartSec=5s against the default
10s/5-burst window is only two restarts per interval.
It also made op-mode connect interface pppoe0 unusable (it refuses without the
peers file), and put every failover behind a priority-322 commit where one
unrelated invalid node fails the whole thing.
The gate
pppoe0 is configured identically and enabled on both routers, so the peers
file always exists. Dialling is gated by
/etc/systemd/system/ppp@pppoe0.service.d/10-vrrp-wan-gate.conf:
ConditionPathExists=/run/vrrp-wan/may-dial
ConditionPathExists=/etc/ppp/peers/pppoe0
StartLimitIntervalSec=600
StartLimitBurst=6
/run is tmpfs, so the gate is shut at boot and neither box can dial before VRRP
has decided. This is load-bearing, not a nicety: with the node enabled,
interfaces_pppoe.py restarts ppp on every commit touching the pppoe subtree
when the daemon is not running — so the backup actively tries to dial whenever
anything commits (pulumi up, a hand commit, the boot-time config load). The
gate is the only thing making that a no-op, which is why vrrp-wan-reconcile
refuses to bless a box whose drop-in is missing: /etc is per-image, so a
VyOS upgrade silently removes the protection, and failing closed turns that into
"PPPoE never dials" rather than "both routers dial".
may-dial is a lease, not a flag. ConditionPathExists is evaluated at
start only — it can prevent a dial, never revoke one. vrrp-wan-reconcile
renews it every 30s; vrrp-wan-guard runs every 5s and only ever revokes, on
either "I do not hold the VIP" or "the lease is stale".
Measured in labsim
Clean failover (force-fault) |
pppoe0 moves in 20-26s, reproducible |
| 10 gig down → PPPoE | route falls to pppoe0; LAN back online in 5s |
| Stale lease | guard hangs up within ~5s |
| Missing peers file | NRestarts=0 — no loop |
| Flap holdoff vs live session | session survives a forged 900s holdoff |
| Invariant | never more than one router dialled, in any run |
On the last row, precisely: the AC did report two simdsl sessions during the
session-control=disable run — one live, one orphaned from the router that had
just been destroyed, which that policy does not clean up. Only one live router
was ever dialled. The count is a proxy for the real invariant and only a valid
one while the AC enforces single-session, so it is reported as a WARN rather
than silenced: an orphaned session still occupies the single slot at a real ISP,
and that is exactly what made deny take 141–148s.
Two failure modes found by running it, not by reading it
The flap damper tore down a healthy WAN. ppp_dial() checked the hold-off
and returned before renewing may-dial. That file is a lease the guard
expires after LEASE_TTL, so tripping the damper stopped the renewal and the
guard hung up pppoe0 on the master ~80s later:
DIAL FLAP: >=6 attempts in 600s -- holding off 900s
GUARD: lease stale (81s > 75s) -- hanging up pppoe0
A damper meant to suppress repeated dials was destroying an established session instead. An active session now renews the lease and returns before every other check; everything below only decides whether to start a new session. T12 is the regression test.
A missing peers file is silent. /etc/ppp/peers/pppoe0 is both pppd's
options file and the gate's second condition, and it is only written by a commit
that touches the pppoe subtree. Without it systemd logs
skipped because of an unmet condition check exactly once and then nothing —
a router that cannot dial at all looks identical to a healthy backup.
ppp_dial() now says so on every tick, and distinguishes the two causes:
configured-but-not-rendered (re-commit the subtree) versus no pppoe0 in the
config at all.
The second cause is the one to watch in production: a commit that was never
saved reverts on reboot and takes pppoe0 with it. That is exactly how the
sim secondary lost its WAN and spent hours looking like an ISP problem. After
any hand commit to the pppoe subtree, save — or the next reboot produces a
standby that can never take over.
Deploying — steps 1–7 done 2026-09-06, 8–9 outstanding
sudo /config/vyos-known-good saveon both.migration/vrrp-wan-install --vip 192.168.1.1 --host vyos@10.0.1.253then the same for.252. Then--checkon both. No config change yet — verify nothing dials.- Confirm
/config/wan-secretsis present and identical on both. - vyos002 first (the non-master), on its own commit —
interfaces pppoeis priority 322 and one bad node fails everything:delete interfaces pppoe pppoe0 disable,commit-confirm 10. - Verify vyos002 did not dial — check on the wire, not from state:
sudo tcpdump -i bond0.51 -nn pppoedshould show no PADI. Thenconfirm;save. - vyos001: nothing to change; it already has
pppoe0enabled. - Confirm
vif 53 disableis in bothconfig.boots. It is not. vyos002 has it; vyos001 does not (itsconfig.bootdates from 2026-09-02 and predates this work). See "Outstanding: vyos001's config.boot" below — it self-heals, but it should still be fixed. - Only now merge
migration/pulumi-override-pppoe-gated.jsonintokubernetes-deploymentinfra/vyos/subtrees/overrides.json. It is staged here, unapplied, on purpose: another agent runspulumi upon that repo, so merging it is a production change made by someone else at a time you do not choose. Removingpppoe0 disablefrom vyos002 before the gate exists there lets it dial on the next commit and take the single Vodafone session off vyos001. Runnpm run vyos:export && npm run vyos:renderfirst so the model follows whichever router actually holds the WAN. - Add the drop-in re-install to the VyOS image-upgrade runbook.
Step 4 is the one that matters most and is worth stopping on. vyos002 has been
in FAULT on all six groups for over three days — verified again while
writing this, alongside vyos001 holding 192.168.1.1 on bond0.1 and
pppoe0 up on 83.106.5.72. Until vyos002 reaches BACKUP there is no standby
at all: if vyos001 died today nothing would pick up the gateway VIPs. The
existing vrrp-health-check-wan-present override predicted exactly this in its
own reason text — "with the primary genuinely dead the secondary stays FAULT
and nothing holds the gateway. The fix for that is WAN-follows-master, which is
a separate change." This is that change.
Proven in production — controlled drill, 2026-09-06
migration/wan-drill force-faulted vyos001 and timed a real failover:
TAKEOVER OK: vyos002 held the VIP and reached the internet in 52s
FAILBACK OK: 36s
vyos002 took the VIPs at t+18s and had both WANs by t+52s. Failback put vyos001 back with a WAN in 36s. Total interruption ≈ 88s across two deliberate transitions.
The cloned-MAC lease transfers. This was the largest untested item in the
whole design — whether the 10 gig ISP would re-issue 87.192.101.48 to
f0:9f:c2:12:9b:4f arriving on a different switch port. It did, same address,
within the takeover window. That risk is now closed.
Vodafone did not refuse the re-dial, so its session-control behaves like
replace rather than the hostile deny. GRACE=300 was sized against the
sim's 148s deny case and is therefore comfortable — but it should stay where
it is, because one drill on one evening does not establish the ISP's policy
under all conditions.
Vodafone hands out a different IPv4 on every dial: 83.106.5.72 →
90.251.153.180 (vyos002) → 90.251.142.103 (vyos001, after failback).
Nothing may be pinned to the PPPoE address. Checked: the HE IPv6 tunnel is
pinned to 87.192.101.48, which is the 10 gig (bond0.53) and stable
across failover, so tun0 survived untouched and IPv6 stayed up at 15.5ms.
Anything added later that hardcodes a WAN IP must use the 10 gig one, not
pppoe0's.
Production takeover (52s) is about twice the sim's replace figure (26s), which
is the expected direction: the VP2440s commit under kea, BGP and conntrack while
the sim routers are idle.
Outstanding: vyos001's config.boot lacks vif 53 disable
The convention is that both config.boots hold vif 53 disable, so a
reboot in any order comes up unable to claim the cloned MAC and the reconciler
then enables it on whichever box holds the VIP. vyos002 satisfies this.
vyos001 does not — its config.boot is from 2026-09-02 and predates this work,
so this is pre-existing rather than introduced here.
What it now exposes: if vyos001 reboots while vyos002 is master and holding the
10 gig, vyos001 comes up with bond0.53 enabled and the same cloned MAC
f0:9f:c2:12:9b:4f is briefly live on both boxes. This mattered less before
today, because vyos002 was stuck in FAULT and could never be master. It can be
now.
It self-heals within 30s: vyos001 comes up BACKUP (priority 200 but
no-preempt), the reconciler sees "not master but bond0.53 enabled", and
commits disable. So the exposure is a ≤30s duplicate MAC on the WAN segment,
not a permanent split.
Fixing it properly is not free, which is why it is listed rather than done:
config.boot can only be written by save, and save writes the running
state — so you must set … vif 53 disable, commit (which takes the 10 gig
down, dropping the default route onto pppoe0), save, then delete … disable
and commit without saving. That is a brief, deliberate WAN interruption on
the live master and belongs in a maintenance window. Hand-editing config.boot
avoids the blip but risks an unbootable router, which is a worse trade.
Rollback, from either box: set interfaces pppoe pppoe0 disable on both and
rm /run/vrrp-wan/may-dial. That restores today's behaviour exactly.
What the sim cannot prove
-
Vodafone's
session-control. The matrix now brackets it properly. Destroying the master and timing the survivor's session:policy takeover replace(accel-ppp default)26s / 26s deny(hostile)148s / 141s disable21s / 20s denyis the sizing case: the AC refuses the survivor until its own dead-peer timer frees the dead session, and the poller caught two dial attempts being rejected before one succeeded.GRACEis set from that — seemigration/vrrp-wan.conf. This still cannot tell you which policy Vodafone runs, and account rate-limiting or lockout on repeated dials has no sim analogue at all; the flap damper (6 dials / 600s → 15 min hold-off) exists for that.Treat the 148s as a floor rather than a worst case. These are idle 2-vCPU VMs, and the AC shares an OVS bridge with the routers, so
virsh destroyremoves the port and accel-ppp sees the peer physically vanish. A real BRAS reached over DSL does not learn that our router died — it waits out its own timers, which are longer and not ours to know.Worth knowing how close this came to being missed: until 2026-09-06 the matrix set the policy with
vbash -c 'source script-template; configure; ...; commit', which never starts a config session.commitfailed to stderr, the helper discarded it, and all three iterations ran against the default while printing the mode they were supposedly testing. It reporteddenyat 25s. The real figure is 148s. -
Whether Vodafone honours our LCP Terminate / PADT on a graceful stop.
-
The cloned-MAC lease.Answered 2026-09-06: the drill moved it and the ISP re-issued87.192.101.48tof0:9f:c2:12:9b:4fon vyos002's port within the takeover window. See "Proven in production" above. -
Whether VyOS can dial Vodafone at all.Answered 2026-09-06: both routers dialled successfully during the drill. MTU/MSS under sustained load is still unmeasured, and Vodafone hands out a different IPv4 every dial. -
Timing under load. The sim routers are idle 2-vCPU VMs; commit latency on the VP2440s under kea + BGP + conntrack will be worse, and commit latency is the dominant term in the
bond0.53half of a failover.
A model hazard to fix before applying
The imported baseline records the running state, not the safe one: vyos001
has no vif 53 disable (it is master), vyos002 does. So an apply performed while
vyos002 held the VIP would enable vyos001's WAN as well, putting the cloned MAC
on both boxes. An override must assert vif 53 disable on both routers, with
the reconciler re-enabling whichever holds the VIP. Note the consequence: an
apply then briefly disables the current master's 10 gig until the reconciler
restores it (≤30s).