Files
lab/migration/PPPOE-HA.md
Michal 5ed0e4888a
Some checks failed
CI/CD / typecheck (push) Failing after 9s
CI/CD / test (push) Failing after 9s
CI/CD / lint (push) Failing after 24s
CI/CD / build (push) Has been skipped
CI/CD / publish-rpm (push) Has been skipped
CI/CD / publish-deb (push) Has been skipped
labsim: both matrices green end to end, numbers reproduced
Final confirming run with GRACE=300 and every fix installed:

  --all   ALL PASS  (T0 T3 T5 T8 T11 T12)
  --hard  ALL PASS  (replace / deny / disable, each policy verified)

Second independent measurement of the hard failover, which is what makes the
figures trustworthy rather than anecdotal:

                run 1   run 2
  replace        26s     26s
  deny          148s    141s
  disable        21s     20s

`deny` sits at ~141-148s across both, so it is the AC's dead-peer behaviour
and not a one-off; GRACE=300 keeps roughly 2x margin.

Corrected an overclaim in PPPOE-HA.md while confirming it: the invariant row
read "AC never showed two simdsl sessions", and under session-control=disable
it did -- one live, one orphaned from the destroyed router. Only ever one LIVE
router dialled, which is the invariant that matters. Said so plainly rather
than leaving a table that reads better than the evidence.
2026-09-06 00:42:01 +01:00

10 KiB
Raw Blame History

PPPoE high availability

Proven in labsim. Not applied to production.

What it does

One consumer ISP account, two routers. The 10 gig lease is bound to a cloned MAC (f0:9f:c2:12:9b:4f, the retired USG's) and the Vodafone line to a single credential, so neither may be live on both boxes. The WAN follows VRRP mastership — but the two halves use different control planes, and that is the whole design:

plane why
bond0.53 (10 gig) VyOS config (disable) only config can move a MAC
pppoe0 (Vodafone) systemd unit gate see below

Why PPPoE cannot live on the config plane

interfaces_pppoe.py treats disable and delete identically: both unlink /etc/ppp/peers/pppoe0, call PPPoEIf.remove() (withdrawing the FRR default route) and stop the unit. That path is pppd's own options file (ExecStart=/usr/sbin/pppd call %I), so the resting state destroyed exactly what the promotion path needed. ppp@pppoe0 then restart-looped against the missing file — 47 restarts observed, zero sessions at the access concentrator — and never tripped systemd's limiter, because RestartSec=5s against the default 10s/5-burst window is only two restarts per interval.

It also made op-mode connect interface pppoe0 unusable (it refuses without the peers file), and put every failover behind a priority-322 commit where one unrelated invalid node fails the whole thing.

The gate

pppoe0 is configured identically and enabled on both routers, so the peers file always exists. Dialling is gated by /etc/systemd/system/ppp@pppoe0.service.d/10-vrrp-wan-gate.conf:

ConditionPathExists=/run/vrrp-wan/may-dial
ConditionPathExists=/etc/ppp/peers/pppoe0
StartLimitIntervalSec=600
StartLimitBurst=6

/run is tmpfs, so the gate is shut at boot and neither box can dial before VRRP has decided. This is load-bearing, not a nicety: with the node enabled, interfaces_pppoe.py restarts ppp on every commit touching the pppoe subtree when the daemon is not running — so the backup actively tries to dial whenever anything commits (pulumi up, a hand commit, the boot-time config load). The gate is the only thing making that a no-op, which is why vrrp-wan-reconcile refuses to bless a box whose drop-in is missing: /etc is per-image, so a VyOS upgrade silently removes the protection, and failing closed turns that into "PPPoE never dials" rather than "both routers dial".

may-dial is a lease, not a flag. ConditionPathExists is evaluated at start only — it can prevent a dial, never revoke one. vrrp-wan-reconcile renews it every 30s; vrrp-wan-guard runs every 5s and only ever revokes, on either "I do not hold the VIP" or "the lease is stale".

Measured in labsim

Clean failover (force-fault) pppoe0 moves in 20-26s, reproducible
10 gig down → PPPoE route falls to pppoe0; LAN back online in 5s
Stale lease guard hangs up within ~5s
Missing peers file NRestarts=0 — no loop
Flap holdoff vs live session session survives a forged 900s holdoff
Invariant never more than one router dialled, in any run

On the last row, precisely: the AC did report two simdsl sessions during the session-control=disable run — one live, one orphaned from the router that had just been destroyed, which that policy does not clean up. Only one live router was ever dialled. The count is a proxy for the real invariant and only a valid one while the AC enforces single-session, so it is reported as a WARN rather than silenced: an orphaned session still occupies the single slot at a real ISP, and that is exactly what made deny take 141148s.

Two failure modes found by running it, not by reading it

The flap damper tore down a healthy WAN. ppp_dial() checked the hold-off and returned before renewing may-dial. That file is a lease the guard expires after LEASE_TTL, so tripping the damper stopped the renewal and the guard hung up pppoe0 on the master ~80s later:

DIAL FLAP: >=6 attempts in 600s -- holding off 900s
GUARD: lease stale (81s > 75s) -- hanging up pppoe0

A damper meant to suppress repeated dials was destroying an established session instead. An active session now renews the lease and returns before every other check; everything below only decides whether to start a new session. T12 is the regression test.

A missing peers file is silent. /etc/ppp/peers/pppoe0 is both pppd's options file and the gate's second condition, and it is only written by a commit that touches the pppoe subtree. Without it systemd logs skipped because of an unmet condition check exactly once and then nothing — a router that cannot dial at all looks identical to a healthy backup. ppp_dial() now says so on every tick, and distinguishes the two causes: configured-but-not-rendered (re-commit the subtree) versus no pppoe0 in the config at all.

The second cause is the one to watch in production: a commit that was never saved reverts on reboot and takes pppoe0 with it. That is exactly how the sim secondary lost its WAN and spent hours looking like an ISP problem. After any hand commit to the pppoe subtree, save — or the next reboot produces a standby that can never take over.

Deploying (not yet done)

  1. sudo /config/vyos-known-good save on both.
  2. migration/vrrp-wan-install --vip 192.168.1.1 --host vyos@10.0.1.253 then the same for .252. Then --check on both. No config change yet — verify nothing dials.
  3. Confirm /config/wan-secrets is present and identical on both.
  4. vyos002 first (the non-master), on its own commit — interfaces pppoe is priority 322 and one bad node fails everything: delete interfaces pppoe pppoe0 disable, commit-confirm 10.
  5. Verify vyos002 did not dial — check on the wire, not from state: sudo tcpdump -i bond0.51 -nn pppoed should show no PADI. Then confirm; save.
  6. vyos001: nothing to change; it already has pppoe0 enabled.
  7. Confirm vif 53 disable is in both config.boots.
  8. Only now merge migration/pulumi-override-pppoe-gated.json into kubernetes-deployment infra/vyos/subtrees/overrides.json. It is staged here, unapplied, on purpose: another agent runs pulumi up on that repo, so merging it is a production change made by someone else at a time you do not choose. Removing pppoe0 disable from vyos002 before the gate exists there lets it dial on the next commit and take the single Vodafone session off vyos001. Run npm run vyos:export && npm run vyos:render first so the model follows whichever router actually holds the WAN.
  9. Add the drop-in re-install to the VyOS image-upgrade runbook.

Step 4 is the one that matters most and is worth stopping on. vyos002 has been in FAULT on all six groups for over three days — verified again while writing this, alongside vyos001 holding 192.168.1.1 on bond0.1 and pppoe0 up on 83.106.5.72. Until vyos002 reaches BACKUP there is no standby at all: if vyos001 died today nothing would pick up the gateway VIPs. The existing vrrp-health-check-wan-present override predicted exactly this in its own reason text — "with the primary genuinely dead the secondary stays FAULT and nothing holds the gateway. The fix for that is WAN-follows-master, which is a separate change." This is that change.

Rollback, from either box: set interfaces pppoe pppoe0 disable on both and rm /run/vrrp-wan/may-dial. That restores today's behaviour exactly.

What the sim cannot prove

  • Vodafone's session-control. The matrix now brackets it properly. Destroying the master and timing the survivor's session:

    policy takeover
    replace (accel-ppp default) 26s / 26s
    deny (hostile) 148s / 141s
    disable 21s / 20s

    deny is the sizing case: the AC refuses the survivor until its own dead-peer timer frees the dead session, and the poller caught two dial attempts being rejected before one succeeded. GRACE is set from that — see migration/vrrp-wan.conf. This still cannot tell you which policy Vodafone runs, and account rate-limiting or lockout on repeated dials has no sim analogue at all; the flap damper (6 dials / 600s → 15 min hold-off) exists for that.

    Treat the 148s as a floor rather than a worst case. These are idle 2-vCPU VMs, and the AC shares an OVS bridge with the routers, so virsh destroy removes the port and accel-ppp sees the peer physically vanish. A real BRAS reached over DSL does not learn that our router died — it waits out its own timers, which are longer and not ours to know.

    Worth knowing how close this came to being missed: until 2026-09-06 the matrix set the policy with vbash -c 'source script-template; configure; ...; commit', which never starts a config session. commit failed to stderr, the helper discarded it, and all three iterations ran against the default while printing the mode they were supposedly testing. It reported deny at 25s. The real figure is 148s.

  • Whether Vodafone honours our LCP Terminate / PADT on a graceful stop.

  • The cloned-MAC lease — whether the 10 gig ISP re-issues 87.192.101.48 to f0:9f:c2:12:9b:4f arriving on a different switch port. That risk belongs to bond0.53, not PPPoE, and is the largest untested item in the failover.

  • Real dial time and MTU/MSS under load. PPPoE was proven on the USG; VyOS dialling Vodafone has never been done.

  • Timing under load. The sim routers are idle 2-vCPU VMs; commit latency on the VP2440s under kea + BGP + conntrack will be worse, and commit latency is the dominant term in the bond0.53 half of a failover.

A model hazard to fix before applying

The imported baseline records the running state, not the safe one: vyos001 has no vif 53 disable (it is master), vyos002 does. So an apply performed while vyos002 held the VIP would enable vyos001's WAN as well, putting the cloned MAC on both boxes. An override must assert vif 53 disable on both routers, with the reconciler re-enabling whichever holds the VIP. Note the consequence: an apply then briefly disables the current master's 10 gig until the reconciler restores it (≤30s).