Files
lab/migration/PPPOE-HA.md
Michal d626750010
Some checks failed
CI/CD / lint (push) Failing after 9s
CI/CD / typecheck (push) Failing after 9s
CI/CD / test (push) Failing after 9s
CI/CD / build (push) Has been skipped
CI/CD / publish-rpm (push) Has been skipped
CI/CD / publish-deb (push) Has been skipped
labsim: hard failover and reboot safety hold; runbook for the production apply
Hard failover proven by destroying the master outright (virsh destroy -- no PADT,
the case a graceful stop cannot cover): the survivor took the VIPs, dialled, and
the access concentrator showed exactly ONE simdsl session from its MAC for the
whole seven minutes. When the destroyed router came back it did NOT dial --
ConditionResult=no, ActiveState=inactive, NRestarts=0, pppoe0 absent, may_dial=no
-- and the AC session count stayed at 1. That is the gate doing the one job it
exists for, on the path that previously had no protection at all.

Two more harness bugs of the same family as the last three, both of which
reported a working system as broken:

  - ppp_on() returned an EMPTY string for an unreachable router, and `[ "" = 0 ]`
    is false, so the hard-failover wait sat for its full timeout waiting for a
    DESTROYED box to report zero -- long after the survivor had taken over
    correctly. Absent now means 0.
  - a VM restart recreates its taps under new names and the OVS bond keeps the
    old ones: lacp dies, VLAN 1 goes with it, and the box returns reachable on
    some VLANs and not others. That looked exactly like a failed failover. It is
    the same stale-membership fault ovs_bond_router already detects, but nothing
    ran it after a restart; the harness now does.

migration/PPPOE-HA.md is the runbook: what the design is, why `disable` cannot
work, the measured numbers, the deploy order (vyos002 first, on its own commit,
verified on the wire with tcpdump rather than from state), and the one-line
rollback.

It also records a model hazard found while writing it. The imported baseline
captures the RUNNING state, not the safe one -- vyos001 has no `vif 53 disable`
because it happens to be master, vyos002 does. An apply performed while vyos002
held the VIP would therefore enable vyos001's WAN too, putting the cloned MAC on
both boxes. An override must assert `vif 53 disable` on BOTH, accepting that an
apply then briefly disables the current master's 10 gig until the reconciler
restores it.

Still not applied to production.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01DMVzWZgiKW2wquf5z8S1yH
2026-09-05 19:21:13 +01:00

5.7 KiB

PPPoE high availability

Proven in labsim. Not applied to production.

What it does

One consumer ISP account, two routers. The 10 gig lease is bound to a cloned MAC (f0:9f:c2:12:9b:4f, the retired USG's) and the Vodafone line to a single credential, so neither may be live on both boxes. The WAN follows VRRP mastership — but the two halves use different control planes, and that is the whole design:

plane why
bond0.53 (10 gig) VyOS config (disable) only config can move a MAC
pppoe0 (Vodafone) systemd unit gate see below

Why PPPoE cannot live on the config plane

interfaces_pppoe.py treats disable and delete identically: both unlink /etc/ppp/peers/pppoe0, call PPPoEIf.remove() (withdrawing the FRR default route) and stop the unit. That path is pppd's own options file (ExecStart=/usr/sbin/pppd call %I), so the resting state destroyed exactly what the promotion path needed. ppp@pppoe0 then restart-looped against the missing file — 47 restarts observed, zero sessions at the access concentrator — and never tripped systemd's limiter, because RestartSec=5s against the default 10s/5-burst window is only two restarts per interval.

It also made op-mode connect interface pppoe0 unusable (it refuses without the peers file), and put every failover behind a priority-322 commit where one unrelated invalid node fails the whole thing.

The gate

pppoe0 is configured identically and enabled on both routers, so the peers file always exists. Dialling is gated by /etc/systemd/system/ppp@pppoe0.service.d/10-vrrp-wan-gate.conf:

ConditionPathExists=/run/vrrp-wan/may-dial
ConditionPathExists=/etc/ppp/peers/pppoe0
StartLimitIntervalSec=600
StartLimitBurst=6

/run is tmpfs, so the gate is shut at boot and neither box can dial before VRRP has decided. This is load-bearing, not a nicety: with the node enabled, interfaces_pppoe.py restarts ppp on every commit touching the pppoe subtree when the daemon is not running — so the backup actively tries to dial whenever anything commits (pulumi up, a hand commit, the boot-time config load). The gate is the only thing making that a no-op, which is why vrrp-wan-reconcile refuses to bless a box whose drop-in is missing: /etc is per-image, so a VyOS upgrade silently removes the protection, and failing closed turns that into "PPPoE never dials" rather than "both routers dial".

may-dial is a lease, not a flag. ConditionPathExists is evaluated at start only — it can prevent a dial, never revoke one. vrrp-wan-reconcile renews it every 30s; vrrp-wan-guard runs every 5s and only ever revokes, on either "I do not hold the VIP" or "the lease is stale".

Measured in labsim

Clean failover (force-fault) pppoe0 moves in 26s, reproducible
10 gig down → PPPoE route falls to pppoe0; LAN back online in 5s
Stale lease guard hangs up within ~5s
Missing peers file NRestarts=0 — no loop
Invariant AC never showed two simdsl sessions

Deploying (not yet done)

  1. sudo /config/vyos-known-good save on both.
  2. migration/vrrp-wan-install --vip 192.168.1.1 --host vyos@10.0.1.253 then the same for .252. Then --check on both. No config change yet — verify nothing dials.
  3. Confirm /config/wan-secrets is present and identical on both.
  4. vyos002 first (the non-master), on its own commit — interfaces pppoe is priority 322 and one bad node fails everything: delete interfaces pppoe pppoe0 disable, commit-confirm 10.
  5. Verify vyos002 did not dial — check on the wire, not from state: sudo tcpdump -i bond0.51 -nn pppoed should show no PADI. Then confirm; save.
  6. vyos001: nothing to change; it already has pppoe0 enabled.
  7. Confirm vif 53 disable is in both config.boots.
  8. Add the drop-in re-install to the VyOS image-upgrade runbook.

Rollback, from either box: set interfaces pppoe pppoe0 disable on both and rm /run/vrrp-wan/may-dial. That restores today's behaviour exactly.

What the sim cannot prove

  • Vodafone's session-control. The sim's accel-ppp defaults to replace, so a new auth kills the old session immediately. A real BRAS may deny and hold the session for its own dead-peer timer. The matrix runs all three modes to bracket the risk; it cannot tell you which one you will meet, and account rate-limiting or lockout on repeated dials has no sim analogue at all. The flap damper (6 dials / 600s → 15 min hold-off) exists for that.
  • Whether Vodafone honours our LCP Terminate / PADT on a graceful stop.
  • The cloned-MAC lease — whether the 10 gig ISP re-issues 87.192.101.48 to f0:9f:c2:12:9b:4f arriving on a different switch port. That risk belongs to bond0.53, not PPPoE, and is the largest untested item in the failover.
  • Real dial time and MTU/MSS under load. PPPoE was proven on the USG; VyOS dialling Vodafone has never been done.
  • Timing under load. The sim routers are idle 2-vCPU VMs; commit latency on the VP2440s under kea + BGP + conntrack will be worse, and commit latency is the dominant term in the bond0.53 half of a failover.

A model hazard to fix before applying

The imported baseline records the running state, not the safe one: vyos001 has no vif 53 disable (it is master), vyos002 does. So an apply performed while vyos002 held the VIP would enable vyos001's WAN as well, putting the cloned MAC on both boxes. An override must assert vif 53 disable on both routers, with the reconciler re-enabling whichever holds the VIP. Note the consequence: an apply then briefly disables the current master's 10 gig until the reconciler restores it (≤30s).