# PPPoE high availability Proven in labsim. **Not applied to production.** ## What it does One consumer ISP account, two routers. The 10 gig lease is bound to a cloned MAC (`f0:9f:c2:12:9b:4f`, the retired USG's) and the Vodafone line to a single credential, so neither may be live on both boxes. The WAN follows VRRP mastership — but the two halves use different control planes, and that is the whole design: | | plane | why | |---|---|---| | `bond0.53` (10 gig) | VyOS **config** (`disable`) | only config can move a MAC | | `pppoe0` (Vodafone) | **systemd** unit gate | see below | ## Why PPPoE cannot live on the config plane `interfaces_pppoe.py` treats `disable` and `delete` identically: both **unlink `/etc/ppp/peers/pppoe0`**, call `PPPoEIf.remove()` (withdrawing the FRR default route) and stop the unit. That path is pppd's own options file (`ExecStart=/usr/sbin/pppd call %I`), so the resting state destroyed exactly what the promotion path needed. `ppp@pppoe0` then restart-looped against the missing file — 47 restarts observed, zero sessions at the access concentrator — and never tripped systemd's limiter, because `RestartSec=5s` against the default 10s/5-burst window is only two restarts per interval. It also made op-mode `connect interface pppoe0` unusable (it refuses without the peers file), and put every failover behind a priority-322 commit where one unrelated invalid node fails the whole thing. ## The gate `pppoe0` is configured identically and **enabled on both** routers, so the peers file always exists. Dialling is gated by `/etc/systemd/system/ppp@pppoe0.service.d/10-vrrp-wan-gate.conf`: ```ini ConditionPathExists=/run/vrrp-wan/may-dial ConditionPathExists=/etc/ppp/peers/pppoe0 StartLimitIntervalSec=600 StartLimitBurst=6 ``` `/run` is tmpfs, so the gate is shut at boot and neither box can dial before VRRP has decided. **This is load-bearing, not a nicety:** with the node enabled, `interfaces_pppoe.py` restarts ppp on *every* commit touching the pppoe subtree when the daemon is not running — so the backup actively tries to dial whenever anything commits (`pulumi up`, a hand commit, the boot-time config load). The gate is the only thing making that a no-op, which is why `vrrp-wan-reconcile` **refuses to bless a box whose drop-in is missing**: `/etc` is per-image, so a VyOS upgrade silently removes the protection, and failing closed turns that into "PPPoE never dials" rather than "both routers dial". `may-dial` is a **lease, not a flag**. `ConditionPathExists` is evaluated at start only — it can prevent a dial, never revoke one. `vrrp-wan-reconcile` renews it every 30s; `vrrp-wan-guard` runs every 5s and only ever revokes, on either "I do not hold the VIP" or "the lease is stale". ## Measured in labsim | | | |---|---| | Clean failover (`force-fault`) | `pppoe0` moves in **20-26s**, reproducible | | 10 gig down → PPPoE | route falls to `pppoe0`; LAN back online in **5s** | | Stale lease | guard hangs up within ~5s | | Missing peers file | `NRestarts=0` — no loop | | Flap holdoff vs live session | session survives a forged 900s holdoff | | Invariant | never more than one router dialled, in any run | On the last row, precisely: the AC *did* report two `simdsl` sessions during the `session-control=disable` run — one live, one orphaned from the router that had just been destroyed, which that policy does not clean up. Only one live router was ever dialled. The count is a proxy for the real invariant and only a valid one while the AC enforces single-session, so it is reported as a `WARN` rather than silenced: an orphaned session still occupies the single slot at a real ISP, and that is exactly what made `deny` take 141–148s. ## Two failure modes found by running it, not by reading it **The flap damper tore down a healthy WAN.** `ppp_dial()` checked the hold-off and returned *before* renewing `may-dial`. That file is a lease the guard expires after `LEASE_TTL`, so tripping the damper stopped the renewal and the guard hung up `pppoe0` **on the master** ~80s later: ``` DIAL FLAP: >=6 attempts in 600s -- holding off 900s GUARD: lease stale (81s > 75s) -- hanging up pppoe0 ``` A damper meant to suppress repeated *dials* was destroying an established session instead. An active session now renews the lease and returns before every other check; everything below only decides whether to start a **new** session. T12 is the regression test. **A missing peers file is silent.** `/etc/ppp/peers/pppoe0` is both pppd's options file and the gate's second condition, and it is only written by a commit that touches the pppoe subtree. Without it systemd logs `skipped because of an unmet condition check` exactly once and then nothing — a router that cannot dial at all looks identical to a healthy backup. `ppp_dial()` now says so on every tick, and distinguishes the two causes: configured-but-not-rendered (re-commit the subtree) versus no `pppoe0` in the config at all. The second cause is the one to watch in production: **a commit that was never `save`d reverts on reboot and takes `pppoe0` with it.** That is exactly how the sim secondary lost its WAN and spent hours looking like an ISP problem. After any hand commit to the pppoe subtree, `save` — or the next reboot produces a standby that can never take over. ## Deploying (not yet done) 1. `sudo /config/vyos-known-good save` on both. 2. `migration/vrrp-wan-install --vip 192.168.1.1 --host vyos@10.0.1.253` then the same for `.252`. Then `--check` on both. **No config change yet** — verify nothing dials. 3. Confirm `/config/wan-secrets` is present and identical on both. 4. **vyos002 first** (the non-master), on its own commit — `interfaces pppoe` is priority 322 and one bad node fails everything: `delete interfaces pppoe pppoe0 disable`, `commit-confirm 10`. 5. Verify vyos002 did **not** dial — check on the wire, not from state: `sudo tcpdump -i bond0.51 -nn pppoed` should show no PADI. Then `confirm`; `save`. 6. vyos001: nothing to change; it already has `pppoe0` enabled. 7. Confirm `vif 53 disable` is in **both** `config.boot`s. 8. **Only now** merge `migration/pulumi-override-pppoe-gated.json` into `kubernetes-deployment` `infra/vyos/subtrees/overrides.json`. It is staged here, unapplied, on purpose: another agent runs `pulumi up` on that repo, so merging it *is* a production change made by someone else at a time you do not choose. Removing `pppoe0 disable` from vyos002 before the gate exists there lets it dial on the next commit and take the single Vodafone session off vyos001. Run `npm run vyos:export && npm run vyos:render` first so the model follows whichever router actually holds the WAN. 9. Add the drop-in re-install to the VyOS image-upgrade runbook. Step 4 is the one that matters most and is worth stopping on. vyos002 has been in **FAULT on all six groups for over three days** — verified again while writing this, alongside vyos001 holding `192.168.1.1` on `bond0.1` and `pppoe0` up on `83.106.5.72`. Until vyos002 reaches BACKUP there is no standby at all: if vyos001 died today nothing would pick up the gateway VIPs. The existing `vrrp-health-check-wan-present` override predicted exactly this in its own reason text — *"with the primary genuinely dead the secondary stays FAULT and nothing holds the gateway. The fix for that is WAN-follows-master, which is a separate change."* This is that change. **Rollback**, from either box: `set interfaces pppoe pppoe0 disable` on both and `rm /run/vrrp-wan/may-dial`. That restores today's behaviour exactly. ## What the sim cannot prove - **Vodafone's `session-control`.** The matrix now brackets it properly. Destroying the master and timing the survivor's session: | policy | takeover | |---|---| | `replace` (accel-ppp default) | 26s / 26s | | **`deny`** (hostile) | **148s / 141s** | | `disable` | 21s / 20s | `deny` is the sizing case: the AC refuses the survivor until its own dead-peer timer frees the dead session, and the poller caught two dial attempts being rejected before one succeeded. `GRACE` is set from that — see `migration/vrrp-wan.conf`. This still cannot tell you which policy Vodafone runs, and account rate-limiting or lockout on repeated dials has no sim analogue at all; the flap damper (6 dials / 600s → 15 min hold-off) exists for that. Treat the 148s as a floor rather than a worst case. These are idle 2-vCPU VMs, and the AC shares an OVS bridge with the routers, so `virsh destroy` removes the port and accel-ppp sees the peer physically vanish. A real BRAS reached over DSL does not learn that our router died — it waits out its own timers, which are longer and not ours to know. Worth knowing how close this came to being missed: until 2026-09-06 the matrix set the policy with `vbash -c 'source script-template; configure; ...; commit'`, which never starts a config session. `commit` failed to stderr, the helper discarded it, and all three iterations ran against the default while printing the mode they were supposedly testing. It reported `deny` at 25s. The real figure is 148s. - **Whether Vodafone honours our LCP Terminate / PADT** on a graceful stop. - **The cloned-MAC lease** — whether the 10 gig ISP re-issues `87.192.101.48` to `f0:9f:c2:12:9b:4f` arriving on a different switch port. That risk belongs to `bond0.53`, not PPPoE, and is the largest untested item in the failover. - **Real dial time and MTU/MSS under load.** PPPoE was proven on the *USG*; VyOS dialling Vodafone has never been done. - **Timing under load.** The sim routers are idle 2-vCPU VMs; commit latency on the VP2440s under kea + BGP + conntrack will be worse, and commit latency is the dominant term in the `bond0.53` half of a failover. ## A model hazard to fix before applying The imported baseline records the **running** state, not the safe one: vyos001 has no `vif 53 disable` (it is master), vyos002 does. So an apply performed while vyos002 held the VIP would enable vyos001's WAN as well, putting the cloned MAC on both boxes. An override must assert `vif 53 disable` on **both** routers, with the reconciler re-enabling whichever holds the VIP. Note the consequence: an apply then briefly disables the current master's 10 gig until the reconciler restores it (≤30s).