Some checks failed
Final confirming run with GRACE=300 and every fix installed:
--all ALL PASS (T0 T3 T5 T8 T11 T12)
--hard ALL PASS (replace / deny / disable, each policy verified)
Second independent measurement of the hard failover, which is what makes the
figures trustworthy rather than anecdotal:
run 1 run 2
replace 26s 26s
deny 148s 141s
disable 21s 20s
`deny` sits at ~141-148s across both, so it is the AC's dead-peer behaviour
and not a one-off; GRACE=300 keeps roughly 2x margin.
Corrected an overclaim in PPPOE-HA.md while confirming it: the invariant row
read "AC never showed two simdsl sessions", and under session-control=disable
it did -- one live, one orphaned from the destroyed router. Only ever one LIVE
router dialled, which is the invariant that matters. Said so plainly rather
than leaving a table that reads better than the evidence.
199 lines
10 KiB
Markdown
199 lines
10 KiB
Markdown
# PPPoE high availability
|
||
|
||
Proven in labsim. **Not applied to production.**
|
||
|
||
## What it does
|
||
|
||
One consumer ISP account, two routers. The 10 gig lease is bound to a cloned MAC
|
||
(`f0:9f:c2:12:9b:4f`, the retired USG's) and the Vodafone line to a single
|
||
credential, so neither may be live on both boxes. The WAN follows VRRP
|
||
mastership — but the two halves use different control planes, and that is the
|
||
whole design:
|
||
|
||
| | plane | why |
|
||
|---|---|---|
|
||
| `bond0.53` (10 gig) | VyOS **config** (`disable`) | only config can move a MAC |
|
||
| `pppoe0` (Vodafone) | **systemd** unit gate | see below |
|
||
|
||
## Why PPPoE cannot live on the config plane
|
||
|
||
`interfaces_pppoe.py` treats `disable` and `delete` identically: both **unlink
|
||
`/etc/ppp/peers/pppoe0`**, call `PPPoEIf.remove()` (withdrawing the FRR default
|
||
route) and stop the unit. That path is pppd's own options file
|
||
(`ExecStart=/usr/sbin/pppd call %I`), so the resting state destroyed exactly what
|
||
the promotion path needed. `ppp@pppoe0` then restart-looped against the missing
|
||
file — 47 restarts observed, zero sessions at the access concentrator — and
|
||
never tripped systemd's limiter, because `RestartSec=5s` against the default
|
||
10s/5-burst window is only two restarts per interval.
|
||
|
||
It also made op-mode `connect interface pppoe0` unusable (it refuses without the
|
||
peers file), and put every failover behind a priority-322 commit where one
|
||
unrelated invalid node fails the whole thing.
|
||
|
||
## The gate
|
||
|
||
`pppoe0` is configured identically and **enabled on both** routers, so the peers
|
||
file always exists. Dialling is gated by
|
||
`/etc/systemd/system/ppp@pppoe0.service.d/10-vrrp-wan-gate.conf`:
|
||
|
||
```ini
|
||
ConditionPathExists=/run/vrrp-wan/may-dial
|
||
ConditionPathExists=/etc/ppp/peers/pppoe0
|
||
StartLimitIntervalSec=600
|
||
StartLimitBurst=6
|
||
```
|
||
|
||
`/run` is tmpfs, so the gate is shut at boot and neither box can dial before VRRP
|
||
has decided. **This is load-bearing, not a nicety:** with the node enabled,
|
||
`interfaces_pppoe.py` restarts ppp on *every* commit touching the pppoe subtree
|
||
when the daemon is not running — so the backup actively tries to dial whenever
|
||
anything commits (`pulumi up`, a hand commit, the boot-time config load). The
|
||
gate is the only thing making that a no-op, which is why `vrrp-wan-reconcile`
|
||
**refuses to bless a box whose drop-in is missing**: `/etc` is per-image, so a
|
||
VyOS upgrade silently removes the protection, and failing closed turns that into
|
||
"PPPoE never dials" rather than "both routers dial".
|
||
|
||
`may-dial` is a **lease, not a flag**. `ConditionPathExists` is evaluated at
|
||
start only — it can prevent a dial, never revoke one. `vrrp-wan-reconcile`
|
||
renews it every 30s; `vrrp-wan-guard` runs every 5s and only ever revokes, on
|
||
either "I do not hold the VIP" or "the lease is stale".
|
||
|
||
## Measured in labsim
|
||
|
||
| | |
|
||
|---|---|
|
||
| Clean failover (`force-fault`) | `pppoe0` moves in **20-26s**, reproducible |
|
||
| 10 gig down → PPPoE | route falls to `pppoe0`; LAN back online in **5s** |
|
||
| Stale lease | guard hangs up within ~5s |
|
||
| Missing peers file | `NRestarts=0` — no loop |
|
||
| Flap holdoff vs live session | session survives a forged 900s holdoff |
|
||
| Invariant | never more than one router dialled, in any run |
|
||
|
||
On the last row, precisely: the AC *did* report two `simdsl` sessions during the
|
||
`session-control=disable` run — one live, one orphaned from the router that had
|
||
just been destroyed, which that policy does not clean up. Only one live router
|
||
was ever dialled. The count is a proxy for the real invariant and only a valid
|
||
one while the AC enforces single-session, so it is reported as a `WARN` rather
|
||
than silenced: an orphaned session still occupies the single slot at a real ISP,
|
||
and that is exactly what made `deny` take 141–148s.
|
||
|
||
## Two failure modes found by running it, not by reading it
|
||
|
||
**The flap damper tore down a healthy WAN.** `ppp_dial()` checked the hold-off
|
||
and returned *before* renewing `may-dial`. That file is a lease the guard
|
||
expires after `LEASE_TTL`, so tripping the damper stopped the renewal and the
|
||
guard hung up `pppoe0` **on the master** ~80s later:
|
||
|
||
```
|
||
DIAL FLAP: >=6 attempts in 600s -- holding off 900s
|
||
GUARD: lease stale (81s > 75s) -- hanging up pppoe0
|
||
```
|
||
|
||
A damper meant to suppress repeated *dials* was destroying an established
|
||
session instead. An active session now renews the lease and returns before
|
||
every other check; everything below only decides whether to start a **new**
|
||
session. T12 is the regression test.
|
||
|
||
**A missing peers file is silent.** `/etc/ppp/peers/pppoe0` is both pppd's
|
||
options file and the gate's second condition, and it is only written by a commit
|
||
that touches the pppoe subtree. Without it systemd logs
|
||
`skipped because of an unmet condition check` exactly once and then nothing —
|
||
a router that cannot dial at all looks identical to a healthy backup.
|
||
`ppp_dial()` now says so on every tick, and distinguishes the two causes:
|
||
configured-but-not-rendered (re-commit the subtree) versus no `pppoe0` in the
|
||
config at all.
|
||
|
||
The second cause is the one to watch in production: **a commit that was never
|
||
`save`d reverts on reboot and takes `pppoe0` with it.** That is exactly how the
|
||
sim secondary lost its WAN and spent hours looking like an ISP problem. After
|
||
any hand commit to the pppoe subtree, `save` — or the next reboot produces a
|
||
standby that can never take over.
|
||
|
||
## Deploying (not yet done)
|
||
|
||
1. `sudo /config/vyos-known-good save` on both.
|
||
2. `migration/vrrp-wan-install --vip 192.168.1.1 --host vyos@10.0.1.253` then the
|
||
same for `.252`. Then `--check` on both. **No config change yet** — verify
|
||
nothing dials.
|
||
3. Confirm `/config/wan-secrets` is present and identical on both.
|
||
4. **vyos002 first** (the non-master), on its own commit — `interfaces pppoe` is
|
||
priority 322 and one bad node fails everything:
|
||
`delete interfaces pppoe pppoe0 disable`, `commit-confirm 10`.
|
||
5. Verify vyos002 did **not** dial — check on the wire, not from state:
|
||
`sudo tcpdump -i bond0.51 -nn pppoed` should show no PADI. Then `confirm`; `save`.
|
||
6. vyos001: nothing to change; it already has `pppoe0` enabled.
|
||
7. Confirm `vif 53 disable` is in **both** `config.boot`s.
|
||
8. **Only now** merge `migration/pulumi-override-pppoe-gated.json` into
|
||
`kubernetes-deployment` `infra/vyos/subtrees/overrides.json`. It is staged
|
||
here, unapplied, on purpose: another agent runs `pulumi up` on that repo, so
|
||
merging it *is* a production change made by someone else at a time you do not
|
||
choose. Removing `pppoe0 disable` from vyos002 before the gate exists there
|
||
lets it dial on the next commit and take the single Vodafone session off
|
||
vyos001. Run `npm run vyos:export && npm run vyos:render` first so the model
|
||
follows whichever router actually holds the WAN.
|
||
9. Add the drop-in re-install to the VyOS image-upgrade runbook.
|
||
|
||
Step 4 is the one that matters most and is worth stopping on. vyos002 has been
|
||
in **FAULT on all six groups for over three days** — verified again while
|
||
writing this, alongside vyos001 holding `192.168.1.1` on `bond0.1` and
|
||
`pppoe0` up on `83.106.5.72`. Until vyos002 reaches BACKUP there is no standby
|
||
at all: if vyos001 died today nothing would pick up the gateway VIPs. The
|
||
existing `vrrp-health-check-wan-present` override predicted exactly this in its
|
||
own reason text — *"with the primary genuinely dead the secondary stays FAULT
|
||
and nothing holds the gateway. The fix for that is WAN-follows-master, which is
|
||
a separate change."* This is that change.
|
||
|
||
**Rollback**, from either box: `set interfaces pppoe pppoe0 disable` on both and
|
||
`rm /run/vrrp-wan/may-dial`. That restores today's behaviour exactly.
|
||
|
||
## What the sim cannot prove
|
||
|
||
- **Vodafone's `session-control`.** The matrix now brackets it properly.
|
||
Destroying the master and timing the survivor's session:
|
||
|
||
| policy | takeover |
|
||
|---|---|
|
||
| `replace` (accel-ppp default) | 26s / 26s |
|
||
| **`deny`** (hostile) | **148s / 141s** |
|
||
| `disable` | 21s / 20s |
|
||
|
||
`deny` is the sizing case: the AC refuses the survivor until its own
|
||
dead-peer timer frees the dead session, and the poller caught two dial
|
||
attempts being rejected before one succeeded. `GRACE` is set from that — see
|
||
`migration/vrrp-wan.conf`. This still cannot tell you which policy Vodafone
|
||
runs, and account rate-limiting or lockout on repeated dials has no sim
|
||
analogue at all; the flap damper (6 dials / 600s → 15 min hold-off) exists
|
||
for that.
|
||
|
||
Treat the 148s as a floor rather than a worst case. These are idle 2-vCPU
|
||
VMs, and the AC shares an OVS bridge with the routers, so `virsh destroy`
|
||
removes the port and accel-ppp sees the peer physically vanish. A real BRAS
|
||
reached over DSL does not learn that our router died — it waits out its own
|
||
timers, which are longer and not ours to know.
|
||
|
||
Worth knowing how close this came to being missed: until 2026-09-06 the
|
||
matrix set the policy with `vbash -c 'source script-template; configure;
|
||
...; commit'`, which never starts a config session. `commit` failed to
|
||
stderr, the helper discarded it, and all three iterations ran against the
|
||
default while printing the mode they were supposedly testing. It reported
|
||
`deny` at 25s. The real figure is 148s.
|
||
- **Whether Vodafone honours our LCP Terminate / PADT** on a graceful stop.
|
||
- **The cloned-MAC lease** — whether the 10 gig ISP re-issues `87.192.101.48` to
|
||
`f0:9f:c2:12:9b:4f` arriving on a different switch port. That risk belongs to
|
||
`bond0.53`, not PPPoE, and is the largest untested item in the failover.
|
||
- **Real dial time and MTU/MSS under load.** PPPoE was proven on the *USG*;
|
||
VyOS dialling Vodafone has never been done.
|
||
- **Timing under load.** The sim routers are idle 2-vCPU VMs; commit latency on
|
||
the VP2440s under kea + BGP + conntrack will be worse, and commit latency is
|
||
the dominant term in the `bond0.53` half of a failover.
|
||
|
||
## A model hazard to fix before applying
|
||
|
||
The imported baseline records the **running** state, not the safe one: vyos001
|
||
has no `vif 53 disable` (it is master), vyos002 does. So an apply performed while
|
||
vyos002 held the VIP would enable vyos001's WAN as well, putting the cloned MAC
|
||
on both boxes. An override must assert `vif 53 disable` on **both** routers, with
|
||
the reconciler re-enabling whichever holds the VIP. Note the consequence: an
|
||
apply then briefly disables the current master's 10 gig until the reconciler
|
||
restores it (≤30s).
|