Some checks failed
Second drill of the day, after IPv6-follows-master went in. First time the IPv6 behaviour of a failover is a measurement rather than an assertion -- the thing this whole review started from was a 15.5ms figure read outside the window. TAKEOVER OK: 37s IPv6 followed in 37s (v4 37s, gap 0s) FAILBACK OK: 32s IPv6 back in 44s tun0 src : 87.192.101.48 -> 87.192.101.48 unchanged HE updates from vyos001: 0 HE updates from vyos002: 0 Zero HE API calls across a full takeover and failback is the invariant the design rests on, and it now has evidence: the 10 gig lease follows the cloned MAC, so the tunnel endpoint is the same address on whichever router holds the WAN and there is nothing to tell Hurricane Electric. This also exercised the proto-41 accept rule added to vyos002 earlier today. Without it the drill would have shown IPv6 failing to return while tun0 was up and radvd running -- which is precisely how the original gap hid. The two directions are NOT symmetric and the doc says so: IPv6 arrived in the same 5s sample on takeover but trailed by 12s on failback, because the reconciler enables bond0.53 first and v6_take only raises the tunnel once the source address exists. Bounded by one 30s tick. Also noted: the drill samples every 5s, so "gap 0s" means within the same sample, not simultaneous -- and the 37s vs the morning's 52s is a different run of the same IPv4 mechanism, not an improvement. Vodafone handed out a new address again across the drill (90.251.142.103 -> 90.251.152.236), corroborating that nothing may be pinned to the PPPoE address. Production returned to normal: vyos001 MASTER on all six with both WANs, vyos002 BACKUP with tun0 down and radvd stopped, force-fault clear. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01DMVzWZgiKW2wquf5z8S1yH
361 lines
19 KiB
Markdown
361 lines
19 KiB
Markdown
# PPPoE high availability
|
||
|
||
Proven in labsim. **Deployed to production 2026-09-06** — mechanism on both
|
||
routers, `pppoe0 disable` removed from vyos002, vyos002 out of FAULT and in
|
||
BACKUP, Pulumi model merged, and a controlled failover drill passed
|
||
(takeover 52s, failback 36s). `vyos:verify` reports both routers in sync.
|
||
|
||
## What it does
|
||
|
||
One consumer ISP account, two routers. The 10 gig lease is bound to a cloned MAC
|
||
(`f0:9f:c2:12:9b:4f`, the retired USG's) and the Vodafone line to a single
|
||
credential, so neither may be live on both boxes. The WAN follows VRRP
|
||
mastership — but the two halves use different control planes, and that is the
|
||
whole design:
|
||
|
||
| | plane | why |
|
||
|---|---|---|
|
||
| `bond0.53` (10 gig) | VyOS **config** (`disable`) | only config can move a MAC |
|
||
| `pppoe0` (Vodafone) | **systemd** unit gate | see below |
|
||
|
||
## Why PPPoE cannot live on the config plane
|
||
|
||
`interfaces_pppoe.py` treats `disable` and `delete` identically: both **unlink
|
||
`/etc/ppp/peers/pppoe0`**, call `PPPoEIf.remove()` (withdrawing the FRR default
|
||
route) and stop the unit. That path is pppd's own options file
|
||
(`ExecStart=/usr/sbin/pppd call %I`), so the resting state destroyed exactly what
|
||
the promotion path needed. `ppp@pppoe0` then restart-looped against the missing
|
||
file — 47 restarts observed, zero sessions at the access concentrator — and
|
||
never tripped systemd's limiter, because `RestartSec=5s` against the default
|
||
10s/5-burst window is only two restarts per interval.
|
||
|
||
It also made op-mode `connect interface pppoe0` unusable (it refuses without the
|
||
peers file), and put every failover behind a priority-322 commit where one
|
||
unrelated invalid node fails the whole thing.
|
||
|
||
## The gate
|
||
|
||
`pppoe0` is configured identically and **enabled on both** routers, so the peers
|
||
file always exists. Dialling is gated by
|
||
`/etc/systemd/system/ppp@pppoe0.service.d/10-vrrp-wan-gate.conf`:
|
||
|
||
```ini
|
||
ConditionPathExists=/run/vrrp-wan/may-dial
|
||
ConditionPathExists=/etc/ppp/peers/pppoe0
|
||
StartLimitIntervalSec=600
|
||
StartLimitBurst=6
|
||
```
|
||
|
||
`/run` is tmpfs, so the gate is shut at boot and neither box can dial before VRRP
|
||
has decided. **This is load-bearing, not a nicety:** with the node enabled,
|
||
`interfaces_pppoe.py` restarts ppp on *every* commit touching the pppoe subtree
|
||
when the daemon is not running — so the backup actively tries to dial whenever
|
||
anything commits (`pulumi up`, a hand commit, the boot-time config load). The
|
||
gate is the only thing making that a no-op, which is why `vrrp-wan-reconcile`
|
||
**refuses to bless a box whose drop-in is missing**: `/etc` is per-image, so a
|
||
VyOS upgrade silently removes the protection, and failing closed turns that into
|
||
"PPPoE never dials" rather than "both routers dial".
|
||
|
||
`may-dial` is a **lease, not a flag**. `ConditionPathExists` is evaluated at
|
||
start only — it can prevent a dial, never revoke one. `vrrp-wan-reconcile`
|
||
renews it every 30s; `vrrp-wan-guard` runs every 5s and only ever revokes, on
|
||
either "I do not hold the VIP" or "the lease is stale".
|
||
|
||
## Measured in labsim
|
||
|
||
| | |
|
||
|---|---|
|
||
| Clean failover (`force-fault`) | `pppoe0` moves in **20-26s**, reproducible |
|
||
| 10 gig down → PPPoE | route falls to `pppoe0`; LAN back online in **5s** |
|
||
| Stale lease | guard hangs up within ~5s |
|
||
| Missing peers file | `NRestarts=0` — no loop |
|
||
| Flap holdoff vs live session | session survives a forged 900s holdoff |
|
||
| Invariant | never more than one router dialled, in any run |
|
||
|
||
On the last row, precisely: the AC *did* report two `simdsl` sessions during the
|
||
`session-control=disable` run — one live, one orphaned from the router that had
|
||
just been destroyed, which that policy does not clean up. Only one live router
|
||
was ever dialled. The count is a proxy for the real invariant and only a valid
|
||
one while the AC enforces single-session, so it is reported as a `WARN` rather
|
||
than silenced: an orphaned session still occupies the single slot at a real ISP,
|
||
and that is exactly what made `deny` take 141–148s.
|
||
|
||
## Two failure modes found by running it, not by reading it
|
||
|
||
**The flap damper tore down a healthy WAN.** `ppp_dial()` checked the hold-off
|
||
and returned *before* renewing `may-dial`. That file is a lease the guard
|
||
expires after `LEASE_TTL`, so tripping the damper stopped the renewal and the
|
||
guard hung up `pppoe0` **on the master** ~80s later:
|
||
|
||
```
|
||
DIAL FLAP: >=6 attempts in 600s -- holding off 900s
|
||
GUARD: lease stale (81s > 75s) -- hanging up pppoe0
|
||
```
|
||
|
||
A damper meant to suppress repeated *dials* was destroying an established
|
||
session instead. An active session now renews the lease and returns before
|
||
every other check; everything below only decides whether to start a **new**
|
||
session. T12 is the regression test.
|
||
|
||
**A missing peers file is silent.** `/etc/ppp/peers/pppoe0` is both pppd's
|
||
options file and the gate's second condition, and it is only written by a commit
|
||
that touches the pppoe subtree. Without it systemd logs
|
||
`skipped because of an unmet condition check` exactly once and then nothing —
|
||
a router that cannot dial at all looks identical to a healthy backup.
|
||
`ppp_dial()` now says so on every tick, and distinguishes the two causes:
|
||
configured-but-not-rendered (re-commit the subtree) versus no `pppoe0` in the
|
||
config at all.
|
||
|
||
The second cause is the one to watch in production: **a commit that was never
|
||
`save`d reverts on reboot and takes `pppoe0` with it.** That is exactly how the
|
||
sim secondary lost its WAN and spent hours looking like an ISP problem. After
|
||
any hand commit to the pppoe subtree, `save` — or the next reboot produces a
|
||
standby that can never take over.
|
||
|
||
## Deploying — steps 1–9 done 2026-09-06
|
||
|
||
1. `sudo /config/vyos-known-good save` on both.
|
||
2. `migration/vrrp-wan-install --vip 192.168.1.1 --host vyos@10.0.1.253` then the
|
||
same for `.252`. Then `--check` on both. **No config change yet** — verify
|
||
nothing dials.
|
||
3. Confirm `/config/wan-secrets` is present and identical on both.
|
||
4. **vyos002 first** (the non-master), on its own commit — `interfaces pppoe` is
|
||
priority 322 and one bad node fails everything:
|
||
`delete interfaces pppoe pppoe0 disable`, `commit-confirm 10`.
|
||
5. Verify vyos002 did **not** dial — check on the wire, not from state:
|
||
`sudo tcpdump -i bond0.51 -nn pppoed` should show no PADI. Then `confirm`; `save`.
|
||
6. vyos001: nothing to change; it already has `pppoe0` enabled.
|
||
7. Confirm `vif 53 disable` is in **both** `config.boot`s. vyos001's lacked it;
|
||
fixed with `migration/vif53-pin-boot-disable` — see below.
|
||
8. **Done 2026-09-06** (`kubernetes-deployment@45033dd`, on `main`). Both
|
||
overrides merged, transition-scripts applied to both boxes by hand rather
|
||
than left as drift, and `vyos:verify` is clean: 533 / 512 nodes, zero drift.
|
||
The staging file `migration/pulumi-override-pppoe-gated.json` is kept as the
|
||
record of why the ordering mattered. It was staged, unapplied, on purpose: another agent runs `pulumi up` on that repo, so
|
||
merging it *is* a production change made by someone else at a time you do not
|
||
choose. Removing `pppoe0 disable` from vyos002 before the gate exists there
|
||
lets it dial on the next commit and take the single Vodafone session off
|
||
vyos001. Run `npm run vyos:export && npm run vyos:render` first so the model
|
||
follows whichever router actually holds the WAN.
|
||
9. Add the drop-in re-install to the VyOS image-upgrade runbook.
|
||
**Done** — `migration/VYOS-IMAGE-UPGRADE.md`. An upgrade replaces `/etc` and
|
||
so removes the gate; the reconciler fails closed, giving "PPPoE never
|
||
dials" rather than "both routers dial".
|
||
|
||
Step 4 is the one that matters most and is worth stopping on. vyos002 has been
|
||
in **FAULT on all six groups for over three days** — verified again while
|
||
writing this, alongside vyos001 holding `192.168.1.1` on `bond0.1` and
|
||
`pppoe0` up on `83.106.5.72`. Until vyos002 reaches BACKUP there is no standby
|
||
at all: if vyos001 died today nothing would pick up the gateway VIPs. The
|
||
existing `vrrp-health-check-wan-present` override predicted exactly this in its
|
||
own reason text — *"with the primary genuinely dead the secondary stays FAULT
|
||
and nothing holds the gateway. The fix for that is WAN-follows-master, which is
|
||
a separate change."* This is that change.
|
||
|
||
## Proven in production — controlled drill, 2026-09-06
|
||
|
||
`migration/wan-drill` force-faulted vyos001 and timed a real failover:
|
||
|
||
```
|
||
TAKEOVER OK: vyos002 held the VIP and reached the internet in 52s
|
||
FAILBACK OK: 36s
|
||
```
|
||
|
||
vyos002 took the VIPs at t+18s and had **both** WANs by t+52s. Failback put
|
||
vyos001 back with a WAN in 36s. Total interruption ≈ 88s across two deliberate
|
||
transitions.
|
||
|
||
**The cloned-MAC lease transfers.** This was the largest untested item in the
|
||
whole design — whether the 10 gig ISP would re-issue `87.192.101.48` to
|
||
`f0:9f:c2:12:9b:4f` arriving on a different switch port. It did, same address,
|
||
within the takeover window. That risk is now closed.
|
||
|
||
**Vodafone did not refuse the re-dial**, so its `session-control` behaves like
|
||
`replace` rather than the hostile `deny`. `GRACE=300` was sized against the
|
||
sim's 148s `deny` case and is therefore comfortable — but it should stay where
|
||
it is, because one drill on one evening does not establish the ISP's policy
|
||
under all conditions.
|
||
|
||
**Vodafone hands out a different IPv4 on every dial**: `83.106.5.72` →
|
||
`90.251.153.180` (vyos002) → `90.251.142.103` (vyos001, after failback).
|
||
Nothing may be pinned to the PPPoE address. The HE IPv6 tunnel is pinned to
|
||
`87.192.101.48`, which is the **10 gig** (`bond0.53`) and stable across
|
||
failover. Anything added later that hardcodes a WAN IP must use the 10 gig one,
|
||
not `pppoe0`'s.
|
||
|
||
> **CORRECTED 2026-09-06.** This paragraph originally continued "so `tun0`
|
||
> survived untouched and IPv6 stayed up at 15.5ms." **That conclusion was
|
||
> wrong, and it should never have been recorded as a result.** The premise is
|
||
> right — the endpoint address is stable — but it does not follow. During
|
||
> takeover the reconciler disables `bond0.53` on the demoted box, so
|
||
> `87.192.101.48` *leaves vyos001 and appears on vyos002*, and vyos002 has no
|
||
> `tun0` at all: no tunnel, no `he-tunnel-follow`, no `/config/he-secrets`, no
|
||
> VLAN 9 prefix, no `route6 ::/0`. Inbound protocol 41 from HE lands on a router
|
||
> with nothing to decapsulate it. vyos001's own journal for the drill window
|
||
> reads `08:36:01 he-tunnel-follow: no default route; refusing to guess`.
|
||
>
|
||
> The reading was taken either side of the window, not through it — `wan-drill`
|
||
> contained **no IPv6 check of any kind**, and neither did any other part of the
|
||
> mechanism. That is now fixed: the drill probes IPv6 in both timing loops and
|
||
> asserts that a router-level failover makes **zero** HE API calls. Until a
|
||
> drill produces that figure, the IPv6 behaviour of a failover is *unmeasured*,
|
||
> not "fine".
|
||
>
|
||
> The gap itself was real: **IPv6 was single-homed on vyos001 while the WAN
|
||
> beneath it was HA.** **Closed and measured 2026-09-06** — see the drill below.
|
||
|
||
Production takeover (52s) is about twice the sim's `replace` figure (26s), which
|
||
is the expected direction: the VP2440s commit under kea, BGP and conntrack while
|
||
the sim routers are idle.
|
||
|
||
## IPv6 follows the WAN — measured, 2026-09-06
|
||
|
||
The second drill of the day, run after IPv6-follows-master was deployed. This is
|
||
the first time the IPv6 behaviour of a failover has been a **measurement** rather
|
||
than an assertion:
|
||
|
||
```
|
||
TAKEOVER OK: vyos002 held the VIP and reached the internet in 37s
|
||
IPv6 followed in 37s (v4 37s, gap 0s)
|
||
FAILBACK OK in 32s
|
||
IPv6 back in 44s
|
||
tun0 src : 87.192.101.48 -> 87.192.101.48 OK: unchanged across the drill
|
||
HE updates from vyos001: 0
|
||
HE updates from vyos002: 0
|
||
```
|
||
|
||
Full log: `migration/drill-evidence/wan-drill-2026-09-06-ipv6.log`.
|
||
|
||
**Zero HE API calls across a full takeover and failback.** This is the invariant
|
||
the design rests on and it now has evidence: the 10 gig lease follows the cloned
|
||
MAC, so the tunnel endpoint is the *same address* on whichever router holds the
|
||
WAN, and there is nothing to tell Hurricane Electric. Only the within-box fall
|
||
back to PPPoE needs an HE update.
|
||
|
||
**IPv6 is no longer the laggard, but the two directions are not symmetric.** On
|
||
takeover it arrived in the same 5s sample as IPv4; on failback it trailed by 12s.
|
||
That asymmetry is the reconciler's tick, not a fault: on promotion it enables
|
||
`bond0.53` first, and `v6_take` only raises the tunnel once the source address
|
||
actually exists, so it can land on the following 30s tick. Bound is one tick.
|
||
Note the sampling granularity — the drill polls every 5s, so "gap 0s" means
|
||
"within the same sample", not "simultaneous".
|
||
|
||
Takeover was 37s here against 52s in the morning drill. Do not read that as an
|
||
IPv6 improvement; it is the same IPv4 mechanism on a different run, and the
|
||
morning figure included the first-ever cloned-MAC lease transfer.
|
||
|
||
**Vodafone confirmed the every-dial-a-new-address behaviour again**: `pppoe0`
|
||
came back as `90.251.152.236`, having been `90.251.142.103` before the drill.
|
||
|
||
## config.boot pins `vif 53 disable` on both — fixed 2026-09-06
|
||
|
||
The convention is that **both** `config.boot`s hold `vif 53 disable`, so a reboot
|
||
in any order comes up unable to claim the cloned MAC and the reconciler enables
|
||
it on whichever box holds the VIP. vyos001's did not; its `config.boot` predated
|
||
this work.
|
||
|
||
There is no clean way to express "boot disabled, run enabled" in VyOS: **`save`
|
||
writes the RUNNING config, not the candidate.** Setting the node, saving and
|
||
discarding was tested in labsim and `config.boot` came back *without* `disable`,
|
||
the WAN untouched. So the node must genuinely be disabled, saved, and
|
||
re-enabled. `migration/vif53-pin-boot-disable` does exactly that, is idempotent,
|
||
and no-ops on a box that already has it.
|
||
|
||
Cost, measured on vyos001: a **32s** window (10s in the sim — production commits
|
||
under kea/BGP/conntrack are slower), `bond0.53` re-leased `87.192.101.48` 5s
|
||
after re-enable, and `vyos-failover` restored the primary route about a minute
|
||
later:
|
||
|
||
```
|
||
09:32:23 ip route del 0.0.0.0/0 ... dev bond0.53
|
||
09:32:32 Check fail for route 0.0.0.0/0 interface "bond0.53"
|
||
09:33:23 ip route add 0.0.0.0/0 via 87.192.96.1 dev bond0.53 metric 1 proto failover
|
||
```
|
||
|
||
The house rode `pppoe0` for that minute rather than losing the internet, which
|
||
is the T5 path working. Note the shape of that recovery before reading a fresh
|
||
`ip route show` as a regression: for ~60s after the bounce the default really is
|
||
on `pppoe0`, because `vyos-failover` only re-adds the `bond0.53` route once its
|
||
probes pass again.
|
||
|
||
**A race worth knowing about.** The first sim run collided with
|
||
`vrrp-wan-reconcile`'s own commit — *"Configuration system temporarily locked due
|
||
to another commit in progress"* — and the `save` landed while the **re-enable did
|
||
not**, leaving the master with its 10 gig down. The script now takes the
|
||
reconciler's `/run/vrrp-wan.lock` (which the reconciler skips a tick rather than
|
||
block on), with `9>&-` so the config session's unionfs child cannot inherit it.
|
||
Even the bad run ended correctly — the reconciler logged *"MASTER with bond0.53
|
||
disabled -> enabling"* and repaired it in 4s — so a half-completed run is
|
||
survivable by design. The script no longer leans on that, and verifies the
|
||
re-enable rather than reporting a success it did not achieve.
|
||
|
||
**Rollback**, from either box: `set interfaces pppoe pppoe0 disable` on both and
|
||
`rm /run/vrrp-wan/may-dial`. That restores today's behaviour exactly.
|
||
|
||
## What the sim cannot prove
|
||
|
||
- **Vodafone's `session-control`.** The matrix now brackets it properly.
|
||
Destroying the master and timing the survivor's session:
|
||
|
||
| policy | takeover |
|
||
|---|---|
|
||
| `replace` (accel-ppp default) | 26s / 26s |
|
||
| **`deny`** (hostile) | **148s / 141s** |
|
||
| `disable` | 21s / 20s |
|
||
|
||
`deny` is the sizing case: the AC refuses the survivor until its own
|
||
dead-peer timer frees the dead session, and the poller caught two dial
|
||
attempts being rejected before one succeeded. `GRACE` is set from that — see
|
||
`migration/vrrp-wan.conf`. This still cannot tell you which policy Vodafone
|
||
runs, and account rate-limiting or lockout on repeated dials has no sim
|
||
analogue at all; the flap damper (6 dials / 600s → 15 min hold-off) exists
|
||
for that.
|
||
|
||
Treat the 148s as a floor rather than a worst case. These are idle 2-vCPU
|
||
VMs, and the AC shares an OVS bridge with the routers, so `virsh destroy`
|
||
removes the port and accel-ppp sees the peer physically vanish. A real BRAS
|
||
reached over DSL does not learn that our router died — it waits out its own
|
||
timers, which are longer and not ours to know.
|
||
|
||
Worth knowing how close this came to being missed: until 2026-09-06 the
|
||
matrix set the policy with `vbash -c 'source script-template; configure;
|
||
...; commit'`, which never starts a config session. `commit` failed to
|
||
stderr, the helper discarded it, and all three iterations ran against the
|
||
default while printing the mode they were supposedly testing. It reported
|
||
`deny` at 25s. The real figure is 148s.
|
||
- **Whether Vodafone honours our LCP Terminate / PADT** on a graceful stop.
|
||
- ~~**The cloned-MAC lease.**~~ **Answered 2026-09-06**: the drill moved it and
|
||
the ISP re-issued `87.192.101.48` to `f0:9f:c2:12:9b:4f` on vyos002's port
|
||
within the takeover window. See "Proven in production" above.
|
||
- ~~**Whether VyOS can dial Vodafone at all.**~~ **Answered 2026-09-06**: both
|
||
routers dialled successfully during the drill. MTU/MSS under sustained load
|
||
is still unmeasured, and Vodafone hands out a different IPv4 every dial.
|
||
- **Timing under load.** The sim routers are idle 2-vCPU VMs; commit latency on
|
||
the VP2440s under kea + BGP + conntrack will be worse, and commit latency is
|
||
the dominant term in the `bond0.53` half of a failover.
|
||
|
||
## How the model handles the asymmetry — resolved
|
||
|
||
An earlier draft of this file said an override "must assert `vif 53 disable` on
|
||
**both** routers". **That advice was wrong and has been removed**; do not
|
||
reintroduce it. `vif 53 disable` is *runtime* state owned by
|
||
`vrrp-wan-reconcile`, keyed on who holds the management VIP, so pinning it in
|
||
the model would fight the reconciler on every apply and would briefly disable
|
||
the live master's 10 gig each time.
|
||
|
||
What is actually done, and why it is safe:
|
||
|
||
- **Runtime (Pulumi): follow reality.** Run `npm run vyos:export && npm run
|
||
vyos:render` immediately before any `pulumi up` touching vyos. Whichever
|
||
router currently holds the WAN keeps it; the apply is a no-op on that node.
|
||
There is deliberately **no** override for `vif 53 disable`.
|
||
- **Boot (`config.boot`): hardcode safe.** Both routers pin `vif 53 disable`,
|
||
so a reboot in any order comes up unable to claim the cloned MAC and the
|
||
reconciler enables it on whoever holds the VIP. See the section above.
|
||
- **Install time (PXE, nothing to follow).** `migration/vyos-mode-delta.py`
|
||
emits `vif 53 disable` for a box with no WAN, and deliberately does *not*
|
||
emit `pppoe pppoe0 disable`.
|
||
|
||
Verified 2026-09-06: `vyos:verify` reports both routers in sync, 533 and 512
|
||
nodes, zero drift.
|