vrrp-wan: size GRACE from the measured hostile failover, not the theory
Some checks failed
CI/CD / typecheck (push) Failing after 9s
CI/CD / test (push) Failing after 9s
CI/CD / lint (push) Failing after 24s
CI/CD / build (push) Has been skipped
CI/CD / publish-rpm (push) Has been skipped
CI/CD / publish-deb (push) Has been skipped

With the matrix actually setting session-control, T4 timed a destroyed
master's takeover at:

    replace    26s
    deny      148s   <-- sizing case
    disable    21s

`deny` is the case GRACE exists for: the AC refuses the survivor until its own
dead-peer timer frees the dead session. The session poller caught it happening
-- the destroyed router's session stayed in the table while the survivor's
dials appeared and were rejected, twice, before one took at 148s.

148s is well past the lcp-echo-interval(30) x failure(3) = 90s budget that 180
was sized against, leaving 32s of margin. GRACE=300 is ~2x the worst observed.
Treat 148s as a floor, not a worst case: these are idle 2-vCPU VMs, and the AC
shares an OVS bridge with the routers, so `virsh destroy` removes the port and
accel-ppp sees the peer vanish. A real BRAS over DSL never learns our router
died and waits out longer timers of its own.

The cost is stated in the conf: GRACE is also how long an alive-but-unroutable
master holds every VIP before yielding. bond0.53 covers most of that -- a DHCP
lease satisfies the check in seconds -- so GRACE only dominates when PPPoE is
the last path. Kept the health check's fallback in step, since keepalived runs
it with no environment and that number decides mastership if the conf is ever
missing.

check_invariant no longer fails blind on the AC's session count. That count is
only a proxy for "two of our routers dialled", and only while the AC enforces
single-session; under `disable` it does not, so a destroyed router's session
lingers and the count reads 2 with exactly one live router dialled. The real
invariant -- at most one router holds pppoe0 -- is now the failing one, and the
stale session is reported as a WARN rather than silenced, because it still
occupies the slot at a real ISP and is precisely what made `deny` take 148s.
This commit is contained in:
Michal
2026-09-06 00:25:00 +01:00
parent 9221c71ff0
commit 6987b324f1
7 changed files with 151 additions and 28 deletions

View File

@@ -115,19 +115,60 @@ standby that can never take over.
`sudo tcpdump -i bond0.51 -nn pppoed` should show no PADI. Then `confirm`; `save`.
6. vyos001: nothing to change; it already has `pppoe0` enabled.
7. Confirm `vif 53 disable` is in **both** `config.boot`s.
8. Add the drop-in re-install to the VyOS image-upgrade runbook.
8. **Only now** merge `migration/pulumi-override-pppoe-gated.json` into
`kubernetes-deployment` `infra/vyos/subtrees/overrides.json`. It is staged
here, unapplied, on purpose: another agent runs `pulumi up` on that repo, so
merging it *is* a production change made by someone else at a time you do not
choose. Removing `pppoe0 disable` from vyos002 before the gate exists there
lets it dial on the next commit and take the single Vodafone session off
vyos001. Run `npm run vyos:export && npm run vyos:render` first so the model
follows whichever router actually holds the WAN.
9. Add the drop-in re-install to the VyOS image-upgrade runbook.
Step 4 is the one that matters most and is worth stopping on. vyos002 has been
in **FAULT on all six groups for over three days** — verified again while
writing this, alongside vyos001 holding `192.168.1.1` on `bond0.1` and
`pppoe0` up on `83.106.5.72`. Until vyos002 reaches BACKUP there is no standby
at all: if vyos001 died today nothing would pick up the gateway VIPs. The
existing `vrrp-health-check-wan-present` override predicted exactly this in its
own reason text — *"with the primary genuinely dead the secondary stays FAULT
and nothing holds the gateway. The fix for that is WAN-follows-master, which is
a separate change."* This is that change.
**Rollback**, from either box: `set interfaces pppoe pppoe0 disable` on both and
`rm /run/vrrp-wan/may-dial`. That restores today's behaviour exactly.
## What the sim cannot prove
- **Vodafone's `session-control`.** The sim's accel-ppp defaults to `replace`, so
a new auth kills the old session immediately. A real BRAS may `deny` and hold
the session for its own dead-peer timer. The matrix runs all three modes to
bracket the risk; it cannot tell you which one you will meet, and account
rate-limiting or lockout on repeated dials has no sim analogue at all. The
flap damper (6 dials / 600s → 15 min hold-off) exists for that.
- **Vodafone's `session-control`.** The matrix now brackets it properly.
Destroying the master and timing the survivor's session:
| policy | takeover |
|---|---|
| `replace` (accel-ppp default) | 26s |
| **`deny`** (hostile) | **148s** |
| `disable` | 21s |
`deny` is the sizing case: the AC refuses the survivor until its own
dead-peer timer frees the dead session, and the poller caught two dial
attempts being rejected before one succeeded. `GRACE` is set from that — see
`migration/vrrp-wan.conf`. This still cannot tell you which policy Vodafone
runs, and account rate-limiting or lockout on repeated dials has no sim
analogue at all; the flap damper (6 dials / 600s → 15 min hold-off) exists
for that.
Treat the 148s as a floor rather than a worst case. These are idle 2-vCPU
VMs, and the AC shares an OVS bridge with the routers, so `virsh destroy`
removes the port and accel-ppp sees the peer physically vanish. A real BRAS
reached over DSL does not learn that our router died — it waits out its own
timers, which are longer and not ours to know.
Worth knowing how close this came to being missed: until 2026-09-06 the
matrix set the policy with `vbash -c 'source script-template; configure;
...; commit'`, which never starts a config session. `commit` failed to
stderr, the helper discarded it, and all three iterations ran against the
default while printing the mode they were supposedly testing. It reported
`deny` at 25s. The real figure is 148s.
- **Whether Vodafone honours our LCP Terminate / PADT** on a graceful stop.
- **The cloned-MAC lease** — whether the 10 gig ISP re-issues `87.192.101.48` to
`f0:9f:c2:12:9b:4f` arriving on a different switch port. That risk belongs to

View File

@@ -40,7 +40,14 @@ VIP="${VRRP_WAN_VIP:-192.168.1.1}"
# the boundary: the new master fails its own check, sheds the VIPs, and the peer
# -- in the same position -- does likewise. Both end in FAULT, which is worse
# than the outage this check exists to prevent.
GRACE="${GRACE:-180}"
#
# The fallback tracks vrrp-wan.conf, where the reasoning and the measurements
# live. Short version: labsim T4 timed a destroyed master's takeover at 26s
# (replace), 148s (deny) and 21s (disable); `deny` is the sizing case and 180
# left only 32s over it. Keep the two in step -- keepalived runs this script
# with no environment, so if vrrp-wan.conf is ever missing THIS number is the
# one that decides mastership.
GRACE="${GRACE:-300}"
# A deliberate hand-over lever.
#

View File

@@ -20,10 +20,38 @@ WAN_VIF=53
#
# Must exceed the ISP's stale-session hold-down, or a hard failover blows the
# window and BOTH routers end up in FAULT -- worse than the outage the check
# exists to prevent. accel-ppp's default dead-peer budget is
# lcp-echo-interval(30) x lcp-echo-failure(3) = 90s, so 90 sat exactly on the
# boundary. Measured worst case: see labsim/wan-failover-evidence/.
GRACE=180
# exists to prevent.
#
# MEASURED, labsim T4, master destroyed with `virsh destroy`, time until the
# survivor held a PPPoE session (labsim/wan-failover-evidence/T4-*):
#
# session-control=replace 26s
# session-control=deny 148s <-- worst
# session-control=disable 21s
#
# `deny` is the hostile case and the only one that matters for sizing: the AC
# refuses the survivor until its own dead-peer timer frees the dead session.
# The poller caught it happening -- the destroyed router's session stayed in the
# table while the survivor's dial attempts appeared and were rejected, twice,
# before it finally got in at 148s.
#
# 148s also lands well past the theoretical lcp-echo-interval(30) x
# failure(3) = 90s budget that 180 was originally sized against, which left only
# 32s of margin. 300 gives roughly 2x the worst observed, on IDLE 2-vCPU sim
# VMs; the VP2440s under kea, BGP and conntrack will be slower, and Vodafone's
# actual policy and timers are unknown.
#
# The cost is real and worth stating: this is also how long a master that is
# alive but genuinely cannot route keeps holding every VIP before yielding --
# the 2026-09-02 outage shape. That case is mostly covered by bond0.53, which
# satisfies the check within seconds of getting a DHCP lease; GRACE only
# dominates when PPPoE is the only path left.
#
# Do not lower this below the worst measured handover without re-running
# `labsim/labsim-pppoe-ha-test.sh --hard`. Before 2026-09-06 that matrix never
# actually set session-control and reported `deny` at 25s -- a number that did
# not exist.
GRACE=300
# may-dial is a LEASE, not a flag. vrrp-wan-reconcile renews its mtime every
# tick; vrrp-wan-guard revokes it once it goes stale. A plain flag survives the