Files
lab/migration/ppp-vrrp-gate.conf

51 lines
2.6 KiB
Plaintext
Raw Permalink Normal View History

vyos: move PPPoE off the config plane onto a gated systemd unit PPPoE HA could not work as written, and the reason is structural rather than a bug: `set interfaces pppoe pppoe0 disable` and `delete` are handled identically by interfaces_pppoe.py -- both UNLINK /etc/ppp/peers/pppoe0. That path is pppd's own options file, so the resting state destroyed exactly what the promotion path needed, and `ppp@pppoe0` restart-looped against it (observed: 47 restarts, zero sessions at the access concentrator). It also made op-mode `connect interface pppoe0` unusable, and put every failover behind a priority-322 commit where one unrelated invalid node fails the whole thing -- which has already taken the 10 gig down once. pppoe0 is now configured identically and ENABLED on both routers, so the peers file always exists, and dialling is gated by a drop-in on the unit: ConditionPathExists=/run/vrrp-wan/may-dial ConditionPathExists=/etc/ppp/peers/pppoe0 /run is tmpfs, so the gate is shut at boot and neither box can dial before VRRP has decided. That matters more than it first appears: with the node enabled, interfaces_pppoe.py restarts ppp on EVERY commit touching the pppoe subtree when the daemon is not running -- so the backup actively tries to dial whenever anything commits. The gate is the only thing making that a no-op, which is why vrrp-wan-reconcile now refuses to bless a box whose drop-in is missing: /etc is per-image, and a VyOS upgrade would otherwise silently remove the protection. may-dial is a LEASE, not a flag. ConditionPathExists is evaluated at start only -- it can prevent a dial, never revoke one -- so a reconciler that stops running while its box is demoted would keep the one ISP session for ever. The reconciler renews the lease; a new 5s vrrp-wan-guard revokes it, and only ever revokes. It fired correctly first time: "GUARD: lease stale (81s > 75s)". Also: remove-then-stop on release (the file's absence blocks a NEW start that a concurrent commit would trigger); a flap damper, because two routers that both believe they hold the VIP will both dial and each dial kills the other's session -- against a real ISP that is how an account gets rate-limited; and a guard on `cfg` returning empty under commit-lock contention, which had already produced one spurious "releasing" on a box that needed nothing. GRACE 90 -> 180. accel-ppp's dead-peer budget is lcp-echo-interval(30) x failure(3) = 90s, so the old value sat exactly on the boundary: a hard failover into an AC that does not replace the stale session would fail its own check, shed the VIPs, and leave both routers in FAULT. The sim could not have tested any of this. Both routers now get the identical WAN -- the secondary had none "because two PPPoE clients sharing one credential is a different failure mode than anything production has", which is backwards: that IS production. It also left the pair incomparable, ten NAT rules against none. Safety now comes from resting state, not asymmetry. Three more things the sim was hiding: - the drift check's secondary regex omitted interfaces pppoe/bonding, nat source and protocols failover, so it reported "in sync" for a box with no WAN at all; - the VRRP health-check and transition-script hooks existed on both live VMs and in NEITHER generator -- the mechanism under test was pure undetected drift; - labsim-vyos's only default route was the libvirt-NAT scaffold, so every "the LAN still has internet" verdict on it was answered by eth2 rather than the WAN. --drop-scaffold applied; the earlier DHCP-failover proof is being re-run because of it. vrrp-wan-install ends the other half of that: the sim's previous proof came from scripts hand-`sed`-ed in place, so the tested behaviour was not the committed behaviour. `--check` now makes that a hard failure. First green run: master holds both WANs, backup released, and the AC reports exactly ONE session. sim-net-apply.sh check: all four in sync. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01DMVzWZgiKW2wquf5z8S1yH
2026-09-05 18:50:16 +01:00
# Installed to /etc/systemd/system/ppp@pppoe0.service.d/10-vrrp-wan-gate.conf
#
# This drop-in is the ONLY thing preventing both routers from dialling the one
# ISP credential at the same time. Do not remove it without reading this.
#
# pppoe0 is configured identically and ENABLED on both routers, because the
# alternative -- `set interfaces pppoe pppoe0 disable` -- unlinks
# /etc/ppp/peers/pppoe0 (interfaces_pppoe.py treats `disable` and `delete`
# identically), and pppd's options file IS that path. A promotion then had to
# re-render it via a full config commit at priority 322, where one unrelated
# invalid node fails the whole commit and takes the 10 gig down with it. It also
# made op-mode `connect interface pppoe0` unusable, since that refuses when the
# peers file is absent.
#
# With the node enabled, interfaces_pppoe.py's apply() does this on EVERY commit
# that touches the pppoe subtree:
#
# if not is_systemd_service_running('ppp@pppoe0.service') or shutdown_required:
# call('systemctl restart ppp@pppoe0.service')
#
# -- i.e. the backup actively tries to dial whenever anything commits. A
# `pulumi up`, a `sim-net-apply.sh apply`, or the boot-time config load are all
# that commit. This gate is what makes that a no-op.
#
# /run is tmpfs, so the gate is shut at boot on both boxes and neither can dial
# before VRRP has decided. ppp@.service is already After=vyos-router.service, so
# no extra ordering is needed.
[Unit]
# Both must hold; multiple ConditionPathExists are ANDed.
# may-dial -- vrrp-wan-reconcile has blessed this box (a renewed lease)
# /etc/ppp/peers -- refuse to start pppd against a missing options file, which
# is what produced a restart loop of 47 and counting on
# 2026-09-05. A failed Condition is NOT a failure: the job
# succeeds, the unit stays inactive, and `systemctl start`
# exits 0 -- so callers must check is-active, never rc.
ConditionPathExists=/run/vrrp-wan/may-dial
ConditionPathExists=/etc/ppp/peers/pppoe0
# Belt to that brace. The stock unit is Restart=on-failure/RestartSec=5s against
# systemd's default StartLimitIntervalSec=10s/Burst=5 -- two restarts per window,
# so the limiter can never trip and a doomed pppd retries for ever.
StartLimitIntervalSec=600
StartLimitBurst=6
[Service]
RestartSec=15
# A hung pppd must be resolved inside the failover budget. The stock 90s means a
# demoted router could still hold the session while the new master is dialling.
# 20s still allows a clean LCP Terminate + PADT in the normal case.
TimeoutStopSec=20