"No NAT? How are we supposed to get internet?" -- a fair question that exposed a
worse design than I had admitted. Internet did work, but only via vyos001: NAT
and the entire WAN were gated behind --with-wan, so vyos002 would have held the
LAN VIPs and routed between VLANs with no path to the outside at all. Failover
would have preserved addressing and lost the internet.
The fix rests on a checked fact rather than an assumption: VyOS WARNS but still
commits when a NAT rule names an interface that does not exist
("Interface bond0.53 for source NAT rule 900 does not exist!"). Verified on a
real VyOS before relying on it.
So both boxes now get the identical WAN, NAT, port-forward and firewall config,
and the backup's two WAN interfaces are simply set `disable`. The cloned WAN MAC
is therefore never live on two boxes at once, while everything needed to route
and masquerade is already in place. The two deltas are now byte-identical apart
from VRRP priority, own/peer addresses, DHCP HA role, the conntrack /30 -- and
the two disable lines.
Taking over the internet path becomes deleting two lines rather than
reconstructing NAT under pressure:
delete interfaces bonding bond0 vif 53 disable
delete interfaces pppoe pppoe0 disable
Both boxes now: 21 NAT rules, 58 firewall rules, full PPPoE. Backup delta
validated against a real VyOS config with the disable lines present -- commits
clean. Runbook updated with the takeover procedure and the warning that it must
only be done when vyos001 is genuinely down.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01DMVzWZgiKW2wquf5z8S1yH
8.5 KiB
Cutover runbook — USG to VyOS
Print this. During the cutover there is no internet, so there is no assistant and no web search. Everything you need is on this page and on the boxes themselves.
Use these addresses. Not the other ones.
| use this | do NOT use | |
|---|---|---|
| vyos001 (MASTER) | 10.0.1.252 |
|
| vyos002 (BACKUP) | 10.0.1.253 |
ssh vyos@10.0.1.252 — by IP, not by name.
The 192.168.8.x addresses stop working the instant the USG is unplugged.
That is not a maybe. Your workstation is on LoT (10.0.0.210/23) and reaching
192.168.8.x requires routing through the USG:
ip route get 192.168.8.143 -> via 10.0.0.1 <- the USG. Gone.
ip route get 10.0.1.252 -> dev lanbr0 <- same L2. Survives.
10.0.1.252 and .253 are on the LoT VLAN, the same broadcast domain as your
workstation, so they need no gateway at all. They are the only remote path that
survives the cutover.
Between unplugging the USG and finishing the switch there is no inter-VLAN routing. In that window:
- the JetKVMs are unreachable from your workstation (they are on Management and kvm) — they are not a fallback during the gap
- Tailscale is down with the internet
- your workstation keeps
10.0.0.210(86400s lease) and can still resolve via10.0.0.194, which is also link-scope
If LoT SSH fails, the next step is physical console, not the network.
| JetKVMs (after routing is restored) | 192.168.1.28, 192.168.1.29, 192.168.3.6 |
| Switch script | /config/vyos-unifi-switch on each box |
| Login | user vyos |
If something is wrong, do this
sudo /config/vyos-unifi-switch unifi
Then reconnect the USG. That command runs no health checks, asks nothing and cannot refuse. It restores a byte-exact copy of the configuration the box had before the cutover — verified by diff, not by assumption.
You do not have to be quick. If you do nothing at all after
vyos-unifi-switch vyos, the box reverts by itself within 10 minutes. Verified:
config returns to the previous state and the box does not reboot
(uptime and boot-id unchanged across an auto-revert).
What has actually been tested
Proven on the labsim router (same VyOS version, isolated OVS bridge with no physical NIC), by loading vyos001's real running config and applying the real production delta:
- All 317 commands accepted, and the whole delta commits (
COMMIT OK). unifimode restores the previous config byte-exact (diff clean).- Auto-revert fires when the commit is not confirmed: config returns to the
saved state and the box does not reboot —
uptimeand boot-id unchanged across the revert. - Failed health checks trigger an immediate revert rather than waiting out the timer.
Two bugs were found this way and would each have failed the entire switch,
since the delta commits as one unit: bond0.51 did not exist for PPPoE to
reference, and translation port rejects a port list.
Not tested, and untestable in advance:
- PPPoE. The line permits one session and the USG holds it. The first real attempt is during the cutover.
- The commit on the real boxes. The rehearsal ran with vyos001's
eth2andeth3stanzas stripped, because the sim VM has only two NICs. Those are plain interface configs that already work on the real hardware, but they were not part of what committed.
Before you unplug anything
- Tether your workstation to your phone if you want the assistant available. Cutting the USG cuts your internet, not your LAN.
- On both boxes, confirm the machinery is present:
sudo /config/vyos-unifi-switch status ls -la /config/modes/ # unifi.boot + to-vyos.commands ls -la /config/wan-secrets # must be 0600statusmust reportmode: unifi. Ifunifi.bootis missing, stop — there is no way back without it. - Confirm the revert action is
reload, notreboot:Must showshow configuration commands | match commit-confirmaction 'reload'. Without it a failed switch reboots the firewall instead of reverting it. The switch script refuses to run if this is missing, but check anyway.
The cutover
-
Open both SSH sessions BEFORE you unplug anything, and leave them open:
ssh vyos@10.0.1.253 # vyos002, BACKUP ssh vyos@10.0.1.252 # vyos001, MASTERIf either will not connect, stop. Do not unplug the USG.
-
Physically disconnect the USG. Not just powered off — disconnected. The switch script refuses to run while anything still answers on a gateway address, because two devices on
.1is the worst available outcome. You cannot switch first and unplug after, for exactly that reason. -
In the vyos002 (BACKUP) session, first:
sudo /config/vyos-unifi-switch vyos -
Watch the health checks. They cover PPPoE, the default route, kea, the DNS forwarder and reachability. On failure the script reverts immediately and tells you so.
-
If vyos002 came up clean, repeat on vyos001 (MASTER).
-
Check a real client: does it get an address, and is it the same address as before? Every active client has a reservation, so it should be.
What will probably go wrong first
The WAN. There are two, and they behave differently:
| line | VLAN | transport | notes | |
|---|---|---|---|---|
| WAN2 | 10 gig ISP | 53 | DHCP, public 87.192.101.48/21 |
primary, distance 1 |
| WAN1 | Vodafone | 51 | PPPoE, ~900/700 Mbit | failover, distance 10 |
VyOS clones the USG's WAN2 MAC (f0:9f:c2:12:9b:4f) on bond0.53, which is how
it keeps the existing public lease rather than asking for a new one.
Both boxes carry the identical WAN and NAT config. vyos002's WAN interfaces are simply held administratively down, so the cloned MAC is never live on two boxes at once. To move the internet path to vyos002:
configure
delete interfaces bonding bond0 vif 53 disable
delete interfaces pppoe pppoe0 disable
commit; save
Two lines. Do it only when vyos001 is genuinely down or disconnected — two boxes holding that MAC at once is exactly what the disable prevents.
PPPoE is no longer an unknown: it was proven on the USG before cutover
(pppoe0 came up with 90.241.226.213, MTU 1492). What remains untested is
VyOS dialling it, and whether the ISP hands the same lease to the cloned MAC.
If the WAN check fails:
show interfaces pppoe pppoe0
sudo journalctl -u ppp@pppoe0 -n 50 --no-pager
Check the credential in /config/wan-secrets and that VLAN 51 actually reaches
the box. If it will not come up, run vyos-unifi-switch unifi, reconnect the
USG, and debug with the internet back on.
Things that are true and easy to forget
- WiFi keeps working, but through VyOS. The SSIDs stay in UniFi and the APs are untouched, but 37 of 83 active clients are wireless and every one is on LoT — they get their addresses from VyOS now.
- DHCP leases last 24h (86400s). A device that does not renew promptly keeps its old address for a while. That is fine, not a symptom.
- The firewalls resolve via
8.8.8.8/8.8.4.4— matching the DNS the USG used on its WAN. This means their own name resolution now depends on the internet being up, so between unplugging the USG and PPPoE establishing, the boxes have no DNS at all. That is expected and harmless: they only need DNS for NTP hostnames, and the switch's own health checks use it precisely to prove the WAN came up. Nothing in the switch itself resolves a name. - Internal
ad.itaz.eunames still resolve through Google, because that zone is published publicly with private addresses in it (nas001→10.0.0.194,kvm-macstudio1→192.168.3.8). Convenient here; worth knowing it is public. - The USG was a DNS resolver for every VLAN except LoT. VyOS now runs
dns forwardingin its place. If names stop resolving but IPs still work, that is where to look. eth2andbond0.2are both in192.168.8.0/23. It works, but if you see odd source-address behaviour on the management NIC, that is why.
Afterwards
Once it has been stable for a day:
- Re-run
migration/unifi-export.py— the UniFi controller is no longer the source of truth for DHCP, and the export will drift. - The VPN rules (ESP, UDP 500/4500) are carried over but the VPN itself still terminated on the USG. Decide whether it moves.
labsimstill holds a deliberatekvm→k8sdrop rule from earlier testing.