Files
lab/migration/CUTOVER.md
Michal f81c94af43 feat(migration): dual WAN, cloned MAC, and the new 10.8.0.0/23 Private VLAN
Two corrections from reading the live USG instead of trusting UniFi's fields,
which report wan_type=dhcp for both WANs and are simply wrong:

  - There are TWO WANs, not one. WAN2 is the 10 gig ISP on VLAN 53, plain DHCP
    with a PUBLIC address (87.192.101.48/21, gw 87.192.96.1) on the USG's eth2 --
    and it is what actually carries traffic. WAN1 is Vodafone PPPoE on VLAN 51,
    the failover. The delta had PPPoE as the only WAN, which would have left the
    primary line unconfigured.
  - The DHCP lease is bound to MAC, so bond0.53 now clones the USG's WAN2 MAC
    (f0:9f:c2:12:9b:4f). That is how VyOS keeps the existing public lease rather
    than negotiating a new one -- or getting none, if the ISP allows one per
    line. Distances: 10 gig at 1, Vodafone at 10.

Only ONE box may hold the cloned MAC, so --with-wan gates the entire WAN, NAT
and firewall section. vyos001 gets it (320 set lines); vyos002 gets none (234,
zero WAN/NAT/firewall) and routes the LAN only. Pretending both could hold it
would have meant a duplicate MAC on VLAN 53 and a flapping switch table.

Private was rebuilt at 10.8.0.0/23 (VLAN 9) after the old 10.0.8.0/23 was
deleted. bond0.9 and the VRRP group were moved to 10.8.0.252/.253 with VIP
10.8.0.254 on both boxes, and the delta now targets 10.8.0.1.

Creating that network first required breaking a deadlock in UniFi: every LAN
write was rejected with api.err.WanIpOverlapped / 0.0.0.0/0, because WAN1 was
set to DHCP on a line that only speaks PPPoE, so it sat at 0.0.0.0 forever and
the validator treated that as a subnet overlapping everything. Verified
server-side, not a UI bug -- the API rejected it identically. Setting
wan_type=pppoe let it dial (90.241.226.213, MTU 1492), which cleared the phantom
overlap and incidentally PROVED the Vodafone credentials and line work, which
had been listed as untestable before cutover.

dhcp-options no-default-route-dns does not exist; the valid set is client-id,
default-route-distance, host-name, mtu, no-default-route, reject, user-class,
vendor-class-id. Caught by validating the delta against vyos001's real config on
the labsim router before installing.

After adding the network, the gateway's dhcpd.conf was checked with
`dhcpd3 -t -cf` (valid) and confirmed to contain the new subnet only after a
force-provision -- controller state is not device state.

Both boxes: mode unifi, VRRP unchanged, unifi.boot re-captured (232 lines,
carrying the new VLAN 9), no config drift.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01DMVzWZgiKW2wquf5z8S1yH
2026-08-16 18:16:25 +01:00

8.2 KiB

Cutover runbook — USG to VyOS

Print this. During the cutover there is no internet, so there is no assistant and no web search. Everything you need is on this page and on the boxes themselves.

Use these addresses. Not the other ones.

use this do NOT use
vyos001 (MASTER) 10.0.1.252 192.168.8.143
vyos002 (BACKUP) 10.0.1.253 192.168.8.144

ssh vyos@10.0.1.252 — by IP, not by name.

The 192.168.8.x addresses stop working the instant the USG is unplugged. That is not a maybe. Your workstation is on LoT (10.0.0.210/23) and reaching 192.168.8.x requires routing through the USG:

ip route get 192.168.8.143  ->  via 10.0.0.1     <- the USG. Gone.
ip route get 10.0.1.252     ->  dev lanbr0       <- same L2. Survives.

10.0.1.252 and .253 are on the LoT VLAN, the same broadcast domain as your workstation, so they need no gateway at all. They are the only remote path that survives the cutover.

Between unplugging the USG and finishing the switch there is no inter-VLAN routing. In that window:

  • the JetKVMs are unreachable from your workstation (they are on Management and kvm) — they are not a fallback during the gap
  • Tailscale is down with the internet
  • your workstation keeps 10.0.0.210 (86400s lease) and can still resolve via 10.0.0.194, which is also link-scope

If LoT SSH fails, the next step is physical console, not the network.

JetKVMs (after routing is restored) 192.168.1.28, 192.168.1.29, 192.168.3.6
Switch script /config/vyos-unifi-switch on each box
Login user vyos

If something is wrong, do this

sudo /config/vyos-unifi-switch unifi

Then reconnect the USG. That command runs no health checks, asks nothing and cannot refuse. It restores a byte-exact copy of the configuration the box had before the cutover — verified by diff, not by assumption.

You do not have to be quick. If you do nothing at all after vyos-unifi-switch vyos, the box reverts by itself within 10 minutes. Verified: config returns to the previous state and the box does not reboot (uptime and boot-id unchanged across an auto-revert).


What has actually been tested

Proven on the labsim router (same VyOS version, isolated OVS bridge with no physical NIC), by loading vyos001's real running config and applying the real production delta:

  • All 317 commands accepted, and the whole delta commits (COMMIT OK).
  • unifi mode restores the previous config byte-exact (diff clean).
  • Auto-revert fires when the commit is not confirmed: config returns to the saved state and the box does not reboot — uptime and boot-id unchanged across the revert.
  • Failed health checks trigger an immediate revert rather than waiting out the timer.

Two bugs were found this way and would each have failed the entire switch, since the delta commits as one unit: bond0.51 did not exist for PPPoE to reference, and translation port rejects a port list.

Not tested, and untestable in advance:

  • PPPoE. The line permits one session and the USG holds it. The first real attempt is during the cutover.
  • The commit on the real boxes. The rehearsal ran with vyos001's eth2 and eth3 stanzas stripped, because the sim VM has only two NICs. Those are plain interface configs that already work on the real hardware, but they were not part of what committed.

Before you unplug anything

  1. Tether your workstation to your phone if you want the assistant available. Cutting the USG cuts your internet, not your LAN.
  2. On both boxes, confirm the machinery is present:
    sudo /config/vyos-unifi-switch status
    ls -la /config/modes/          # unifi.boot + to-vyos.commands
    ls -la /config/wan-secrets     # must be 0600
    
    status must report mode: unifi. If unifi.boot is missing, stop — there is no way back without it.
  3. Confirm the revert action is reload, not reboot:
    show configuration commands | match commit-confirm
    
    Must show action 'reload'. Without it a failed switch reboots the firewall instead of reverting it. The switch script refuses to run if this is missing, but check anyway.

The cutover

  1. Open both SSH sessions BEFORE you unplug anything, and leave them open:

    ssh vyos@10.0.1.253      # vyos002, BACKUP
    ssh vyos@10.0.1.252      # vyos001, MASTER
    

    If either will not connect, stop. Do not unplug the USG.

  2. Physically disconnect the USG. Not just powered off — disconnected. The switch script refuses to run while anything still answers on a gateway address, because two devices on .1 is the worst available outcome. You cannot switch first and unplug after, for exactly that reason.

  3. In the vyos002 (BACKUP) session, first:

    sudo /config/vyos-unifi-switch vyos
    
  4. Watch the health checks. They cover PPPoE, the default route, kea, the DNS forwarder and reachability. On failure the script reverts immediately and tells you so.

  5. If vyos002 came up clean, repeat on vyos001 (MASTER).

  6. Check a real client: does it get an address, and is it the same address as before? Every active client has a reservation, so it should be.

What will probably go wrong first

The WAN. There are two, and they behave differently:

line VLAN transport notes
WAN2 10 gig ISP 53 DHCP, public 87.192.101.48/21 primary, distance 1
WAN1 Vodafone 51 PPPoE, ~900/700 Mbit failover, distance 10

VyOS clones the USG's WAN2 MAC (f0:9f:c2:12:9b:4f) on bond0.53, which is how it keeps the existing public lease rather than asking for a new one. Only vyos001 carries the WAN — the MAC must be unique, so vyos002 routes the LAN and holds the VIPs but has no internet path until the WAN is moved deliberately.

PPPoE is no longer an unknown: it was proven on the USG before cutover (pppoe0 came up with 90.241.226.213, MTU 1492). What remains untested is VyOS dialling it, and whether the ISP hands the same lease to the cloned MAC.

If the WAN check fails:

show interfaces pppoe pppoe0
sudo journalctl -u ppp@pppoe0 -n 50 --no-pager

Check the credential in /config/wan-secrets and that VLAN 51 actually reaches the box. If it will not come up, run vyos-unifi-switch unifi, reconnect the USG, and debug with the internet back on.

Things that are true and easy to forget

  • WiFi keeps working, but through VyOS. The SSIDs stay in UniFi and the APs are untouched, but 37 of 83 active clients are wireless and every one is on LoT — they get their addresses from VyOS now.
  • DHCP leases last 24h (86400s). A device that does not renew promptly keeps its old address for a while. That is fine, not a symptom.
  • The firewalls resolve via 8.8.8.8 / 8.8.4.4 — matching the DNS the USG used on its WAN. This means their own name resolution now depends on the internet being up, so between unplugging the USG and PPPoE establishing, the boxes have no DNS at all. That is expected and harmless: they only need DNS for NTP hostnames, and the switch's own health checks use it precisely to prove the WAN came up. Nothing in the switch itself resolves a name.
  • Internal ad.itaz.eu names still resolve through Google, because that zone is published publicly with private addresses in it (nas00110.0.0.194, kvm-macstudio1192.168.3.8). Convenient here; worth knowing it is public.
  • The USG was a DNS resolver for every VLAN except LoT. VyOS now runs dns forwarding in its place. If names stop resolving but IPs still work, that is where to look.
  • eth2 and bond0.2 are both in 192.168.8.0/23. It works, but if you see odd source-address behaviour on the management NIC, that is why.

Afterwards

Once it has been stable for a day:

  • Re-run migration/unifi-export.py — the UniFi controller is no longer the source of truth for DHCP, and the export will drift.
  • The VPN rules (ESP, UDP 500/4500) are carried over but the VPN itself still terminated on the USG. Decide whether it moves.
  • labsim still holds a deliberate kvm→k8s drop rule from earlier testing.