eth3 <-> eth3 direct cable is in. Both ends negotiated 2500Mb full duplex -- these are 2.5 GbE ports, not the 1G I had assumed. Carrier alone proves nothing, so the link was tested end to end with temporary kernel-level addresses (never committed to VyOS config, removed afterwards): 3/3 packets, 0% loss, 0.371ms average. A cable can show carrier and still not pass traffic; now it is known to. Deltas regenerated with --conntrack-link and installed on both boxes: vyos001 nat=21 fw=58 conntrack=9 eth3=10.255.255.1/30 disable=0 vyos002 nat=21 fw=58 conntrack=9 eth3=10.255.255.2/30 disable=2 Identical apart from the peer /30, the VRRP/DHCP-HA roles, and the two disable lines holding vyos002's WAN down. Both still report mode=unifi and nothing about their behaviour has changed -- eth3 carries no address in the running config, and conntrack-sync appears only in the delta, which is applied at cutover. Not verified: multicast on the peer link. `ping -I eth3 224.0.0.1` drew no responders, but that is the all-hosts group which VyOS need not answer, so it proves nothing either way. conntrack-sync's own multicast (225.0.0.50) was proven working in labsim over bond0.10, and this is a point-to-point link, so the risk is low -- but it is untested on this specific cable and worth watching at cutover. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01DMVzWZgiKW2wquf5z8S1yH
210 lines
8.6 KiB
Markdown
210 lines
8.6 KiB
Markdown
# Cutover runbook — USG to VyOS
|
|
|
|
**Print this.** During the cutover there is no internet, so there is no
|
|
assistant and no web search. Everything you need is on this page and on the
|
|
boxes themselves.
|
|
|
|
## Use these addresses. Not the other ones.
|
|
|
|
| | use this | do NOT use |
|
|
|---|---|---|
|
|
| vyos001 (MASTER) | **`10.0.1.252`** | ~~192.168.8.143~~ |
|
|
| vyos002 (BACKUP) | **`10.0.1.253`** | ~~192.168.8.144~~ |
|
|
|
|
`ssh vyos@10.0.1.252` — by IP, not by name.
|
|
|
|
**The `192.168.8.x` addresses stop working the instant the USG is unplugged.**
|
|
That is not a maybe. Your workstation is on LoT (`10.0.0.210/23`) and reaching
|
|
`192.168.8.x` requires routing *through the USG*:
|
|
|
|
```
|
|
ip route get 192.168.8.143 -> via 10.0.0.1 <- the USG. Gone.
|
|
ip route get 10.0.1.252 -> dev lanbr0 <- same L2. Survives.
|
|
```
|
|
|
|
`10.0.1.252` and `.253` are on the LoT VLAN, the same broadcast domain as your
|
|
workstation, so they need no gateway at all. They are the only remote path that
|
|
survives the cutover.
|
|
|
|
**Between unplugging the USG and finishing the switch there is no inter-VLAN
|
|
routing.** In that window:
|
|
|
|
- the **JetKVMs are unreachable** from your workstation (they are on Management
|
|
and kvm) — they are *not* a fallback during the gap
|
|
- **Tailscale is down** with the internet
|
|
- your workstation keeps `10.0.0.210` (86400s lease) and can still resolve via
|
|
`10.0.0.194`, which is also link-scope
|
|
|
|
If LoT SSH fails, the next step is physical console, not the network.
|
|
|
|
| | |
|
|
|---|---|
|
|
| JetKVMs (after routing is restored) | `192.168.1.28`, `192.168.1.29`, `192.168.3.6` |
|
|
| Switch script | `/config/vyos-unifi-switch` on each box |
|
|
| Peer link | `eth3` ↔ `eth3` direct cable, 2.5 GbE, `10.255.255.0/30` — conntrack state sync |
|
|
| Login | user `vyos` |
|
|
|
|
---
|
|
|
|
## If something is wrong, do this
|
|
|
|
```
|
|
sudo /config/vyos-unifi-switch unifi
|
|
```
|
|
|
|
Then reconnect the USG. That command runs no health checks, asks nothing and
|
|
cannot refuse. It restores a byte-exact copy of the configuration the box had
|
|
before the cutover — verified by diff, not by assumption.
|
|
|
|
**You do not have to be quick.** If you do nothing at all after
|
|
`vyos-unifi-switch vyos`, the box reverts by itself within 10 minutes. Verified:
|
|
config returns to the previous state and the box does **not** reboot
|
|
(`uptime` and boot-id unchanged across an auto-revert).
|
|
|
|
---
|
|
|
|
## What has actually been tested
|
|
|
|
Proven on the labsim router (same VyOS version, isolated OVS bridge with no
|
|
physical NIC), by loading **vyos001's real running config** and applying the
|
|
**real production delta**:
|
|
|
|
- All 317 commands accepted, and the whole delta **commits** (`COMMIT OK`).
|
|
- `unifi` mode restores the previous config **byte-exact** (diff clean).
|
|
- Auto-revert fires when the commit is not confirmed: config returns to the
|
|
saved state and the box does **not** reboot — `uptime` and boot-id unchanged
|
|
across the revert.
|
|
- Failed health checks trigger an immediate revert rather than waiting out the
|
|
timer.
|
|
|
|
Two bugs were found this way and would each have failed the entire switch,
|
|
since the delta commits as one unit: `bond0.51` did not exist for PPPoE to
|
|
reference, and `translation port` rejects a port list.
|
|
|
|
**Not tested, and untestable in advance:**
|
|
|
|
- **PPPoE.** The line permits one session and the USG holds it. The first real
|
|
attempt is during the cutover.
|
|
- **The commit on the real boxes.** The rehearsal ran with vyos001's `eth2` and
|
|
`eth3` stanzas stripped, because the sim VM has only two NICs. Those are
|
|
plain interface configs that already work on the real hardware, but they were
|
|
not part of what committed.
|
|
|
|
## Before you unplug anything
|
|
|
|
1. Tether your workstation to your phone if you want the assistant available.
|
|
Cutting the USG cuts your internet, not your LAN.
|
|
2. On **both** boxes, confirm the machinery is present:
|
|
```
|
|
sudo /config/vyos-unifi-switch status
|
|
ls -la /config/modes/ # unifi.boot + to-vyos.commands
|
|
ls -la /config/wan-secrets # must be 0600
|
|
```
|
|
`status` must report `mode: unifi`. If `unifi.boot` is missing, **stop** —
|
|
there is no way back without it.
|
|
3. Confirm the revert action is `reload`, not `reboot`:
|
|
```
|
|
show configuration commands | match commit-confirm
|
|
```
|
|
Must show `action 'reload'`. Without it a failed switch **reboots** the
|
|
firewall instead of reverting it. The switch script refuses to run if this
|
|
is missing, but check anyway.
|
|
|
|
## The cutover
|
|
|
|
0. **Open both SSH sessions BEFORE you unplug anything**, and leave them open:
|
|
```
|
|
ssh vyos@10.0.1.253 # vyos002, BACKUP
|
|
ssh vyos@10.0.1.252 # vyos001, MASTER
|
|
```
|
|
If either will not connect, stop. Do not unplug the USG.
|
|
|
|
1. **Physically disconnect the USG.** Not just powered off — disconnected. The
|
|
switch script refuses to run while anything still answers on a gateway
|
|
address, because two devices on `.1` is the worst available outcome.
|
|
You cannot switch first and unplug after, for exactly that reason.
|
|
|
|
2. In the **vyos002 (BACKUP)** session, first:
|
|
```
|
|
sudo /config/vyos-unifi-switch vyos
|
|
```
|
|
3. Watch the health checks. They cover PPPoE, the default route, kea, the DNS
|
|
forwarder and reachability. On failure the script reverts immediately and
|
|
tells you so.
|
|
4. If vyos002 came up clean, repeat on **vyos001 (MASTER)**.
|
|
5. Check a real client: does it get an address, and is it the *same* address as
|
|
before? Every active client has a reservation, so it should be.
|
|
|
|
## What will probably go wrong first
|
|
|
|
**The WAN.** There are two, and they behave differently:
|
|
|
|
| | line | VLAN | transport | notes |
|
|
|---|---|---|---|---|
|
|
| WAN2 | 10 gig ISP | **53** | DHCP, public `87.192.101.48/21` | primary, distance 1 |
|
|
| WAN1 | Vodafone | **51** | PPPoE, ~900/700 Mbit | failover, distance 10 |
|
|
|
|
VyOS clones the USG's WAN2 MAC (`f0:9f:c2:12:9b:4f`) on `bond0.53`, which is how
|
|
it keeps the existing public lease rather than asking for a new one.
|
|
|
|
**Both boxes carry the identical WAN and NAT config.** vyos002's WAN interfaces
|
|
are simply held administratively down, so the cloned MAC is never live on two
|
|
boxes at once. To move the internet path to vyos002:
|
|
|
|
```
|
|
configure
|
|
delete interfaces bonding bond0 vif 53 disable
|
|
delete interfaces pppoe pppoe0 disable
|
|
commit; save
|
|
```
|
|
|
|
Two lines. Do it only when vyos001 is genuinely down or disconnected — two boxes
|
|
holding that MAC at once is exactly what the disable prevents.
|
|
|
|
PPPoE is no longer an unknown: it was proven on the USG before cutover
|
|
(`pppoe0` came up with `90.241.226.213`, MTU 1492). What remains untested is
|
|
VyOS dialling it, and whether the ISP hands the same lease to the cloned MAC.
|
|
|
|
If the WAN check fails:
|
|
|
|
```
|
|
show interfaces pppoe pppoe0
|
|
sudo journalctl -u ppp@pppoe0 -n 50 --no-pager
|
|
```
|
|
|
|
Check the credential in `/config/wan-secrets` and that VLAN 51 actually reaches
|
|
the box. If it will not come up, run `vyos-unifi-switch unifi`, reconnect the
|
|
USG, and debug with the internet back on.
|
|
|
|
## Things that are true and easy to forget
|
|
|
|
- **WiFi keeps working, but through VyOS.** The SSIDs stay in UniFi and the APs
|
|
are untouched, but 37 of 83 active clients are wireless and every one is on
|
|
LoT — they get their addresses from VyOS now.
|
|
- **DHCP leases last 24h (86400s).** A device that does not renew promptly keeps
|
|
its old address for a while. That is fine, not a symptom.
|
|
- **The firewalls resolve via `8.8.8.8` / `8.8.4.4`** — matching the DNS the USG
|
|
used on its WAN. This means their own name resolution now depends on the
|
|
*internet* being up, so between unplugging the USG and PPPoE establishing,
|
|
the boxes have no DNS at all. That is expected and harmless: they only need
|
|
DNS for NTP hostnames, and the switch's own health checks use it precisely to
|
|
prove the WAN came up. Nothing in the switch itself resolves a name.
|
|
- Internal `ad.itaz.eu` names still resolve through Google, because that zone is
|
|
published publicly with private addresses in it (`nas001` → `10.0.0.194`,
|
|
`kvm-macstudio1` → `192.168.3.8`). Convenient here; worth knowing it is public.
|
|
- **The USG was a DNS resolver** for every VLAN except LoT. VyOS now runs
|
|
`dns forwarding` in its place. If names stop resolving but IPs still work,
|
|
that is where to look.
|
|
- **`eth2` and `bond0.2` are both in `192.168.8.0/23`.** It works, but if you
|
|
see odd source-address behaviour on the management NIC, that is why.
|
|
|
|
## Afterwards
|
|
|
|
Once it has been stable for a day:
|
|
|
|
- Re-run `migration/unifi-export.py` — the UniFi controller is no longer the
|
|
source of truth for DHCP, and the export will drift.
|
|
- The VPN rules (ESP, UDP 500/4500) are carried over but the VPN itself still
|
|
terminated on the USG. Decide whether it moves.
|
|
- `labsim` still holds a deliberate `kvm→k8s` drop rule from earlier testing.
|