Files
lab/labsim/README.md

213 lines
9.4 KiB
Markdown
Raw Normal View History

feat(labsim): libvirt replica of the lab network with LACP + VyOS routing A throwaway copy of the production VLAN topology so routing and firewall changes can be tested before they touch the real network. Same VLAN IDs and roles as UniFi, deliberately different ranges (172.31.<vlan>.0/24) so nothing here can be mistaken for production. - OVS fabric: real 802.1Q. Access port per micro VM, host leg per VLAN (.2, for SSH only — NOT the VMs' default route, so inter-VLAN tests exercise the router rather than the host's routing table), and a trunk portgroup with VLAN 1 declared nativeMode='untagged'. - Six Alpine micro VMs (256MB, copy-on-write overlays on one 176MB image), SSH + a hello-world HTTP page naming the VLAN. - VyOS router installed to disk unattended over the console, with the SAME config shape as the VP2440s: two NICs in an LACP bond carrying the trunk, VLAN 1 native, bond0.<vlan> holding the .1 gateway on each. - labsim-matrix.py: full-mesh ICMP/TCP22/TCP80 probe, ~0.2s, --watch highlights cells that changed since the last sweep. Guest-side probe is python3 (already present via cloud-init) so nothing is installed on VMs that have no internet. - Prometheus + Grafana (anonymous auth, no login) with a provisioned dashboard: heatmap plus a state timeline showing exactly when a path flipped. Verified end to end: one VyOS rule took sum(labsim_reachable) from 90 to 84, blocking precisely kvm<->k8s across all three protocols. Traps found building this, all now encoded in the scripts: - virtio-net breaks 802.3ad: the guest's bonding driver reports slaves "MII Status: down" despite carrier=1 and never sends an LACPDU, so the bond sits in AD_STATE_DEFAULTED. e1000e fixes it with no other change. Matches the netdev thread "bonding (IEEE 802.3ad) not working with qemu/virtio". - OVS defaults bonds to active-backup, which does not speak LACP at all — bond_mode=balance-tcp is required. - LACP deadlock: OVS holds members disabled until negotiation while the partner needs carrier before it will send LACPDUs. lacp-fallback-ab breaks it. - LACPDUs are untagged, so a trunk with no native VLAN has nowhere to put them. - --boot cdrom,hd re-runs the ISO on every restart, so every commit+save went to a live system that evaporated. Install now switches the VM to boot hd. - cloud-init on Alpine: users stay locked without lock_passwd:false, one failing runcmd aborts the rest, busybox here has no httpd applet, and start-stop-daemon --exec /usr/bin/python3 matches cloud-init's own python3. - The user-data heredoc is unquoted, so backticks in a COMMENT were executed by the host shell and their output corrupted the YAML. build_seed now validates with yaml.safe_load before building the ISO. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01DMVzWZgiKW2wquf5z8S1yH
2026-08-13 00:42:39 +01:00
# labsim — libvirt replica of the lab network
A throwaway copy of the production VLAN topology for testing routing, firewall
rules and failover **without touching the real network**. Same VLAN IDs and
roles as UniFi, deliberately different IP ranges so nothing can be confused for
production.
## Topology
Each VLAN is its own isolated libvirt network with one tiny Alpine VM on it.
| VLAN | Name | Sim subnet | VM address | Mirrors production |
|-----:|------|------------|-----------|--------------------|
| 1 | management | 172.31.1.0/24 | 172.31.1.10 | 192.168.1.0/24 |
| 2 | k8s | 172.31.2.0/24 | 172.31.2.10 | 192.168.8.0/23 |
| 3 | kvm | 172.31.3.0/24 | 172.31.3.10 | 192.168.3.0/24 |
| 9 | private | 172.31.9.0/24 | 172.31.9.10 | 10.0.9.0/23 |
feat(migration): export UniFi config and generate VyOS DHCP+DNS from it Groundwork for replacing the USG with the VyOS pair without anything on the network noticing. Three pieces: migration/unifi-export.py pulls 13 endpoints off the classic controller into timestamped JSON plus a normalised inventory: 11 networks, 31 DHCP reservations, 4 port forwards, 2 firewall rules, 0 static routes. Two things this turned up that a naive export would have lost: - 23 of the 31 reservations carry no network_id at all -- UniFi simply does not store the binding -- so they are resolved by subnet containment instead. Without that, three quarters of the reservations have no subnet to be placed in. - 30 of the 31 sit INSIDE the DHCP pool, which UniFi's dhcpd tolerates and which is flagged as a warning rather than discovered at cutover. migration/unifi-to-vyos.py turns that inventory into VyOS `set` commands for DHCP and DNS only -- the services the USG owns that VyOS must reproduce. Not a general converter. --prod and --sim come from one code path so the config proven in the sim and the config applied to the firewalls cannot drift. Prod mode hard-fails if any reservation is missing, since a silent drop is the failure mode that matters. DNS is included because the USG resolves for 5 of 6 VLANs today: UniFi hands out the gateway's own address whenever dhcpd_dns is empty, verified by labmaster resolving against 192.168.8.1. Replacing the USG without a forwarder would take DNS away from those VLANs entirely. labsim/labsim-dhcp-test.sh proves it by booting throwaway VMs with real production MACs -- the one piece of production config that transplants verbatim. Safe because ovs-labsim has no physical NIC, so those MACs cannot reach the real LAN. Result on VyOS 2026.08 (kea), 4/4: printer1 got 172.31.10.46 from inside the pool, sonoff-matter got 172.31.11.67 across the /23 boundary, Hubitat got its out-of-pool .2, and an unreserved MAC got an unreserved address. kea honours in-pool host reservations -- the open question blocking the cutover. Supporting changes to labsim: - VLAN 10 widened to /23. Every reservation is in LoT and LoT spans 10.0.0.x and 10.0.1.x, which a /24 cannot represent. - LoT's host leg moved to .3, because 10.0.0.2 is a real reservation (Hubitat) that maps onto the host's own address. - vlans.conf gained optional masklen and host_octet fields, defaulting to 24 and 2 so the other five VLANs are untouched. - Fixed /etc/network/interfaces hardcoding 255.255.255.0. That file is what actually takes effect on these Alpine guests -- cloud-init's network-config is ignored -- so any non-/24 VLAN was silently wrong. Two generator bugs found by VyOS rejecting the output: static-mapping names are validated as hostnames, so underscores fail; and two devices named "espressif" plus two named "thebeast" collided into single names, which would have overwritten one reservation with another's address. The raw export holds WiFi passphrases and the WAN PPPoE credentials and is gitignored. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01DMVzWZgiKW2wquf5z8S1yH
2026-08-15 01:33:00 +01:00
| 10 | lot | **172.31.10.0/23** | 172.31.10.10 | 10.0.0.0/23 |
feat(labsim): libvirt replica of the lab network with LACP + VyOS routing A throwaway copy of the production VLAN topology so routing and firewall changes can be tested before they touch the real network. Same VLAN IDs and roles as UniFi, deliberately different ranges (172.31.<vlan>.0/24) so nothing here can be mistaken for production. - OVS fabric: real 802.1Q. Access port per micro VM, host leg per VLAN (.2, for SSH only — NOT the VMs' default route, so inter-VLAN tests exercise the router rather than the host's routing table), and a trunk portgroup with VLAN 1 declared nativeMode='untagged'. - Six Alpine micro VMs (256MB, copy-on-write overlays on one 176MB image), SSH + a hello-world HTTP page naming the VLAN. - VyOS router installed to disk unattended over the console, with the SAME config shape as the VP2440s: two NICs in an LACP bond carrying the trunk, VLAN 1 native, bond0.<vlan> holding the .1 gateway on each. - labsim-matrix.py: full-mesh ICMP/TCP22/TCP80 probe, ~0.2s, --watch highlights cells that changed since the last sweep. Guest-side probe is python3 (already present via cloud-init) so nothing is installed on VMs that have no internet. - Prometheus + Grafana (anonymous auth, no login) with a provisioned dashboard: heatmap plus a state timeline showing exactly when a path flipped. Verified end to end: one VyOS rule took sum(labsim_reachable) from 90 to 84, blocking precisely kvm<->k8s across all three protocols. Traps found building this, all now encoded in the scripts: - virtio-net breaks 802.3ad: the guest's bonding driver reports slaves "MII Status: down" despite carrier=1 and never sends an LACPDU, so the bond sits in AD_STATE_DEFAULTED. e1000e fixes it with no other change. Matches the netdev thread "bonding (IEEE 802.3ad) not working with qemu/virtio". - OVS defaults bonds to active-backup, which does not speak LACP at all — bond_mode=balance-tcp is required. - LACP deadlock: OVS holds members disabled until negotiation while the partner needs carrier before it will send LACPDUs. lacp-fallback-ab breaks it. - LACPDUs are untagged, so a trunk with no native VLAN has nowhere to put them. - --boot cdrom,hd re-runs the ISO on every restart, so every commit+save went to a live system that evaporated. Install now switches the VM to boot hd. - cloud-init on Alpine: users stay locked without lock_passwd:false, one failing runcmd aborts the rest, busybox here has no httpd applet, and start-stop-daemon --exec /usr/bin/python3 matches cloud-init's own python3. - The user-data heredoc is unquoted, so backticks in a COMMENT were executed by the host shell and their output corrupted the YAML. build_seed now validates with yaml.safe_load before building the ISO. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01DMVzWZgiKW2wquf5z8S1yH
2026-08-13 00:42:39 +01:00
| 200 | roomates | 172.31.200.0/24 | 172.31.200.10 | 192.168.2.0/24 |
feat(migration): export UniFi config and generate VyOS DHCP+DNS from it Groundwork for replacing the USG with the VyOS pair without anything on the network noticing. Three pieces: migration/unifi-export.py pulls 13 endpoints off the classic controller into timestamped JSON plus a normalised inventory: 11 networks, 31 DHCP reservations, 4 port forwards, 2 firewall rules, 0 static routes. Two things this turned up that a naive export would have lost: - 23 of the 31 reservations carry no network_id at all -- UniFi simply does not store the binding -- so they are resolved by subnet containment instead. Without that, three quarters of the reservations have no subnet to be placed in. - 30 of the 31 sit INSIDE the DHCP pool, which UniFi's dhcpd tolerates and which is flagged as a warning rather than discovered at cutover. migration/unifi-to-vyos.py turns that inventory into VyOS `set` commands for DHCP and DNS only -- the services the USG owns that VyOS must reproduce. Not a general converter. --prod and --sim come from one code path so the config proven in the sim and the config applied to the firewalls cannot drift. Prod mode hard-fails if any reservation is missing, since a silent drop is the failure mode that matters. DNS is included because the USG resolves for 5 of 6 VLANs today: UniFi hands out the gateway's own address whenever dhcpd_dns is empty, verified by labmaster resolving against 192.168.8.1. Replacing the USG without a forwarder would take DNS away from those VLANs entirely. labsim/labsim-dhcp-test.sh proves it by booting throwaway VMs with real production MACs -- the one piece of production config that transplants verbatim. Safe because ovs-labsim has no physical NIC, so those MACs cannot reach the real LAN. Result on VyOS 2026.08 (kea), 4/4: printer1 got 172.31.10.46 from inside the pool, sonoff-matter got 172.31.11.67 across the /23 boundary, Hubitat got its out-of-pool .2, and an unreserved MAC got an unreserved address. kea honours in-pool host reservations -- the open question blocking the cutover. Supporting changes to labsim: - VLAN 10 widened to /23. Every reservation is in LoT and LoT spans 10.0.0.x and 10.0.1.x, which a /24 cannot represent. - LoT's host leg moved to .3, because 10.0.0.2 is a real reservation (Hubitat) that maps onto the host's own address. - vlans.conf gained optional masklen and host_octet fields, defaulting to 24 and 2 so the other five VLANs are untouched. - Fixed /etc/network/interfaces hardcoding 255.255.255.0. That file is what actually takes effect on these Alpine guests -- cloud-init's network-config is ignored -- so any non-/24 VLAN was silently wrong. Two generator bugs found by VyOS rejecting the output: static-mapping names are validated as hostnames, so underscores fail; and two devices named "espressif" plus two named "thebeast" collided into single names, which would have overwritten one reservation with another's address. The raw export holds WiFi passphrases and the WAN PPPoE credentials and is gitignored. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01DMVzWZgiKW2wquf5z8S1yH
2026-08-15 01:33:00 +01:00
The sim subnet encodes the VLAN id: `172.31.<vlan>.0/24`, with one exception.
**VLAN 10 is a `/23`** because every UniFi DHCP reservation lives in LoT and LoT
spans `10.0.0.x` *and* `10.0.1.x`, which a `/24` cannot hold. The mapping stays
readable — `10.0.0.46 → 172.31.10.46`, `10.0.1.67 → 172.31.11.67`.
LoT's host leg is `.3`, not `.2`, because `10.0.0.2` is a real reservation
(Hubitat) that maps onto `172.31.10.2`. `.3` is unreserved and sits below the
DHCP pool, so it can never be handed out.
`vlans.conf` therefore takes two optional trailing fields:
```
vlan_id:name:sim_prefix:real_subnet[:masklen][:host_octet]
```
defaulting to `24` and `2`. k8s and Private are also `/23` in production but
hold no reservations, so they keep their `/24` and their DHCP range is clamped
— reported at generation time, never silently.
feat(labsim): libvirt replica of the lab network with LACP + VyOS routing A throwaway copy of the production VLAN topology so routing and firewall changes can be tested before they touch the real network. Same VLAN IDs and roles as UniFi, deliberately different ranges (172.31.<vlan>.0/24) so nothing here can be mistaken for production. - OVS fabric: real 802.1Q. Access port per micro VM, host leg per VLAN (.2, for SSH only — NOT the VMs' default route, so inter-VLAN tests exercise the router rather than the host's routing table), and a trunk portgroup with VLAN 1 declared nativeMode='untagged'. - Six Alpine micro VMs (256MB, copy-on-write overlays on one 176MB image), SSH + a hello-world HTTP page naming the VLAN. - VyOS router installed to disk unattended over the console, with the SAME config shape as the VP2440s: two NICs in an LACP bond carrying the trunk, VLAN 1 native, bond0.<vlan> holding the .1 gateway on each. - labsim-matrix.py: full-mesh ICMP/TCP22/TCP80 probe, ~0.2s, --watch highlights cells that changed since the last sweep. Guest-side probe is python3 (already present via cloud-init) so nothing is installed on VMs that have no internet. - Prometheus + Grafana (anonymous auth, no login) with a provisioned dashboard: heatmap plus a state timeline showing exactly when a path flipped. Verified end to end: one VyOS rule took sum(labsim_reachable) from 90 to 84, blocking precisely kvm<->k8s across all three protocols. Traps found building this, all now encoded in the scripts: - virtio-net breaks 802.3ad: the guest's bonding driver reports slaves "MII Status: down" despite carrier=1 and never sends an LACPDU, so the bond sits in AD_STATE_DEFAULTED. e1000e fixes it with no other change. Matches the netdev thread "bonding (IEEE 802.3ad) not working with qemu/virtio". - OVS defaults bonds to active-backup, which does not speak LACP at all — bond_mode=balance-tcp is required. - LACP deadlock: OVS holds members disabled until negotiation while the partner needs carrier before it will send LACPDUs. lacp-fallback-ab breaks it. - LACPDUs are untagged, so a trunk with no native VLAN has nowhere to put them. - --boot cdrom,hd re-runs the ISO on every restart, so every commit+save went to a live system that evaporated. Install now switches the VM to boot hd. - cloud-init on Alpine: users stay locked without lock_passwd:false, one failing runcmd aborts the rest, busybox here has no httpd applet, and start-stop-daemon --exec /usr/bin/python3 matches cloud-init's own python3. - The user-data heredoc is unquoted, so backticks in a COMMENT were executed by the host shell and their output corrupted the YAML. build_seed now validates with yaml.safe_load before building the ISO. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01DMVzWZgiKW2wquf5z8S1yH
2026-08-13 00:42:39 +01:00
Address plan, identical on every VLAN:
| Address | Role |
|---------|------|
| `.1` | gateway under test — a router VM you add (not created by default) |
| `.2` | host bridge — how you reach the VMs from this workstation |
| `.10` | the VLAN's micro VM |
| `.254` | reserved for a VRRP VIP, mirroring production |
The host sits at `.2` purely so you can SSH in. It is deliberately **not** the
VMs' default route — that is `.1` — so inter-VLAN tests fail loudly when no
router is present instead of being silently served by the host's own routing
table. libvirt also installs reject rules that stop these networks forwarding
to each other, so traffic between VLANs only works once a router VM bridges
them.
## Usage
```bash
./labsim-up.sh # bring up every VLAN (idempotent)
./labsim-up.sh 2 3 # only VLANs 2 and 3
./labsim-down.sh # destroy VMs + networks, keep the base image
./labsim-down.sh --purge # also delete the downloaded Alpine image
```
Each VM: 256 MB, 1 vCPU, a copy-on-write overlay on one shared 176 MB Alpine
image (so six VMs cost a few MB of disk, not 1 GB).
## Access
```bash
ssh alpine@172.31.2.10 # normal user (password: labsim)
ssh root@172.31.2.10 # privileged — this image has no sudo
curl http://172.31.2.10/ # hello-world page naming the VLAN
```
Console, when the network is the thing that is broken:
```bash
sudo virsh console labsim-2-k8s # root / labsim
```
## Watching it
```bash
./labsim-matrix.py --watch 2 # terminal grid, changed cells highlighted
./monitoring-up.sh # topology page + Prometheus + Grafana
```
feat(migration): export UniFi config and generate VyOS DHCP+DNS from it Groundwork for replacing the USG with the VyOS pair without anything on the network noticing. Three pieces: migration/unifi-export.py pulls 13 endpoints off the classic controller into timestamped JSON plus a normalised inventory: 11 networks, 31 DHCP reservations, 4 port forwards, 2 firewall rules, 0 static routes. Two things this turned up that a naive export would have lost: - 23 of the 31 reservations carry no network_id at all -- UniFi simply does not store the binding -- so they are resolved by subnet containment instead. Without that, three quarters of the reservations have no subnet to be placed in. - 30 of the 31 sit INSIDE the DHCP pool, which UniFi's dhcpd tolerates and which is flagged as a warning rather than discovered at cutover. migration/unifi-to-vyos.py turns that inventory into VyOS `set` commands for DHCP and DNS only -- the services the USG owns that VyOS must reproduce. Not a general converter. --prod and --sim come from one code path so the config proven in the sim and the config applied to the firewalls cannot drift. Prod mode hard-fails if any reservation is missing, since a silent drop is the failure mode that matters. DNS is included because the USG resolves for 5 of 6 VLANs today: UniFi hands out the gateway's own address whenever dhcpd_dns is empty, verified by labmaster resolving against 192.168.8.1. Replacing the USG without a forwarder would take DNS away from those VLANs entirely. labsim/labsim-dhcp-test.sh proves it by booting throwaway VMs with real production MACs -- the one piece of production config that transplants verbatim. Safe because ovs-labsim has no physical NIC, so those MACs cannot reach the real LAN. Result on VyOS 2026.08 (kea), 4/4: printer1 got 172.31.10.46 from inside the pool, sonoff-matter got 172.31.11.67 across the /23 boundary, Hubitat got its out-of-pool .2, and an unreserved MAC got an unreserved address. kea honours in-pool host reservations -- the open question blocking the cutover. Supporting changes to labsim: - VLAN 10 widened to /23. Every reservation is in LoT and LoT spans 10.0.0.x and 10.0.1.x, which a /24 cannot represent. - LoT's host leg moved to .3, because 10.0.0.2 is a real reservation (Hubitat) that maps onto the host's own address. - vlans.conf gained optional masklen and host_octet fields, defaulting to 24 and 2 so the other five VLANs are untouched. - Fixed /etc/network/interfaces hardcoding 255.255.255.0. That file is what actually takes effect on these Alpine guests -- cloud-init's network-config is ignored -- so any non-/24 VLAN was silently wrong. Two generator bugs found by VyOS rejecting the output: static-mapping names are validated as hostnames, so underscores fail; and two devices named "espressif" plus two named "thebeast" collided into single names, which would have overwritten one reservation with another's address. The raw export holds WiFi passphrases and the WAN PPPoE credentials and is gitignored. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01DMVzWZgiKW2wquf5z8S1yH
2026-08-15 01:33:00 +01:00
## Testing the DHCP migration
`./labsim-dhcp-test.sh` boots throwaway VMs whose MACs are **real production
MACs** and checks each gets the address UniFi reserved for it. MACs are the one
piece of production config that transplants verbatim, which is what makes this a
test rather than a rehearsal. It is safe because `ovs-labsim` has no physical
NIC — verified with `ovs-vsctl show` — so a production MAC cannot reach the real
LAN.
Apply the config first, from `../migration`:
```bash
python3 unifi-to-vyos.py --mode sim -o /tmp/sim.conf # 6 subnets, 31 mappings
# load onto labsim-vyos, then:
./labsim-dhcp-test.sh
```
**Result on VyOS 2026.08 (kea): all four cases pass.** The one that mattered:
feat(migration): reserve every active client at its current address kea does not inherit UniFi's lease database. At cutover it starts with an empty view of who holds what, so it can hand an address that is currently in use to a different device. Reservations are what carry "this device has this address" across the switch, because they live in config rather than in lease state. unifi-reserve-all.py creates one per active client, dry run by default. 51 written, 51/51 verified live by reading the records back; the controller now holds 85 reservations and the generator emits all 85 with unique, valid hostnames and no duplicate addresses. Active clients with no reservation went from 48 to 4. Three guards, each of which caught something real in the dry run: - VRRP virtual addresses are excluded. UniFi reports them as ordinary client addresses because the firewalls' bond MACs answer for them, and their apparent IP flips between the real interface address and the VIP. Without this, 192.168.1.254 -- the gateway VIP itself -- would have been given a DHCP reservation. - The firewalls' own interface MACs are excluded; those are statically configured routers, not DHCP clients. - Any address claimed by more than one MAC is dropped rather than guessed at. This is how the VIPs surfaced in the first place. Also skipped: addresses already reserved to another MAC, network gateways, and anything on a network that does not serve DHCP (which excludes the WAN transit VLANs automatically). labsim-dhcp-test.sh gained a lease-database flush, and it is not tidiness. Two findings, both of which first appeared as a PASSING test: - Re-running against stale leases, kea gave dynamic addresses to three devices that have reservations. The reservations were present and correct in kea's own config throughout. Kea saw the reserved address as leased to "another client" -- same MAC, different client-id from the earlier boot -- and allocated elsewhere. Cutover starts with an empty lease database so this is a testing artifact, but a reservation is evidently not unconditional once leases exist. - Removing only dhcp4-leases.csv does nothing: kea's memfile backend keeps lease-file-cleanup rotations (.csv.2) and restores from them on start. The verdict logic no longer takes the first matching lease row. Doing so reported an hours-old lease as the current answer and scored three failures as passes, including one where the device had plainly been given a dynamic address. A MAC with more than one lease is now an explicit failure rather than a guess. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01DMVzWZgiKW2wquf5z8S1yH
2026-08-15 23:07:07 +01:00
most UniFi reservations sit *inside* the DHCP pool, and **kea honours in-pool
host reservations** — `printer1` received `172.31.10.46` from within the
feat(migration): export UniFi config and generate VyOS DHCP+DNS from it Groundwork for replacing the USG with the VyOS pair without anything on the network noticing. Three pieces: migration/unifi-export.py pulls 13 endpoints off the classic controller into timestamped JSON plus a normalised inventory: 11 networks, 31 DHCP reservations, 4 port forwards, 2 firewall rules, 0 static routes. Two things this turned up that a naive export would have lost: - 23 of the 31 reservations carry no network_id at all -- UniFi simply does not store the binding -- so they are resolved by subnet containment instead. Without that, three quarters of the reservations have no subnet to be placed in. - 30 of the 31 sit INSIDE the DHCP pool, which UniFi's dhcpd tolerates and which is flagged as a warning rather than discovered at cutover. migration/unifi-to-vyos.py turns that inventory into VyOS `set` commands for DHCP and DNS only -- the services the USG owns that VyOS must reproduce. Not a general converter. --prod and --sim come from one code path so the config proven in the sim and the config applied to the firewalls cannot drift. Prod mode hard-fails if any reservation is missing, since a silent drop is the failure mode that matters. DNS is included because the USG resolves for 5 of 6 VLANs today: UniFi hands out the gateway's own address whenever dhcpd_dns is empty, verified by labmaster resolving against 192.168.8.1. Replacing the USG without a forwarder would take DNS away from those VLANs entirely. labsim/labsim-dhcp-test.sh proves it by booting throwaway VMs with real production MACs -- the one piece of production config that transplants verbatim. Safe because ovs-labsim has no physical NIC, so those MACs cannot reach the real LAN. Result on VyOS 2026.08 (kea), 4/4: printer1 got 172.31.10.46 from inside the pool, sonoff-matter got 172.31.11.67 across the /23 boundary, Hubitat got its out-of-pool .2, and an unreserved MAC got an unreserved address. kea honours in-pool host reservations -- the open question blocking the cutover. Supporting changes to labsim: - VLAN 10 widened to /23. Every reservation is in LoT and LoT spans 10.0.0.x and 10.0.1.x, which a /24 cannot represent. - LoT's host leg moved to .3, because 10.0.0.2 is a real reservation (Hubitat) that maps onto the host's own address. - vlans.conf gained optional masklen and host_octet fields, defaulting to 24 and 2 so the other five VLANs are untouched. - Fixed /etc/network/interfaces hardcoding 255.255.255.0. That file is what actually takes effect on these Alpine guests -- cloud-init's network-config is ignored -- so any non-/24 VLAN was silently wrong. Two generator bugs found by VyOS rejecting the output: static-mapping names are validated as hostnames, so underscores fail; and two devices named "espressif" plus two named "thebeast" collided into single names, which would have overwritten one reservation with another's address. The raw export holds WiFi passphrases and the WAN PPPoE credentials and is gitignored. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01DMVzWZgiKW2wquf5z8S1yH
2026-08-15 01:33:00 +01:00
`.10.11.11.254` pool. That was the open question blocking the cutover.
feat(migration): reserve every active client at its current address kea does not inherit UniFi's lease database. At cutover it starts with an empty view of who holds what, so it can hand an address that is currently in use to a different device. Reservations are what carry "this device has this address" across the switch, because they live in config rather than in lease state. unifi-reserve-all.py creates one per active client, dry run by default. 51 written, 51/51 verified live by reading the records back; the controller now holds 85 reservations and the generator emits all 85 with unique, valid hostnames and no duplicate addresses. Active clients with no reservation went from 48 to 4. Three guards, each of which caught something real in the dry run: - VRRP virtual addresses are excluded. UniFi reports them as ordinary client addresses because the firewalls' bond MACs answer for them, and their apparent IP flips between the real interface address and the VIP. Without this, 192.168.1.254 -- the gateway VIP itself -- would have been given a DHCP reservation. - The firewalls' own interface MACs are excluded; those are statically configured routers, not DHCP clients. - Any address claimed by more than one MAC is dropped rather than guessed at. This is how the VIPs surfaced in the first place. Also skipped: addresses already reserved to another MAC, network gateways, and anything on a network that does not serve DHCP (which excludes the WAN transit VLANs automatically). labsim-dhcp-test.sh gained a lease-database flush, and it is not tidiness. Two findings, both of which first appeared as a PASSING test: - Re-running against stale leases, kea gave dynamic addresses to three devices that have reservations. The reservations were present and correct in kea's own config throughout. Kea saw the reserved address as leased to "another client" -- same MAC, different client-id from the earlier boot -- and allocated elsewhere. Cutover starts with an empty lease database so this is a testing artifact, but a reservation is evidently not unconditional once leases exist. - Removing only dhcp4-leases.csv does nothing: kea's memfile backend keeps lease-file-cleanup rotations (.csv.2) and restores from them on start. The verdict logic no longer takes the first matching lease row. Doing so reported an hours-old lease as the current answer and scored three failures as passes, including one where the device had plainly been given a dynamic address. A MAC with more than one lease is now an explicit failure rather than a guess. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01DMVzWZgiKW2wquf5z8S1yH
2026-08-15 23:07:07 +01:00
### The lease database will lie to you
The script wipes `/config/dhcp/dhcp4-leases.csv*` before every run, and both
halves of that matter:
- **Stale leases defeat reservations.** Re-running against yesterday's leases,
kea handed dynamic addresses to three devices that have reservations. The
reservation was present and correct in `/run/kea/kea-dhcp4.conf` the whole
time. Kea saw the reserved address as already leased to "another client" —
same MAC, but a different client-id from the earlier boot — and allocated
elsewhere. The cutover itself starts with an empty lease database, so this is
a *testing* artifact, but it is worth knowing that a reservation is not an
unconditional guarantee once leases exist.
- **The `*` is load-bearing.** Kea's memfile backend keeps lease-file-cleanup
rotations (`dhcp4-leases.csv.2`) and restores from them on start, so
truncating only the primary file changes nothing.
Both of those first appeared as a *passing* test. The verdict logic now refuses
to score a MAC with more than one lease, because taking the first match had
reported an hours-old lease as the current answer and turned three failures
into apparent passes.
feat(migration): export UniFi config and generate VyOS DHCP+DNS from it Groundwork for replacing the USG with the VyOS pair without anything on the network noticing. Three pieces: migration/unifi-export.py pulls 13 endpoints off the classic controller into timestamped JSON plus a normalised inventory: 11 networks, 31 DHCP reservations, 4 port forwards, 2 firewall rules, 0 static routes. Two things this turned up that a naive export would have lost: - 23 of the 31 reservations carry no network_id at all -- UniFi simply does not store the binding -- so they are resolved by subnet containment instead. Without that, three quarters of the reservations have no subnet to be placed in. - 30 of the 31 sit INSIDE the DHCP pool, which UniFi's dhcpd tolerates and which is flagged as a warning rather than discovered at cutover. migration/unifi-to-vyos.py turns that inventory into VyOS `set` commands for DHCP and DNS only -- the services the USG owns that VyOS must reproduce. Not a general converter. --prod and --sim come from one code path so the config proven in the sim and the config applied to the firewalls cannot drift. Prod mode hard-fails if any reservation is missing, since a silent drop is the failure mode that matters. DNS is included because the USG resolves for 5 of 6 VLANs today: UniFi hands out the gateway's own address whenever dhcpd_dns is empty, verified by labmaster resolving against 192.168.8.1. Replacing the USG without a forwarder would take DNS away from those VLANs entirely. labsim/labsim-dhcp-test.sh proves it by booting throwaway VMs with real production MACs -- the one piece of production config that transplants verbatim. Safe because ovs-labsim has no physical NIC, so those MACs cannot reach the real LAN. Result on VyOS 2026.08 (kea), 4/4: printer1 got 172.31.10.46 from inside the pool, sonoff-matter got 172.31.11.67 across the /23 boundary, Hubitat got its out-of-pool .2, and an unreserved MAC got an unreserved address. kea honours in-pool host reservations -- the open question blocking the cutover. Supporting changes to labsim: - VLAN 10 widened to /23. Every reservation is in LoT and LoT spans 10.0.0.x and 10.0.1.x, which a /24 cannot represent. - LoT's host leg moved to .3, because 10.0.0.2 is a real reservation (Hubitat) that maps onto the host's own address. - vlans.conf gained optional masklen and host_octet fields, defaulting to 24 and 2 so the other five VLANs are untouched. - Fixed /etc/network/interfaces hardcoding 255.255.255.0. That file is what actually takes effect on these Alpine guests -- cloud-init's network-config is ignored -- so any non-/24 VLAN was silently wrong. Two generator bugs found by VyOS rejecting the output: static-mapping names are validated as hostnames, so underscores fail; and two devices named "espressif" plus two named "thebeast" collided into single names, which would have overwritten one reservation with another's address. The raw export holds WiFi passphrases and the WAN PPPoE credentials and is gitignored. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01DMVzWZgiKW2wquf5z8S1yH
2026-08-15 01:33:00 +01:00
Still open: whether kea will hand a *reserved* address to a *different* client
while the reserved device is offline. The negative case here only proves an
unreserved MAC gets an unreserved address.
- **http://localhost:9101/** — live mesh: a node per VLAN, the router in the
middle, one line per pair coloured green/red with the ICMP RTT on it. Hover a
line for per-direction detail. Refreshes every 5s. This is the one to watch
while changing firewall rules.
- **http://localhost:3000/d/labsim-matrix** — Grafana (anonymous, no login) for
*history*: when did a path flip, and how has latency moved.
- **http://localhost:9101/metrics** — `labsim_reachable{src,dst,proto}` and
`labsim_rtt_ms{src,dst}`.
## Routing: BGP, dual WAN, and the ISP VMs
`sim-ha-config.py` covers the LAN side of the routers. `sim-net-config.py`
covers everything that makes this a rehearsal for production *routing*:
| role | VM | what it generates |
|---|---|---|
| `primary` | `labsim-vyos` | BGP + dual WAN + health-checked failover |
| `secondary` | `labsim-vyos2` | BGP only |
| `isp-dhcp` | `labsim-isp-dhcp` | 10gig-equivalent ISP on VLAN 53 |
| `isp-pppoe` | `labsim-isp-pppoe` | Vodafone-equivalent PPPoE ISP on VLAN 51 |
Both ISP VMs are VyOS with two NICs: one on the OVS trunk facing the sim
router, one on libvirt's `default` network, NATing customers to the real
internet. They use RFC 5737 documentation ranges (`203.0.113.0/24`,
`198.51.100.0/24`) so a leaked sim route cannot blackhole anything real.
```sh
./sim-net-apply.sh check # VM state vs what the code says — run this first
./sim-net-apply.sh apply # push generated config over the serial console
```
`check` is the important one. All of this previously existed only as running
state, applied by hand over SSH; rebuilding a VM lost it, and nothing recorded
why any of it was shaped the way it was.
### Known gaps vs production
- **WAN is on the primary router only.** Production has WAN on both. Two PPPoE
clients sharing one credential against a single access concentrator is a
failure mode production does not have, so the sim does not model it. VRRP and
conntrack failover are still exercised.
- **ISP VM interface names are not stable across a rebuild** — `isp-dhcp` came
up as `eth0`/`eth1` and `isp-pppoe` as `eth2`/`eth3` from identical XML.
Check `show interfaces` and pass `--wan-if` / `--uplink-if` rather than
trusting the defaults.
- **`eth2` on the primary router** is a libvirt-NAT uplink predating the ISP
VMs: a third default route with no production equivalent that masks real WAN
failures during a failover test. `--drop-scaffold` removes it.
- **Committing on `isp-pppoe` drops the router's PPPoE session**, and the
client does not redial promptly. After any change there, check `pppoe0` on
the router and `sudo systemctl restart ppp@pppoe0` if it is missing.
feat(labsim): libvirt replica of the lab network with LACP + VyOS routing A throwaway copy of the production VLAN topology so routing and firewall changes can be tested before they touch the real network. Same VLAN IDs and roles as UniFi, deliberately different ranges (172.31.<vlan>.0/24) so nothing here can be mistaken for production. - OVS fabric: real 802.1Q. Access port per micro VM, host leg per VLAN (.2, for SSH only — NOT the VMs' default route, so inter-VLAN tests exercise the router rather than the host's routing table), and a trunk portgroup with VLAN 1 declared nativeMode='untagged'. - Six Alpine micro VMs (256MB, copy-on-write overlays on one 176MB image), SSH + a hello-world HTTP page naming the VLAN. - VyOS router installed to disk unattended over the console, with the SAME config shape as the VP2440s: two NICs in an LACP bond carrying the trunk, VLAN 1 native, bond0.<vlan> holding the .1 gateway on each. - labsim-matrix.py: full-mesh ICMP/TCP22/TCP80 probe, ~0.2s, --watch highlights cells that changed since the last sweep. Guest-side probe is python3 (already present via cloud-init) so nothing is installed on VMs that have no internet. - Prometheus + Grafana (anonymous auth, no login) with a provisioned dashboard: heatmap plus a state timeline showing exactly when a path flipped. Verified end to end: one VyOS rule took sum(labsim_reachable) from 90 to 84, blocking precisely kvm<->k8s across all three protocols. Traps found building this, all now encoded in the scripts: - virtio-net breaks 802.3ad: the guest's bonding driver reports slaves "MII Status: down" despite carrier=1 and never sends an LACPDU, so the bond sits in AD_STATE_DEFAULTED. e1000e fixes it with no other change. Matches the netdev thread "bonding (IEEE 802.3ad) not working with qemu/virtio". - OVS defaults bonds to active-backup, which does not speak LACP at all — bond_mode=balance-tcp is required. - LACP deadlock: OVS holds members disabled until negotiation while the partner needs carrier before it will send LACPDUs. lacp-fallback-ab breaks it. - LACPDUs are untagged, so a trunk with no native VLAN has nowhere to put them. - --boot cdrom,hd re-runs the ISO on every restart, so every commit+save went to a live system that evaporated. Install now switches the VM to boot hd. - cloud-init on Alpine: users stay locked without lock_passwd:false, one failing runcmd aborts the rest, busybox here has no httpd applet, and start-stop-daemon --exec /usr/bin/python3 matches cloud-init's own python3. - The user-data heredoc is unquoted, so backticks in a COMMENT were executed by the host shell and their output corrupted the YAML. build_seed now validates with yaml.safe_load before building the ISO. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01DMVzWZgiKW2wquf5z8S1yH
2026-08-13 00:42:39 +01:00
## Notes for whoever extends this
Things that cost time the first time round, all verified on this image:
- **No `sudo`.** Alpine ships `doas`; cloud-init's `sudo:` directive is inert
here. Use `root@` for privileged work.
- **cloud-init leaves users locked** (`!*` in `/etc/shadow`) unless
`lock_passwd: false`, and sshd then refuses key auth for that user.
- **One failing `runcmd` aborts every command after it.** Each entry is
`|| true` for that reason.
- **busybox here has no `httpd` applet**, and the VMs have no internet to
`apk add` one — so the hello-world server is `python3 -m http.server`
(python3 is already present because cloud-init depends on it).
- **`start-stop-daemon --exec /usr/bin/python3` matches cloud-init's own
python3** at boot and refuses to start anything.
- **busybox `pgrep -f PATTERN` matches its own argv**, so a "skip if already
running" guard always fires. Verified: `guard_exit=0` with nothing listening.
## Not modelled (yet)
VLANs are separate L2 segments rather than one 802.1Q trunk, so this exercises
inter-VLAN routing but not a `bond0.<vif>` trunk config specifically. A router
VM would attach one NIC per VLAN. Adding a tagged-trunk variant is the obvious
next step if the bond/vif config itself needs testing.