Files
lab/labsim/vlans.conf

44 lines
2.0 KiB
Plaintext
Raw Permalink Normal View History

feat(labsim): libvirt replica of the lab network with LACP + VyOS routing A throwaway copy of the production VLAN topology so routing and firewall changes can be tested before they touch the real network. Same VLAN IDs and roles as UniFi, deliberately different ranges (172.31.<vlan>.0/24) so nothing here can be mistaken for production. - OVS fabric: real 802.1Q. Access port per micro VM, host leg per VLAN (.2, for SSH only — NOT the VMs' default route, so inter-VLAN tests exercise the router rather than the host's routing table), and a trunk portgroup with VLAN 1 declared nativeMode='untagged'. - Six Alpine micro VMs (256MB, copy-on-write overlays on one 176MB image), SSH + a hello-world HTTP page naming the VLAN. - VyOS router installed to disk unattended over the console, with the SAME config shape as the VP2440s: two NICs in an LACP bond carrying the trunk, VLAN 1 native, bond0.<vlan> holding the .1 gateway on each. - labsim-matrix.py: full-mesh ICMP/TCP22/TCP80 probe, ~0.2s, --watch highlights cells that changed since the last sweep. Guest-side probe is python3 (already present via cloud-init) so nothing is installed on VMs that have no internet. - Prometheus + Grafana (anonymous auth, no login) with a provisioned dashboard: heatmap plus a state timeline showing exactly when a path flipped. Verified end to end: one VyOS rule took sum(labsim_reachable) from 90 to 84, blocking precisely kvm<->k8s across all three protocols. Traps found building this, all now encoded in the scripts: - virtio-net breaks 802.3ad: the guest's bonding driver reports slaves "MII Status: down" despite carrier=1 and never sends an LACPDU, so the bond sits in AD_STATE_DEFAULTED. e1000e fixes it with no other change. Matches the netdev thread "bonding (IEEE 802.3ad) not working with qemu/virtio". - OVS defaults bonds to active-backup, which does not speak LACP at all — bond_mode=balance-tcp is required. - LACP deadlock: OVS holds members disabled until negotiation while the partner needs carrier before it will send LACPDUs. lacp-fallback-ab breaks it. - LACPDUs are untagged, so a trunk with no native VLAN has nowhere to put them. - --boot cdrom,hd re-runs the ISO on every restart, so every commit+save went to a live system that evaporated. Install now switches the VM to boot hd. - cloud-init on Alpine: users stay locked without lock_passwd:false, one failing runcmd aborts the rest, busybox here has no httpd applet, and start-stop-daemon --exec /usr/bin/python3 matches cloud-init's own python3. - The user-data heredoc is unquoted, so backticks in a COMMENT were executed by the host shell and their output corrupted the YAML. build_seed now validates with yaml.safe_load before building the ISO. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01DMVzWZgiKW2wquf5z8S1yH
2026-08-13 00:42:39 +01:00
# Lab network simulation — VLAN map.
#
# Mirrors the real UniFi topology (same VLAN IDs, same roles) but with
# deliberately DIFFERENT IP ranges so nothing here can collide with, or be
# confused for, production. The sim subnet always encodes the VLAN id:
#
# 172.31.<vlan-id>.0/24
#
# Per-subnet address plan (same shape on every VLAN):
# .1 gateway under test (VyOS/router VM — not created by default)
# .2 host bridge (how you SSH in from this workstation)
# .10 the micro VM for this VLAN
# .254 VRRP VIP (reserved, mirrors production)
#
feat(migration): export UniFi config and generate VyOS DHCP+DNS from it Groundwork for replacing the USG with the VyOS pair without anything on the network noticing. Three pieces: migration/unifi-export.py pulls 13 endpoints off the classic controller into timestamped JSON plus a normalised inventory: 11 networks, 31 DHCP reservations, 4 port forwards, 2 firewall rules, 0 static routes. Two things this turned up that a naive export would have lost: - 23 of the 31 reservations carry no network_id at all -- UniFi simply does not store the binding -- so they are resolved by subnet containment instead. Without that, three quarters of the reservations have no subnet to be placed in. - 30 of the 31 sit INSIDE the DHCP pool, which UniFi's dhcpd tolerates and which is flagged as a warning rather than discovered at cutover. migration/unifi-to-vyos.py turns that inventory into VyOS `set` commands for DHCP and DNS only -- the services the USG owns that VyOS must reproduce. Not a general converter. --prod and --sim come from one code path so the config proven in the sim and the config applied to the firewalls cannot drift. Prod mode hard-fails if any reservation is missing, since a silent drop is the failure mode that matters. DNS is included because the USG resolves for 5 of 6 VLANs today: UniFi hands out the gateway's own address whenever dhcpd_dns is empty, verified by labmaster resolving against 192.168.8.1. Replacing the USG without a forwarder would take DNS away from those VLANs entirely. labsim/labsim-dhcp-test.sh proves it by booting throwaway VMs with real production MACs -- the one piece of production config that transplants verbatim. Safe because ovs-labsim has no physical NIC, so those MACs cannot reach the real LAN. Result on VyOS 2026.08 (kea), 4/4: printer1 got 172.31.10.46 from inside the pool, sonoff-matter got 172.31.11.67 across the /23 boundary, Hubitat got its out-of-pool .2, and an unreserved MAC got an unreserved address. kea honours in-pool host reservations -- the open question blocking the cutover. Supporting changes to labsim: - VLAN 10 widened to /23. Every reservation is in LoT and LoT spans 10.0.0.x and 10.0.1.x, which a /24 cannot represent. - LoT's host leg moved to .3, because 10.0.0.2 is a real reservation (Hubitat) that maps onto the host's own address. - vlans.conf gained optional masklen and host_octet fields, defaulting to 24 and 2 so the other five VLANs are untouched. - Fixed /etc/network/interfaces hardcoding 255.255.255.0. That file is what actually takes effect on these Alpine guests -- cloud-init's network-config is ignored -- so any non-/24 VLAN was silently wrong. Two generator bugs found by VyOS rejecting the output: static-mapping names are validated as hostnames, so underscores fail; and two devices named "espressif" plus two named "thebeast" collided into single names, which would have overwritten one reservation with another's address. The raw export holds WiFi passphrases and the WAN PPPoE credentials and is gitignored. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01DMVzWZgiKW2wquf5z8S1yH
2026-08-15 01:33:00 +01:00
# Format: vlan_id:name:sim_subnet_prefix:real_subnet:[masklen]:[host_octet]
#
# masklen defaults to 24 and host_octet to 2. Both exist for VLAN 10, which is
# the one VLAN that has to be a /23 here: every UniFi DHCP reservation lives in
# LoT, and LoT spans 10.0.0.x AND 10.0.1.x, which a /24 cannot represent. With
# /23 the mapping stays readable — 10.0.0.46 -> 172.31.10.46 and
# 10.0.1.67 -> 172.31.11.67.
#
# LoT's host leg is .3 rather than .2 because 10.0.0.2 is a real reservation
# (Hubitat) and would map straight onto the host's own address. .3 is free in
# production and sits below the DHCP pool (which starts at .11), so it can
# never be handed out.
feat(labsim): libvirt replica of the lab network with LACP + VyOS routing A throwaway copy of the production VLAN topology so routing and firewall changes can be tested before they touch the real network. Same VLAN IDs and roles as UniFi, deliberately different ranges (172.31.<vlan>.0/24) so nothing here can be mistaken for production. - OVS fabric: real 802.1Q. Access port per micro VM, host leg per VLAN (.2, for SSH only — NOT the VMs' default route, so inter-VLAN tests exercise the router rather than the host's routing table), and a trunk portgroup with VLAN 1 declared nativeMode='untagged'. - Six Alpine micro VMs (256MB, copy-on-write overlays on one 176MB image), SSH + a hello-world HTTP page naming the VLAN. - VyOS router installed to disk unattended over the console, with the SAME config shape as the VP2440s: two NICs in an LACP bond carrying the trunk, VLAN 1 native, bond0.<vlan> holding the .1 gateway on each. - labsim-matrix.py: full-mesh ICMP/TCP22/TCP80 probe, ~0.2s, --watch highlights cells that changed since the last sweep. Guest-side probe is python3 (already present via cloud-init) so nothing is installed on VMs that have no internet. - Prometheus + Grafana (anonymous auth, no login) with a provisioned dashboard: heatmap plus a state timeline showing exactly when a path flipped. Verified end to end: one VyOS rule took sum(labsim_reachable) from 90 to 84, blocking precisely kvm<->k8s across all three protocols. Traps found building this, all now encoded in the scripts: - virtio-net breaks 802.3ad: the guest's bonding driver reports slaves "MII Status: down" despite carrier=1 and never sends an LACPDU, so the bond sits in AD_STATE_DEFAULTED. e1000e fixes it with no other change. Matches the netdev thread "bonding (IEEE 802.3ad) not working with qemu/virtio". - OVS defaults bonds to active-backup, which does not speak LACP at all — bond_mode=balance-tcp is required. - LACP deadlock: OVS holds members disabled until negotiation while the partner needs carrier before it will send LACPDUs. lacp-fallback-ab breaks it. - LACPDUs are untagged, so a trunk with no native VLAN has nowhere to put them. - --boot cdrom,hd re-runs the ISO on every restart, so every commit+save went to a live system that evaporated. Install now switches the VM to boot hd. - cloud-init on Alpine: users stay locked without lock_passwd:false, one failing runcmd aborts the rest, busybox here has no httpd applet, and start-stop-daemon --exec /usr/bin/python3 matches cloud-init's own python3. - The user-data heredoc is unquoted, so backticks in a COMMENT were executed by the host shell and their output corrupted the YAML. build_seed now validates with yaml.safe_load before building the ISO. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01DMVzWZgiKW2wquf5z8S1yH
2026-08-13 00:42:39 +01:00
1:management:172.31.1:192.168.1.0/24
2:k8s:172.31.2:192.168.8.0/23
3:kvm:172.31.3:192.168.3.0/24
9:private:172.31.9:10.0.9.0/23
feat(migration): export UniFi config and generate VyOS DHCP+DNS from it Groundwork for replacing the USG with the VyOS pair without anything on the network noticing. Three pieces: migration/unifi-export.py pulls 13 endpoints off the classic controller into timestamped JSON plus a normalised inventory: 11 networks, 31 DHCP reservations, 4 port forwards, 2 firewall rules, 0 static routes. Two things this turned up that a naive export would have lost: - 23 of the 31 reservations carry no network_id at all -- UniFi simply does not store the binding -- so they are resolved by subnet containment instead. Without that, three quarters of the reservations have no subnet to be placed in. - 30 of the 31 sit INSIDE the DHCP pool, which UniFi's dhcpd tolerates and which is flagged as a warning rather than discovered at cutover. migration/unifi-to-vyos.py turns that inventory into VyOS `set` commands for DHCP and DNS only -- the services the USG owns that VyOS must reproduce. Not a general converter. --prod and --sim come from one code path so the config proven in the sim and the config applied to the firewalls cannot drift. Prod mode hard-fails if any reservation is missing, since a silent drop is the failure mode that matters. DNS is included because the USG resolves for 5 of 6 VLANs today: UniFi hands out the gateway's own address whenever dhcpd_dns is empty, verified by labmaster resolving against 192.168.8.1. Replacing the USG without a forwarder would take DNS away from those VLANs entirely. labsim/labsim-dhcp-test.sh proves it by booting throwaway VMs with real production MACs -- the one piece of production config that transplants verbatim. Safe because ovs-labsim has no physical NIC, so those MACs cannot reach the real LAN. Result on VyOS 2026.08 (kea), 4/4: printer1 got 172.31.10.46 from inside the pool, sonoff-matter got 172.31.11.67 across the /23 boundary, Hubitat got its out-of-pool .2, and an unreserved MAC got an unreserved address. kea honours in-pool host reservations -- the open question blocking the cutover. Supporting changes to labsim: - VLAN 10 widened to /23. Every reservation is in LoT and LoT spans 10.0.0.x and 10.0.1.x, which a /24 cannot represent. - LoT's host leg moved to .3, because 10.0.0.2 is a real reservation (Hubitat) that maps onto the host's own address. - vlans.conf gained optional masklen and host_octet fields, defaulting to 24 and 2 so the other five VLANs are untouched. - Fixed /etc/network/interfaces hardcoding 255.255.255.0. That file is what actually takes effect on these Alpine guests -- cloud-init's network-config is ignored -- so any non-/24 VLAN was silently wrong. Two generator bugs found by VyOS rejecting the output: static-mapping names are validated as hostnames, so underscores fail; and two devices named "espressif" plus two named "thebeast" collided into single names, which would have overwritten one reservation with another's address. The raw export holds WiFi passphrases and the WAN PPPoE credentials and is gitignored. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01DMVzWZgiKW2wquf5z8S1yH
2026-08-15 01:33:00 +01:00
10:lot:172.31.10:10.0.0.0/23:23:3
feat(labsim): libvirt replica of the lab network with LACP + VyOS routing A throwaway copy of the production VLAN topology so routing and firewall changes can be tested before they touch the real network. Same VLAN IDs and roles as UniFi, deliberately different ranges (172.31.<vlan>.0/24) so nothing here can be mistaken for production. - OVS fabric: real 802.1Q. Access port per micro VM, host leg per VLAN (.2, for SSH only — NOT the VMs' default route, so inter-VLAN tests exercise the router rather than the host's routing table), and a trunk portgroup with VLAN 1 declared nativeMode='untagged'. - Six Alpine micro VMs (256MB, copy-on-write overlays on one 176MB image), SSH + a hello-world HTTP page naming the VLAN. - VyOS router installed to disk unattended over the console, with the SAME config shape as the VP2440s: two NICs in an LACP bond carrying the trunk, VLAN 1 native, bond0.<vlan> holding the .1 gateway on each. - labsim-matrix.py: full-mesh ICMP/TCP22/TCP80 probe, ~0.2s, --watch highlights cells that changed since the last sweep. Guest-side probe is python3 (already present via cloud-init) so nothing is installed on VMs that have no internet. - Prometheus + Grafana (anonymous auth, no login) with a provisioned dashboard: heatmap plus a state timeline showing exactly when a path flipped. Verified end to end: one VyOS rule took sum(labsim_reachable) from 90 to 84, blocking precisely kvm<->k8s across all three protocols. Traps found building this, all now encoded in the scripts: - virtio-net breaks 802.3ad: the guest's bonding driver reports slaves "MII Status: down" despite carrier=1 and never sends an LACPDU, so the bond sits in AD_STATE_DEFAULTED. e1000e fixes it with no other change. Matches the netdev thread "bonding (IEEE 802.3ad) not working with qemu/virtio". - OVS defaults bonds to active-backup, which does not speak LACP at all — bond_mode=balance-tcp is required. - LACP deadlock: OVS holds members disabled until negotiation while the partner needs carrier before it will send LACPDUs. lacp-fallback-ab breaks it. - LACPDUs are untagged, so a trunk with no native VLAN has nowhere to put them. - --boot cdrom,hd re-runs the ISO on every restart, so every commit+save went to a live system that evaporated. Install now switches the VM to boot hd. - cloud-init on Alpine: users stay locked without lock_passwd:false, one failing runcmd aborts the rest, busybox here has no httpd applet, and start-stop-daemon --exec /usr/bin/python3 matches cloud-init's own python3. - The user-data heredoc is unquoted, so backticks in a COMMENT were executed by the host shell and their output corrupted the YAML. build_seed now validates with yaml.safe_load before building the ISO. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01DMVzWZgiKW2wquf5z8S1yH
2026-08-13 00:42:39 +01:00
200:roomates:172.31.200:192.168.2.0/24
feat(labsim): add WAN transport VLANs so the sim can host fake ISPs A cutover attempt failed on the WAN and nothing had tested it. The reason the sim could not have caught it: labsim modelled every LAN VLAN faithfully and omitted the WAN entirely -- vlans.conf had 1, 2, 3, 9, 10 and 200, never 51 or 53. Worse, the switch script's WAN health checks are conditional on the delta configuring PPPoE, so in the sim they printed "this delta configures no WAN -- skipping all WAN health checks" and passed. The sim proved the delta commits; it never proved the WAN works, and could not have. Adds VLANs 51 (Vodafone/PPPoE) and 53 (10gig/DHCP) to the fabric so a fake ISP can live on each and those checks actually execute. VyOS has service pppoe-server (accel-ppp) natively -- authentication local-users, client-ip-pool, gateway-address -- so a VyOS VM can play the concentrator, and dhcp-server can play the other ISP. vlans.conf gains host_octet 0, meaning "no host leg". A host address on a WAN transport VLAN would misrepresent the segment: the point is that VyOS reaches an ISP, not the host. Also fixes a real gap in ovs_bond_router: it returned early when the bond already existed, so adding a VLAN to vlans.conf never reached an existing bond. Re-runs now reconcile the trunk and say so. That gap is the same SHAPE as the production failure -- interface present, VLAN missing from the trunk, frames silently dropped -- which is precisely the class of bug the sim needs to be able to reproduce rather than embody. Both bonds updated: [2,3,9,10,200] -> [2,3,9,10,51,53,200]. Note on the production diagnosis, which is NOT settled: a passive RX test showed zero frames on 51/53 at the firewall, and an active DHCP DISCOVER (verified to have transmitted, tx +2) drew no reply. That is consistent with the VLANs not being trunked, but equally with the ISP only answering its registered CPE MAC -- which is exactly why the delta clones f0:9f:c2:12:9b:4f, and why it cannot be settled from production while the USG holds that MAC. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01DMVzWZgiKW2wquf5z8S1yH
2026-08-18 00:44:11 +01:00
# WAN transport VLANs, mirroring production. These exist so the sim can run a
# fake ISP on each and the switch script's WAN health checks actually execute
# instead of printing "this delta configures no WAN -- skipping". A cutover
# attempt failed on the WAN with nothing having tested it, because the sim
# modelled every LAN VLAN faithfully and omitted the WAN entirely.
#
# No host leg is wanted here (host_octet 0 means "skip"): the ISP VMs own these
# segments, and a host address on a WAN transport VLAN would be a lie.
51:wan1:172.31.51:vodafone-pppoe(VLAN 51):24:0
53:wan3:172.31.53:10gig-dhcp(VLAN 53):24:0