Some checks failed
First run of the 3-server harness: the config path works (VMs boot, config.yaml rendered by the production generator, k3s installs, etcd forms), but three full control-plane servers starved etcd of CPU on a labsim host running 18 VMs (load 8+). etcd lost leader continuously -> "storage is (re)initializing" -> k3s stuck activating -> datastore reset. Not a harness bug; host capacity + etcd's latency intolerance. Fixes for a clean next run: - LAB-ONLY etcd tuning (heartbeat-interval=500, election-timeout=5000) in the install exec, clearly marked as not-production and not affecting what the conversion test exercises -- the standard remedy for etcd under nested-virt. - Shut the other labsim k8s/dualstack VMs before `up` (done at teardown). Also fixed a real sim drift found en route: the sim MASTER router lacked the NAT masquerade rules the backup had, so LAN had no internet whenever .253 was master. Added masquerade for 172.31.0.0/16 out bond0.53 + pppoe0. Full write-up: labsim/dualstack-evidence/etcd-harness-2026-09-07.md. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01DMVzWZgiKW2wquf5z8S1yH
10 KiB
Executable File
10 KiB
Executable File