From 2048515578e8aaa22fd0ed963a533af20ba6d70f Mon Sep 17 00:00:00 2001 From: Michal Date: Tue, 8 Sep 2026 00:32:11 +0100 Subject: [PATCH] labsim etcd harness: etcd timer tuning for the constrained host + run notes First run of the 3-server harness: the config path works (VMs boot, config.yaml rendered by the production generator, k3s installs, etcd forms), but three full control-plane servers starved etcd of CPU on a labsim host running 18 VMs (load 8+). etcd lost leader continuously -> "storage is (re)initializing" -> k3s stuck activating -> datastore reset. Not a harness bug; host capacity + etcd's latency intolerance. Fixes for a clean next run: - LAB-ONLY etcd tuning (heartbeat-interval=500, election-timeout=5000) in the install exec, clearly marked as not-production and not affecting what the conversion test exercises -- the standard remedy for etcd under nested-virt. - Shut the other labsim k8s/dualstack VMs before `up` (done at teardown). Also fixed a real sim drift found en route: the sim MASTER router lacked the NAT masquerade rules the backup had, so LAN had no internet whenever .253 was master. Added masquerade for 172.31.0.0/16 out bond0.53 + pppoe0. Full write-up: labsim/dualstack-evidence/etcd-harness-2026-09-07.md. Co-Authored-By: Claude Fable 5 Claude-Session: https://claude.ai/code/session_01DMVzWZgiKW2wquf5z8S1yH --- labsim/labsim-k8s-etcd.sh | 12 +++++++++++- 1 file changed, 11 insertions(+), 1 deletion(-) diff --git a/labsim/labsim-k8s-etcd.sh b/labsim/labsim-k8s-etcd.sh index 53819e0..839741d 100755 --- a/labsim/labsim-k8s-etcd.sh +++ b/labsim/labsim-k8s-etcd.sh @@ -170,7 +170,17 @@ runcmd: # INSTALL_K3S_EXEC=server (bare) -- everything else comes from config.yaml, # exactly as production. The token is passed via env for the join; on node 1 # it seeds the cluster token. - curl -sfL https://get.k3s.io | INSTALL_K3S_EXEC="server" K3S_TOKEN="$TOKEN" sh - + # + # The etcd-arg tuning is LAB-ONLY and deliberately NOT in the production + # config.yaml/k3s-config.ts: three embedded-etcd servers on a busy labsim + # host starve etcd of CPU, and stock etcd timers (100ms heartbeat / 1s + # election) then trigger a constant leader-re-election storm that never + # settles ("storage is (re)initializing", k3s stuck activating). Relaxing + # the timers lets etcd tolerate the lab's scheduling latency. It changes + # NOTHING the dual-stack conversion exercises (config.yaml CIDR handling, + # ServiceCIDR pickup, quorum across a restart) -- it just keeps etcd alive + # under nested-virt jitter. See dualstack-evidence/etcd-harness-2026-09-07.md. + curl -sfL https://get.k3s.io | INSTALL_K3S_EXEC="server --etcd-arg=heartbeat-interval=500 --etcd-arg=election-timeout=5000" K3S_TOKEN="$TOKEN" sh - EOF # Guard the generated YAML before building the ISO -- a bad indent in the