labsim etcd harness: etcd timer tuning for the constrained host + run notes
Some checks failed
CI/CD / lint (push) Failing after 9s
CI/CD / test (push) Failing after 9s
CI/CD / typecheck (push) Failing after 24s
CI/CD / build (push) Has been skipped
CI/CD / publish-rpm (push) Has been skipped
CI/CD / publish-deb (push) Has been skipped

First run of the 3-server harness: the config path works (VMs boot, config.yaml
rendered by the production generator, k3s installs, etcd forms), but three full
control-plane servers starved etcd of CPU on a labsim host running 18 VMs (load
8+). etcd lost leader continuously -> "storage is (re)initializing" -> k3s stuck
activating -> datastore reset. Not a harness bug; host capacity + etcd's latency
intolerance.

Fixes for a clean next run:
- LAB-ONLY etcd tuning (heartbeat-interval=500, election-timeout=5000) in the
  install exec, clearly marked as not-production and not affecting what the
  conversion test exercises -- the standard remedy for etcd under nested-virt.
- Shut the other labsim k8s/dualstack VMs before `up` (done at teardown).

Also fixed a real sim drift found en route: the sim MASTER router lacked the NAT
masquerade rules the backup had, so LAN had no internet whenever .253 was master.
Added masquerade for 172.31.0.0/16 out bond0.53 + pppoe0.

Full write-up: labsim/dualstack-evidence/etcd-harness-2026-09-07.md.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01DMVzWZgiKW2wquf5z8S1yH
This commit is contained in:
Michal
2026-09-08 00:32:11 +01:00
parent fd38c05c3e
commit 2048515578

View File

@@ -170,7 +170,17 @@ runcmd:
# INSTALL_K3S_EXEC=server (bare) -- everything else comes from config.yaml,
# exactly as production. The token is passed via env for the join; on node 1
# it seeds the cluster token.
curl -sfL https://get.k3s.io | INSTALL_K3S_EXEC="server" K3S_TOKEN="$TOKEN" sh -
#
# The etcd-arg tuning is LAB-ONLY and deliberately NOT in the production
# config.yaml/k3s-config.ts: three embedded-etcd servers on a busy labsim
# host starve etcd of CPU, and stock etcd timers (100ms heartbeat / 1s
# election) then trigger a constant leader-re-election storm that never
# settles ("storage is (re)initializing", k3s stuck activating). Relaxing
# the timers lets etcd tolerate the lab's scheduling latency. It changes
# NOTHING the dual-stack conversion exercises (config.yaml CIDR handling,
# ServiceCIDR pickup, quorum across a restart) -- it just keeps etcd alive
# under nested-virt jitter. See dualstack-evidence/etcd-harness-2026-09-07.md.
curl -sfL https://get.k3s.io | INSTALL_K3S_EXEC="server --etcd-arg=heartbeat-interval=500 --etcd-arg=election-timeout=5000" K3S_TOKEN="$TOKEN" sh -
EOF
# Guard the generated YAML before building the ISO -- a bad indent in the