Files
lab/labsim/monitoring-up.sh

82 lines
3.1 KiB
Bash
Raw Normal View History

feat(labsim): libvirt replica of the lab network with LACP + VyOS routing A throwaway copy of the production VLAN topology so routing and firewall changes can be tested before they touch the real network. Same VLAN IDs and roles as UniFi, deliberately different ranges (172.31.<vlan>.0/24) so nothing here can be mistaken for production. - OVS fabric: real 802.1Q. Access port per micro VM, host leg per VLAN (.2, for SSH only — NOT the VMs' default route, so inter-VLAN tests exercise the router rather than the host's routing table), and a trunk portgroup with VLAN 1 declared nativeMode='untagged'. - Six Alpine micro VMs (256MB, copy-on-write overlays on one 176MB image), SSH + a hello-world HTTP page naming the VLAN. - VyOS router installed to disk unattended over the console, with the SAME config shape as the VP2440s: two NICs in an LACP bond carrying the trunk, VLAN 1 native, bond0.<vlan> holding the .1 gateway on each. - labsim-matrix.py: full-mesh ICMP/TCP22/TCP80 probe, ~0.2s, --watch highlights cells that changed since the last sweep. Guest-side probe is python3 (already present via cloud-init) so nothing is installed on VMs that have no internet. - Prometheus + Grafana (anonymous auth, no login) with a provisioned dashboard: heatmap plus a state timeline showing exactly when a path flipped. Verified end to end: one VyOS rule took sum(labsim_reachable) from 90 to 84, blocking precisely kvm<->k8s across all three protocols. Traps found building this, all now encoded in the scripts: - virtio-net breaks 802.3ad: the guest's bonding driver reports slaves "MII Status: down" despite carrier=1 and never sends an LACPDU, so the bond sits in AD_STATE_DEFAULTED. e1000e fixes it with no other change. Matches the netdev thread "bonding (IEEE 802.3ad) not working with qemu/virtio". - OVS defaults bonds to active-backup, which does not speak LACP at all — bond_mode=balance-tcp is required. - LACP deadlock: OVS holds members disabled until negotiation while the partner needs carrier before it will send LACPDUs. lacp-fallback-ab breaks it. - LACPDUs are untagged, so a trunk with no native VLAN has nowhere to put them. - --boot cdrom,hd re-runs the ISO on every restart, so every commit+save went to a live system that evaporated. Install now switches the VM to boot hd. - cloud-init on Alpine: users stay locked without lock_passwd:false, one failing runcmd aborts the rest, busybox here has no httpd applet, and start-stop-daemon --exec /usr/bin/python3 matches cloud-init's own python3. - The user-data heredoc is unquoted, so backticks in a COMMENT were executed by the host shell and their output corrupted the YAML. build_seed now validates with yaml.safe_load before building the ISO. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01DMVzWZgiKW2wquf5z8S1yH
2026-08-13 00:42:39 +01:00
#!/bin/bash
# Prometheus + Grafana for the labsim connectivity matrix.
#
# Grafana runs with anonymous auth as Admin — NO LOGIN. That is deliberate for
# a throwaway lab on localhost; do not copy this into anything reachable.
#
# ./monitoring-up.sh start exporter + prometheus + grafana
# ./monitoring-up.sh --down stop and remove them
#
# Grafana: http://localhost:3000 (dashboard "labsim — VLAN connectivity matrix")
# Prometheus: http://localhost:9090
# Exporter: http://localhost:9101/metrics
set -euo pipefail
SCRIPT_DIR="$(cd "$(dirname "$0")" && pwd)"
source "$SCRIPT_DIR/lib.sh"
GRAFANA_PORT="${GRAFANA_PORT:-3000}"
PROM_PORT="${PROM_PORT:-9090}"
EXPORTER_PORT="${EXPORTER_PORT:-9101}"
NET="labsim-mon"
if [ "${1:-}" = "--down" ]; then
pkill -f "labsim-exporter.py" 2>/dev/null || true
podman rm -f labsim-grafana labsim-prometheus >/dev/null 2>&1 || true
podman network rm -f "$NET" >/dev/null 2>&1 || true
log "monitoring stopped"
exit 0
fi
command -v podman >/dev/null 2>&1 || die "podman not installed"
# --- exporter (on the host: it needs SSH access to the VMs) ----------------
if pgrep -f "labsim-exporter.py" >/dev/null 2>&1; then
log "exporter already running on :$EXPORTER_PORT"
else
log "starting exporter on :$EXPORTER_PORT"
nohup "$SCRIPT_DIR/labsim-exporter.py" --port "$EXPORTER_PORT" --interval 15 \
> /tmp/labsim-exporter.log 2>&1 &
sleep 3
fi
curl -sS --max-time 5 "http://127.0.0.1:${EXPORTER_PORT}/metrics" >/dev/null \
|| die "exporter not answering on :$EXPORTER_PORT (see /tmp/labsim-exporter.log)"
podman network exists "$NET" 2>/dev/null || podman network create "$NET" >/dev/null
# --- prometheus -----------------------------------------------------------
podman rm -f labsim-prometheus >/dev/null 2>&1 || true
log "starting prometheus on :$PROM_PORT"
podman run -d --name labsim-prometheus --network "$NET" \
-p "${PROM_PORT}:9090" \
-v "$SCRIPT_DIR/monitoring/prometheus.yml:/etc/prometheus/prometheus.yml:ro,Z" \
--add-host "host.containers.internal:host-gateway" \
docker.io/prom/prometheus:latest >/dev/null
# --- grafana (anonymous, no login) ----------------------------------------
podman rm -f labsim-grafana >/dev/null 2>&1 || true
log "starting grafana on :$GRAFANA_PORT (anonymous auth — no password)"
podman run -d --name labsim-grafana --network "$NET" \
-p "${GRAFANA_PORT}:3000" \
-e GF_AUTH_ANONYMOUS_ENABLED=true \
-e GF_AUTH_ANONYMOUS_ORG_ROLE=Admin \
-e GF_AUTH_DISABLE_LOGIN_FORM=true \
-e GF_AUTH_BASIC_ENABLED=false \
-e GF_SECURITY_ALLOW_EMBEDDING=true \
-e GF_USERS_DEFAULT_THEME=dark \
-v "$SCRIPT_DIR/monitoring/grafana/provisioning:/etc/grafana/provisioning:ro,Z" \
docker.io/grafana/grafana:latest >/dev/null
log "waiting for grafana..."
for _ in $(seq 1 40); do
if curl -sS --max-time 3 "http://127.0.0.1:${GRAFANA_PORT}/api/health" >/dev/null 2>&1; then
break
fi
sleep 3
done
echo
log "Grafana: http://localhost:${GRAFANA_PORT}/d/labsim-matrix (no login)"
log "Prometheus: http://localhost:${PROM_PORT}"
log "Exporter: http://localhost:${EXPORTER_PORT}/metrics"