`mcpctl get instances` showed all eight servers healthy while the UniFi one
had never once reached its controller. The default probe is `tools/list`,
which MCP servers answer from a static in-process table — no credentials, no
upstream, ~3ms. It cannot fail for any reason the user cares about, so it was
reporting `healthy` for every process that managed to start.
Split the two passes:
healthy — readiness: `tools/call` on `healthCheck.tool`. The upstream
answered, so the server can actually do its job.
live — liveness: `tools/list` only. Process up, upstream unverified.
`live` is now the default for any server without a `healthCheck.tool`. It is
not a warning; it is an admission that nothing is watching that server. Probe
events name which probe ran and which tool ("Readiness check (list_sites)
passed"), so the events log distinguishes the two after the fact.
Also:
- `healthCheck.tool` is optional now, so the timings can be tuned without
inventing a readiness probe.
- `create server --health-check-tool/-args/-interval/-timeout/
-failure-threshold`, per the rule that everything applyable is a create
flag. Merges over a `--from-template` healthCheck rather than replacing it.
- `describe instance` explains a `live` verdict instead of leaving it cryptic.
- create.ts held a raw NUL byte in a string literal, which made grep treat the
whole file as binary and silently skip it. Escaped as `\0`.
Verified against the live fleet: with readiness probes configured, my-grafana
went unhealthy (Grafana API 403) and my-node-red degraded (connect timeout to
a Tailscale address) — both had read healthy for months.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0114dg56YmVacyqhp5fitcTb
6.4 KiB
Reliability: don't let a bad LLM take mcpctl down
The homelab model changes often (fast ↔ thinking, model swaps, backends that drift or go down). mcpctl must stay responsive and honest through all of it.
Principle
LLM-optional operations must be time-bounded, fall back deterministically, and report the degradation — never hang and never degrade silently.
- Bounded: every optional LLM call is wrapped in
withTimeout(Promise.race + anAbortSignalso fetch-based providers actually cancel). A thinking model that streams for minutes can never block the caller. - Deterministic fallback: when the LLM times out or errors, use the non-LLM path (priority/keyword ordering, byte-range pages).
- Loud, not silent: log the reason (
[gate] …,[pagination] …) and tell the user.begin_sessionprepends⚠ Smart prompt-selection unavailable (<reason>)…and setsdegraded: true+degradedReasonon the auditgate_decisionevent.
Applied in: the gate's begin_session prompt selection
(proxymodel/plugins/gate.ts, cap MCPCTL_GATE_LLM_TIMEOUT_MS, default 8s) and
pagination's smart index (llm/pagination.ts, MCPCTL_PAGINATION_LLM_TIMEOUT_MS,
default 10s). read_prompts is LLM-free by design.
Note: the gate's prompt-ranking uses the heavy client provider's own model — it deliberately does not force the project's vLLM model onto it (doing so made every selection fail silently when the model wasn't anthropic-servable).
Instance health: live is not healthy
An MCP server answers tools/list from a static, in-process table. It costs
a few milliseconds, needs no credentials, and reaches no upstream — so it stays
green while the thing the server exists to talk to is unreachable. Treating that
as a health signal is how mcpctl get instances showed eight healthy servers
while the UniFi one had never once reached its controller.
So the probe reports two different passes:
| Status | Probe | Means |
|---|---|---|
healthy |
readiness — tools/call on healthCheck.tool |
The upstream answered. The server can do its job. |
live |
liveness — tools/list only |
The process is up and speaks MCP. Its upstream is unverified. |
degraded |
either, failing | Failing, but under failureThreshold. |
unhealthy |
either, failing | Failed failureThreshold times in a row. |
live is the default for any server with no healthCheck.tool. It is not a
warning — it is an admission that nothing is watching that server's upstream.
Configure a readiness probe on every server. Pick a read-only tool that
genuinely round-trips to the upstream, and verify it passes before configuring
it — a probe naming a local-only tool (get_..._version) or a tool the server
doesn't expose reproduces the same false green it was meant to remove.
mcpctl create server unifi-network --health-check-tool list_sites \
--health-check-interval 60 --health-check-timeout 15 --force
or declaratively — healthCheck round-trips through get -o yaml | apply -f:
healthCheck:
tool: list_sites
arguments: {}
intervalSeconds: 60
timeoutSeconds: 15
failureThreshold: 3
Omit tool to keep liveness while still tuning the timings.
Latency is the tell: a probe answering in single-digit milliseconds is reading a
local table, not crossing a network. The UniFi probe went from 3ms (tools/list,
lying) to 1847ms on its first real list_sites — login, TLS, controller round
trip — and ~40ms once the session was warm.
Two failure modes the probe cannot see for you
The healthy-looking UniFi server was broken twice over, and both are worth checking first when a readiness probe starts failing:
- Egress. MCP server pods default to TCP 80/443 only
(
servers-allow-external-egress). Any upstream on another port — the UniFi controller on:8443— times out on every call. Declare it in Pulumi'smcpctl.serverEgressTargets(name +/32+ ports); don't widen the blanket rule. - Address reachability. A pod cannot reach a Tailscale
100.64.0.0/10address. Config pointing at one connect-timeouts forever. Use LAN IPs.
LLM-essential operations — failover chain
Chat needs an LLM but not a specific one. Instead of failing when the pinned model is down, chat fails over across an ordered chain and reports which model actually answered.
- Chain: an
Llmdeclares fallbacks inextraConfig.fallbacks: string[](Llm names, in order). The dispatcher builds an ordered candidate list — the primary's pool, then each fallback's pool — and tries them in order. - Fails over on real failures, not just transport: a non-2xx status
(e.g. a drifted model's
400) or an empty/invalid completion now advances to the next candidate (chat.service.tsrunOneInference). Streaming fails over pre-first-chunk. - Transparency:
ChatResult/ the SSEfinalframe carryllm,model, andfailedOver. The CLI printsmodel: <llm> (<model>)per turn, and⚠ failed over → answered by <llm> (<model>)when a fallback was used. - Exhaustion is clear: if every candidate fails, the error names the last model + upstream status/body (not "no choice").
Homelab chain: vllm-current (the served vLLM) → an anthropic-fallback server
Llm as the always-up last resort. The chain is owned declaratively by Pulumi
(kubernetes-deployment/deployments/mcpctl/llm-target.ts, fallbacks arg) so a
pulumi up can't wipe it — mcpd replaces extraConfig on update.
A real last resort needs an independent cloud credential. LiteLLM only fronts
the single local vLLM, so any local model shares the same GPU-box failure as the
primary. anthropic-fallback points at api.anthropic.com and only serves once
mcpd Secret anthropic-key holds a genuine sk-ant-api03 API key (from the
Anthropic Console). A Claude subscription OAuth token (sk-ant-oat…, what
claude setup-token mints) does not work: Anthropic gates those to the Claude
Code client — probed directly they 404 on older model ids and 429 on current ones.
mcpd's anthropic adapter does send OAuth tokens via Authorization: Bearer (so
auth passes), but the gating is server-side and unavoidable. Until a real key is
set, the fallback is wired-but-inert: an outage fails over to it and surfaces a
clear anthropic-fallback (model) HTTP 4xx error rather than improving uptime.