Files
mcpctl/docs/reliability.md
Michal 18524c1c57
Some checks failed
CI/CD / typecheck (pull_request) Successful in 1m4s
CI/CD / lint (pull_request) Successful in 2m14s
CI/CD / test (pull_request) Successful in 1m18s
CI/CD / build (pull_request) Successful in 2m30s
CI/CD / smoke (pull_request) Failing after 3m18s
CI/CD / publish (pull_request) Has been skipped
feat(chat): LLM failover chain + show which model answered
Chat is LLM-essential but not model-specific — instead of failing when the
pinned model is down/drifted, it now fails over across an ordered chain and
reports which model actually answered.

- Ordered fallback: an Llm declares `extraConfig.fallbacks: string[]`; the
  dispatcher builds primary-pool → fallback-pool(s) candidates and tries them
  in order (resolveCandidatesWithFallbacks).
- Fail over on real failures, not just transport: runOneInference now advances
  on a non-2xx status (e.g. a 400 from a drifted model) or an empty/invalid
  completion, not only thrown transport errors. Streaming fails over
  pre-first-chunk (already threw on 4xx).
- Transparency: ChatResult + the SSE `final` frame carry {llm, model,
  failedOver}; the CLI prints `model: <llm> (<model>)` each turn and
  `⚠ failed over → answered by …` when a fallback was used.
- Exhaustion names the last model + upstream body (not "no choice").

Tests: 3 failover unit tests (primary 400 → fallback answers + model reported;
primary answers → failedOver=false; all fail → clear aggregated error).
mcpd 948 + CLI 508 green; tsc + lint clean. docs/reliability.md updated.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-07-21 23:39:56 +01:00

2.7 KiB

Reliability: don't let a bad LLM take mcpctl down

The homelab model changes often (fast ↔ thinking, model swaps, backends that drift or go down). mcpctl must stay responsive and honest through all of it.

Principle

LLM-optional operations must be time-bounded, fall back deterministically, and report the degradation — never hang and never degrade silently.

  • Bounded: every optional LLM call is wrapped in withTimeout (Promise.race + an AbortSignal so fetch-based providers actually cancel). A thinking model that streams for minutes can never block the caller.
  • Deterministic fallback: when the LLM times out or errors, use the non-LLM path (priority/keyword ordering, byte-range pages).
  • Loud, not silent: log the reason ([gate] …, [pagination] …) and tell the user. begin_session prepends ⚠ Smart prompt-selection unavailable (<reason>)… and sets degraded: true + degradedReason on the audit gate_decision event.

Applied in: the gate's begin_session prompt selection (proxymodel/plugins/gate.ts, cap MCPCTL_GATE_LLM_TIMEOUT_MS, default 8s) and pagination's smart index (llm/pagination.ts, MCPCTL_PAGINATION_LLM_TIMEOUT_MS, default 10s). read_prompts is LLM-free by design.

Note: the gate's prompt-ranking uses the heavy client provider's own model — it deliberately does not force the project's vLLM model onto it (doing so made every selection fail silently when the model wasn't anthropic-servable).

LLM-essential operations — failover chain

Chat needs an LLM but not a specific one. Instead of failing when the pinned model is down, chat fails over across an ordered chain and reports which model actually answered.

  • Chain: an Llm declares fallbacks in extraConfig.fallbacks: string[] (Llm names, in order). The dispatcher builds an ordered candidate list — the primary's pool, then each fallback's pool — and tries them in order.
  • Fails over on real failures, not just transport: a non-2xx status (e.g. a drifted model's 400) or an empty/invalid completion now advances to the next candidate (chat.service.ts runOneInference). Streaming fails over pre-first-chunk.
  • Transparency: ChatResult / the SSE final frame carry llm, model, and failedOver. The CLI prints model: <llm> (<model>) per turn, and ⚠ failed over → answered by <llm> (<model>) when a fallback was used.
  • Exhaustion is clear: if every candidate fails, the error names the last model + upstream status/body (not "no choice").

Homelab chain: vllm-current (the served vLLM) → an anthropic-fallback server Llm as the always-up last resort.