Files
mcpctl/docs/claude-vllm.md
Michal c5ea39e959 feat(claude): claude-vllm — run Claude Code against the homelab LLM gateway
The gateway at llm.ad.itaz.eu is LiteLLM in front of vLLM, and LiteLLM already
serves the Anthropic Messages API on /v1/messages — verified with a real
completion. So Claude Code needs no bridge: pointing ANTHROPIC_BASE_URL at it is
the whole integration.

`claude-vllm` exists only to stop you pasting four exports each time. It reuses
what another agent is already configured with — ~/.pi/agent, ~/.prime/agent, then
opencode's config, first hit wins — taking the base URL and the credential from
the same source so one gateway's URL is never paired with another's key.

Beyond the obvious ANTHROPIC_* vars it sets two that are easy to miss:

  - ANTHROPIC_SMALL_FAST_MODEL / ANTHROPIC_DEFAULT_HAIKU_MODEL, or the
    background and summarisation calls ask the gateway for a real Haiku it does
    not serve and every one 404s;
  - CLAUDE_CODE_MAX_CONTEXT_TOKENS from the provider's declared contextWindow,
    because Claude Code assumes 200k for models it has no table for — and
    deepseek-v4-* is 393k, so it would auto-compact at half capacity.

On `claude-mcpctl`: mcpctl's LLM layer is a client, not a server. mcpd serves
/api/v1/llms (management) and its adapters call out to providers for gating,
prompt selection and agent chat; nothing serves /v1/messages. Routing Claude
through mcpctl would mean adding an Anthropic-shaped passthrough that re-wraps
LiteLLM — worth doing only if mcpctl in the LLM path buys something of its own
(per-project gating of model calls, prompt audit, budgets), which is a mcpd
endpoint rather than a wrapper script.

Verified: `claude-vllm --model deepseek-v4-fast -- -p "..."` completes against
deepseek on the homelab, with the unknown-model context warning gone.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BVwuCjuMoA13gmzYEfcrNP
2026-08-09 18:39:36 +01:00

4.3 KiB

claude-vllm — Claude Code on the homelab LLM

The short version

claude-vllm                                   # default provider + model
claude-vllm --model deepseek-v4-max
claude-vllm --list                            # what would it use?
claude-vllm --print-env                       # the exports, without running claude
claude-vllm -- -p "summarise this repo"       # anything after -- goes to claude

Why no bridge is needed

The homelab gateway at llm.ad.itaz.eu is LiteLLM in front of vLLM, and LiteLLM already serves the Anthropic Messages API:

$ curl -s https://llm.ad.itaz.eu/v1/messages -X POST \
    -H 'x-api-key: <key>' -H 'anthropic-version: 2023-06-01' \
    -d '{"model":"deepseek-v4-fast","max_tokens":32,
         "messages":[{"role":"user","content":"say only: ok"}]}'
{"type":"message","role":"assistant","model":"deepseek-v4-fast",
 "content":[{"type":"text","text":"ok"}],"stop_reason":"end_turn", …}

That is the exact protocol Claude Code speaks. So pointing ANTHROPIC_BASE_URL at the gateway is the whole integration — no translation layer, no proxy, no mcpctl in the request path.

Why not route it through mcpctl

mcpctl's LLM layer is a client, not a server: mcpd exposes /api/v1/llms (management) and the adapters in src/mcpd/src/services/llm/adapters/ call out to Anthropic/OpenAI for gating, prompt selection and agent chat. Nothing in mcpd or mcplocal serves /v1/messages. A claude-mcpctl would therefore mean adding an Anthropic-shaped passthrough to mcpd that re-wraps what LiteLLM already does — new surface, new failure mode, no new capability.

It would only be worth it if you wanted mcpctl in the LLM path for its own sake: per-project gating of model calls, audit of prompts, or budget enforcement. Those are real features, but they are a passthrough endpoint in mcpd, not a wrapper script.

What it discovers, and from where

claude-vllm exists only so you don't paste four exports each time. It reuses whatever you already configured for another agent — first hit wins:

Order Source Supplies
1 environment (ANTHROPIC_BASE_URL, ANTHROPIC_AUTH_TOKEN, ANTHROPIC_MODEL) anything already set is respected
2 --provider / --model flags provider, model
3 ~/.pi/agent settings.json → default provider/model · models.json → base URL, context window · auth.json → key
4 ~/.prime/agent same shape
5 ~/.config/opencode/opencode.jsonc provider.<name>.options.{baseURL,apiKey}

The base URL and the credential always come from the same source, so one gateway's URL is never paired with another's key.

What it sets, and why each one

Variable Reason
ANTHROPIC_BASE_URL The gateway. A stored OpenAI-style …/v1 base has the suffix stripped — Claude Code appends /v1/messages itself, and /v1/v1/messages 404s.
ANTHROPIC_AUTH_TOKEN + ANTHROPIC_API_KEY The gateway takes either; which one Claude Code sends has changed between releases.
ANTHROPIC_MODEL The model to drive the session.
ANTHROPIC_SMALL_FAST_MODEL, ANTHROPIC_DEFAULT_HAIKU_MODEL Background/summarisation calls otherwise ask for a real Haiku the gateway does not serve, and every one 404s.
CLAUDE_CODE_MAX_CONTEXT_TOKENS Claude Code assumes 200k for models it has no table for. deepseek-v4-* is 393k, so without this it auto-compacts at half capacity. Taken from contextWindow in the provider config.
CLAUDE_CODE_DISABLE_EXPERIMENTAL_BETAS Beta headers the gateway does not implement make it reject otherwise-fine requests.
CLAUDE_CODE_DISABLE_NONESSENTIAL_TRAFFIC No telemetry about a non-Anthropic endpoint.

Expected noise

⚠ claude.ai connectors are disabled because ANTHROPIC_API_KEY or another auth
  source is set and takes precedence over your claude.ai login

Unavoidable and harmless: it fires for any non-claude.ai auth, including ANTHROPIC_AUTH_TOKEN alone (verified). It means your claude.ai org connectors are not loaded for this session — which is the point of running against the homelab.

Verified

$ claude-vllm --model deepseek-v4-fast -- -p "Reply with exactly: PARITY-OK"
claude-vllm: itaz · deepseek-v4-fast · https://llm.ad.itaz.eu
PARITY-OK