diff --git a/README.md b/README.md index 43fcd22..8be6684 100644 --- a/README.md +++ b/README.md @@ -237,6 +237,25 @@ Re-sync skills on their own with See [docs/opencode-extension.md](docs/opencode-extension.md) for full details. +### Run Claude Code on the homelab LLM + +The homelab gateway (LiteLLM in front of vLLM) already serves the Anthropic +Messages API, so Claude Code can talk to it directly — no bridge: + +```bash +claude-vllm # reuse the provider/model/key pi or prime-agent already uses +claude-vllm --model deepseek-v4-max +claude-vllm --list # show what it would use +claude-vllm -- -p "summarise this repo" # args after -- go to claude +``` + +It discovers the endpoint, credential, model and context window from +`~/.pi/agent`, `~/.prime/agent` or opencode's config (first hit wins) and sets +the `ANTHROPIC_*` environment Claude Code needs. + +See [docs/claude-vllm.md](docs/claude-vllm.md), including why routing this +through mcpctl would add surface without adding capability. + ## Declarative Configuration Everything can be defined in YAML and applied with `mcpctl apply`: diff --git a/docs/claude-vllm.md b/docs/claude-vllm.md new file mode 100644 index 0000000..72c3446 --- /dev/null +++ b/docs/claude-vllm.md @@ -0,0 +1,92 @@ +# `claude-vllm` — Claude Code on the homelab LLM + +## The short version + +```bash +claude-vllm # default provider + model +claude-vllm --model deepseek-v4-max +claude-vllm --list # what would it use? +claude-vllm --print-env # the exports, without running claude +claude-vllm -- -p "summarise this repo" # anything after -- goes to claude +``` + +## Why no bridge is needed + +The homelab gateway at `llm.ad.itaz.eu` is **LiteLLM in front of vLLM**, and +LiteLLM already serves the Anthropic Messages API: + +```console +$ curl -s https://llm.ad.itaz.eu/v1/messages -X POST \ + -H 'x-api-key: ' -H 'anthropic-version: 2023-06-01' \ + -d '{"model":"deepseek-v4-fast","max_tokens":32, + "messages":[{"role":"user","content":"say only: ok"}]}' +{"type":"message","role":"assistant","model":"deepseek-v4-fast", + "content":[{"type":"text","text":"ok"}],"stop_reason":"end_turn", …} +``` + +That is the exact protocol Claude Code speaks. So pointing `ANTHROPIC_BASE_URL` +at the gateway is the whole integration — no translation layer, no proxy, no +mcpctl in the request path. + +### Why not route it through mcpctl + +mcpctl's LLM layer is a **client**, not a server: `mcpd` exposes +`/api/v1/llms` (management) and the adapters in +`src/mcpd/src/services/llm/adapters/` *call out* to Anthropic/OpenAI for gating, +prompt selection and agent chat. Nothing in mcpd or mcplocal serves +`/v1/messages`. A `claude-mcpctl` would therefore mean adding an +Anthropic-shaped passthrough to mcpd that re-wraps what LiteLLM already does — +new surface, new failure mode, no new capability. + +It would only be worth it if you wanted mcpctl *in* the LLM path for its own +sake: per-project gating of model calls, audit of prompts, or budget +enforcement. Those are real features, but they are a passthrough endpoint in +mcpd, not a wrapper script. + +## What it discovers, and from where + +`claude-vllm` exists only so you don't paste four exports each time. It reuses +whatever you already configured for another agent — first hit wins: + +| Order | Source | Supplies | +|-------|--------|----------| +| 1 | environment (`ANTHROPIC_BASE_URL`, `ANTHROPIC_AUTH_TOKEN`, `ANTHROPIC_MODEL`) | anything already set is respected | +| 2 | `--provider` / `--model` flags | provider, model | +| 3 | `~/.pi/agent` | `settings.json` → default provider/model · `models.json` → base URL, context window · `auth.json` → key | +| 4 | `~/.prime/agent` | same shape | +| 5 | `~/.config/opencode/opencode.jsonc` | `provider..options.{baseURL,apiKey}` | + +The base URL and the credential always come from the *same* source, so one +gateway's URL is never paired with another's key. + +## What it sets, and why each one + +| Variable | Reason | +|----------|--------| +| `ANTHROPIC_BASE_URL` | The gateway. A stored OpenAI-style `…/v1` base has the suffix stripped — Claude Code appends `/v1/messages` itself, and `/v1/v1/messages` 404s. | +| `ANTHROPIC_AUTH_TOKEN` + `ANTHROPIC_API_KEY` | The gateway takes either; which one Claude Code sends has changed between releases. | +| `ANTHROPIC_MODEL` | The model to drive the session. | +| `ANTHROPIC_SMALL_FAST_MODEL`, `ANTHROPIC_DEFAULT_HAIKU_MODEL` | Background/summarisation calls otherwise ask for a real Haiku the gateway does not serve, and every one 404s. | +| `CLAUDE_CODE_MAX_CONTEXT_TOKENS` | Claude Code assumes 200k for models it has no table for. `deepseek-v4-*` is 393k, so without this it auto-compacts at half capacity. Taken from `contextWindow` in the provider config. | +| `CLAUDE_CODE_DISABLE_EXPERIMENTAL_BETAS` | Beta headers the gateway does not implement make it reject otherwise-fine requests. | +| `CLAUDE_CODE_DISABLE_NONESSENTIAL_TRAFFIC` | No telemetry about a non-Anthropic endpoint. | + +## Expected noise + +``` +⚠ claude.ai connectors are disabled because ANTHROPIC_API_KEY or another auth + source is set and takes precedence over your claude.ai login +``` + +Unavoidable and harmless: it fires for *any* non-claude.ai auth, including +`ANTHROPIC_AUTH_TOKEN` alone (verified). It means your claude.ai org connectors +are not loaded for this session — which is the point of running against the +homelab. + +## Verified + +```console +$ claude-vllm --model deepseek-v4-fast -- -p "Reply with exactly: PARITY-OK" +claude-vllm: itaz · deepseek-v4-fast · https://llm.ad.itaz.eu +PARITY-OK +``` diff --git a/nfpm.yaml b/nfpm.yaml index 047ecc8..3280d1c 100644 --- a/nfpm.yaml +++ b/nfpm.yaml @@ -16,6 +16,10 @@ contents: dst: /usr/bin/mcpctl-local file_info: mode: 0755 + - src: ./stack/claude-vllm + dst: /usr/bin/claude-vllm + file_info: + mode: 0755 - src: ./deploy/mcplocal.service dst: /usr/lib/systemd/user/mcplocal.service file_info: diff --git a/stack/claude-vllm b/stack/claude-vllm new file mode 100755 index 0000000..17942f8 --- /dev/null +++ b/stack/claude-vllm @@ -0,0 +1,178 @@ +#!/usr/bin/env bash +# claude-vllm — run Claude Code against the homelab LLM gateway instead of api.anthropic.com. +# +# The gateway (LiteLLM in front of vLLM) already speaks the Anthropic Messages +# API on /v1/messages, so no bridge or translation layer is needed — Claude Code +# talks to it directly once ANTHROPIC_BASE_URL points there. This script exists +# only to find the endpoint, credential and model you have already configured +# for another agent, instead of making you paste four exports every time. +# +# Discovery order (first hit wins, per field): +# 1. environment already set (ANTHROPIC_BASE_URL / ANTHROPIC_AUTH_TOKEN / ANTHROPIC_MODEL) +# 2. flags (--provider, --model) +# 3. ~/.pi/agent settings.json + models.json + auth.json +# 4. ~/.prime/agent settings.json + models.json + auth.json +# 5. ~/.config/opencode/opencode.jsonc (provider..options.{baseURL,apiKey}) +# +# Usage: +# claude-vllm # default provider + model, then exec claude +# claude-vllm --model deepseek-v4-max +# claude-vllm --provider itaz --model deepseek-v4-fast -- -p "summarise this repo" +# claude-vllm --list # show discoverable providers/models and exit +# claude-vllm --print-env # print the exports and exit (don't run claude) +set -euo pipefail + +PROVIDER="" +MODEL="" +LIST=0 +PRINT_ENV=0 +CLAUDE_ARGS=() + +while [ $# -gt 0 ]; do + case "$1" in + --provider) PROVIDER="${2:?--provider needs a value}"; shift 2 ;; + --model) MODEL="${2:?--model needs a value}"; shift 2 ;; + --list) LIST=1; shift ;; + --print-env) PRINT_ENV=1; shift ;; + -h|--help) sed -n '2,25p' "$0" | sed 's/^# \{0,1\}//'; exit 0 ;; + --) shift; CLAUDE_ARGS+=("$@"); break ;; + *) CLAUDE_ARGS+=("$1"); shift ;; + esac +done + +command -v jq >/dev/null || { echo "claude-vllm: jq is required" >&2; exit 1; } + +PI_HOME="${PI_AGENT_HOME:-$HOME/.pi/agent}" +PRIME_HOME="${PRIME_AGENT_HOME:-$HOME/.prime/agent}" +OC_CONFIG="${XDG_CONFIG_HOME:-$HOME/.config}/opencode/opencode.jsonc" + +# jq tolerates the // comments opencode.jsonc may contain only after we strip +# them; harmless for strict JSON. +read_json() { [ -f "$1" ] && sed 's://[^"]*$::' "$1" | jq -c . 2>/dev/null || echo '{}'; } + +# ── discover ───────────────────────────────────────────────────────────────── +# Each agent home is tried in turn; the first one that yields a base URL wins, +# and the credential is taken from that same home so we never pair one gateway's +# URL with another's key. +discover_from_agent_home() { + local home="$1" settings models auth provider base key model + settings=$(read_json "$home/settings.json") + models=$(read_json "$home/models.json") + auth=$(read_json "$home/auth.json") + + provider="$PROVIDER" + [ -n "$provider" ] || provider=$(jq -r '.defaultProvider // empty' <<<"$settings") + [ -n "$provider" ] || return 1 + + base=$(jq -r --arg p "$provider" '.providers[$p].baseUrl // empty' <<<"$models") + [ -n "$base" ] || return 1 + # `apiKey` in models.json names an env var; the secret itself lives in auth.json. + key=$(jq -r --arg p "$provider" '.[$p].key // empty' <<<"$auth") + + model="$MODEL" + [ -n "$model" ] || model=$(jq -r '.defaultModel // empty' <<<"$settings") + + FOUND_PROVIDER="$provider"; FOUND_BASE="$base"; FOUND_KEY="$key"; FOUND_MODEL="$model" + FOUND_MODELS=$(jq -r --arg p "$provider" '.providers[$p].models[]?.id' <<<"$models") + FOUND_CONTEXT=$(jq -r --arg p "$provider" --arg m "$model" \ + '.providers[$p].models[]? | select(.id==$m) | .contextWindow // empty' <<<"$models") + FOUND_SOURCE="$home" + return 0 +} + +discover_from_opencode() { + local cfg provider base key + cfg=$(read_json "$OC_CONFIG") + provider="$PROVIDER" + if [ -z "$provider" ]; then + provider=$(jq -r '(.model // "") | split("/")[0] // empty' <<<"$cfg") + fi + [ -n "$provider" ] || return 1 + base=$(jq -r --arg p "$provider" '.provider[$p].options.baseURL // empty' <<<"$cfg") + [ -n "$base" ] || return 1 + key=$(jq -r --arg p "$provider" '.provider[$p].options.apiKey // empty' <<<"$cfg") + + FOUND_PROVIDER="$provider"; FOUND_BASE="$base"; FOUND_KEY="$key" + FOUND_MODEL="${MODEL:-$(jq -r '(.model // "") | split("/")[1] // empty' <<<"$cfg")}" + FOUND_MODELS=$(jq -r --arg p "$provider" '.provider[$p].models | keys[]?' <<<"$cfg") + FOUND_CONTEXT=$(jq -r --arg p "$provider" --arg m "$FOUND_MODEL" \ + '.provider[$p].models[$m].limit.context // empty' <<<"$cfg") + FOUND_SOURCE="$OC_CONFIG" + return 0 +} + +FOUND_PROVIDER=""; FOUND_BASE=""; FOUND_KEY=""; FOUND_MODEL=""; FOUND_MODELS=""; FOUND_CONTEXT=""; FOUND_SOURCE="" +discover_from_agent_home "$PI_HOME" \ + || discover_from_agent_home "$PRIME_HOME" \ + || discover_from_opencode \ + || true + +BASE="${ANTHROPIC_BASE_URL:-$FOUND_BASE}" +KEY="${ANTHROPIC_AUTH_TOKEN:-${ANTHROPIC_API_KEY:-$FOUND_KEY}}" +MODEL_ID="${MODEL:-${ANTHROPIC_MODEL:-$FOUND_MODEL}}" + +if [ "$LIST" = 1 ]; then + echo "provider: ${FOUND_PROVIDER:-} (from ${FOUND_SOURCE:-nowhere})" + echo "endpoint: ${BASE:-}" + echo "credential: $([ -n "$KEY" ] && echo "found (${KEY:0:10}…)" || echo "")" + echo "default model: ${MODEL_ID:-}${FOUND_CONTEXT:+ (context ${FOUND_CONTEXT})}" + echo "models:" + [ -n "$FOUND_MODELS" ] && printf ' %s\n' $FOUND_MODELS || echo " " + exit 0 +fi + +if [ -z "$BASE" ]; then + echo "claude-vllm: no LLM endpoint found." >&2 + echo " Looked in $PI_HOME, $PRIME_HOME and $OC_CONFIG." >&2 + echo " Set ANTHROPIC_BASE_URL, or configure a provider in one of those." >&2 + exit 1 +fi +if [ -z "$KEY" ]; then + echo "claude-vllm: found $BASE but no credential for '${FOUND_PROVIDER}'." >&2 + echo " Set ANTHROPIC_AUTH_TOKEN, or add the key to ${FOUND_SOURCE}/auth.json." >&2 + exit 1 +fi + +# Claude Code appends /v1/messages itself, so the stored provider baseUrl's +# trailing /v1 (an OpenAI-style base) has to come off or requests go to +# /v1/v1/messages. +BASE="${BASE%/}"; BASE="${BASE%/v1}" + +export ANTHROPIC_BASE_URL="$BASE" +# Both forms: the gateway accepts either, and which one Claude Code sends has +# changed between releases (Authorization: Bearer vs x-api-key). +export ANTHROPIC_AUTH_TOKEN="$KEY" +export ANTHROPIC_API_KEY="$KEY" +[ -n "$MODEL_ID" ] && export ANTHROPIC_MODEL="$MODEL_ID" +# Without a substitute, the background/summarisation calls ask the gateway for a +# real Haiku it does not serve, and every one of them 404s. +export ANTHROPIC_SMALL_FAST_MODEL="${ANTHROPIC_SMALL_FAST_MODEL:-${CLAUDE_VLLM_FAST_MODEL:-$MODEL_ID}}" +export ANTHROPIC_DEFAULT_HAIKU_MODEL="${ANTHROPIC_DEFAULT_HAIKU_MODEL:-$ANTHROPIC_SMALL_FAST_MODEL}" +# Claude Code only knows the context window of models it ships a table for, and +# assumes 200k for anything else — so a 393k model would auto-compact at half +# its capacity. Tell it the real number when the provider config states one. +if [ -n "${FOUND_CONTEXT:-}" ] && [ -z "${CLAUDE_CODE_MAX_CONTEXT_TOKENS:-}" ]; then + export CLAUDE_CODE_MAX_CONTEXT_TOKENS="$FOUND_CONTEXT" +fi +# Beta headers the gateway does not implement make it reject otherwise fine +# requests. +export CLAUDE_CODE_DISABLE_EXPERIMENTAL_BETAS="${CLAUDE_CODE_DISABLE_EXPERIMENTAL_BETAS:-1}" +# Stops Claude Code phoning home about a non-Anthropic endpoint. +export CLAUDE_CODE_DISABLE_NONESSENTIAL_TRAFFIC="${CLAUDE_CODE_DISABLE_NONESSENTIAL_TRAFFIC:-1}" + +if [ "$PRINT_ENV" = 1 ]; then + for v in ANTHROPIC_BASE_URL ANTHROPIC_MODEL ANTHROPIC_SMALL_FAST_MODEL \ + ANTHROPIC_DEFAULT_HAIKU_MODEL CLAUDE_CODE_MAX_CONTEXT_TOKENS \ + CLAUDE_CODE_DISABLE_EXPERIMENTAL_BETAS \ + CLAUDE_CODE_DISABLE_NONESSENTIAL_TRAFFIC; do + [ -n "${!v:-}" ] || continue + printf 'export %s=%q\n' "$v" "${!v}" + done + # Never printed: the credential. Use --list to confirm one was found. + echo 'export ANTHROPIC_AUTH_TOKEN= # and ANTHROPIC_API_KEY' + exit 0 +fi + +command -v claude >/dev/null || { echo "claude-vllm: claude is not on PATH" >&2; exit 1; } +echo "claude-vllm: ${FOUND_PROVIDER:-custom} · ${ANTHROPIC_MODEL:-} · $ANTHROPIC_BASE_URL" >&2 +exec claude "${CLAUDE_ARGS[@]}"