Compare commits
2 Commits
7fbb827aa5
...
feat/llm-t
| Author | SHA1 | Date | |
|---|---|---|---|
|
|
de906f9c75 | ||
|
|
2da627e77f |
@@ -33,7 +33,7 @@ Key routing rules:
|
||||
- `secret` / `secretbackend` — credentials
|
||||
- `template` — reusable server blueprint
|
||||
- `project` — workspace grouping servers, prompts, agents
|
||||
- `llm` — server-managed LLM provider (api key + endpoint)
|
||||
- `llm` — server-managed LLM provider (api key + endpoint). Never hardcode a served-model id: rows are named for the role they fill (`vllm-current`/fast, `vllm-think`/heavy), Pulumi repoints their `model` on a swap, and consumers resolve by `tier`. See `docs/llm-tiers.md`.
|
||||
- `agent` — LLM persona pinned to one Llm; project attach surfaces project Prompts as system context, project MCP servers as tools, and exposes the agent itself as an MCP virtual server (`agent-<name>/chat`). See `docs/agents.md`, `docs/chat.md`.
|
||||
- `prompt` / `promptrequest` — curated content / legacy pending proposal (use `proposal` for new work).
|
||||
- `skill` — Claude Code skill bundle (SKILL.md + files + typed metadata). Materialised onto disk by `mcpctl skills sync`. See `docs/skills.md`.
|
||||
|
||||
13
README.md
13
README.md
@@ -516,17 +516,19 @@ description "I review security design — ask me after each major change."
|
||||
That's how agents consult each other.
|
||||
|
||||
```bash
|
||||
# 1) point at an LLM. For your in-cluster qwen3-thinking via LiteLLM:
|
||||
# 1) point at an LLM. Name the ROW for the role it fills, not for the model —
|
||||
# the served model changes, the row should not. See docs/llm-tiers.md.
|
||||
mcpctl create secret litellm-key --data API_KEY=sk-…
|
||||
mcpctl create llm qwen3-thinking \
|
||||
mcpctl create llm vllm-current \
|
||||
--type openai \
|
||||
--model qwen3-thinking \
|
||||
--model deepseek-v4-flash \
|
||||
--tier fast \
|
||||
--url http://litellm.nvidia-nim.svc.cluster.local:4000/v1 \
|
||||
--api-key-ref litellm-key/API_KEY
|
||||
|
||||
# 2) create an agent, pinned to that Llm and attached to a project
|
||||
mcpctl create agent reviewer \
|
||||
--llm qwen3-thinking \
|
||||
--llm vllm-current \
|
||||
--project mcpctl-dev \
|
||||
--description "I review security design — ask me after each major change." \
|
||||
--system-prompt-file ./prompts/reviewer.md \
|
||||
@@ -586,7 +588,8 @@ systemctl --user restart mcplocal
|
||||
|
||||
mcpctl get llm
|
||||
# NAME KIND STATUS TYPE MODEL TIER ID
|
||||
# qwen3-thinking public active openai qwen3-thinking fast ...
|
||||
# vllm-current public active openai deepseek-v4-flash fast ...
|
||||
# vllm-think public active openai deepseek-v4-think heavy ...
|
||||
# vllm-local virtual active openai Qwen/Qwen2.5-7B-Instruct-AWQ fast ...
|
||||
|
||||
mcpctl chat-llm vllm-local
|
||||
|
||||
117
docs/llm-tiers.md
Normal file
117
docs/llm-tiers.md
Normal file
@@ -0,0 +1,117 @@
|
||||
# LLM tiers — tracking the served model without hardcoding it
|
||||
|
||||
The homelab's served model changes. `glm-4.6-reap` became `deepseek-v4-flash`;
|
||||
something else will replace it. Every place mcpctl writes a model id down is a
|
||||
place that silently rots when that happens — a row left requesting a suspended
|
||||
model gets `HTTP 400 model not found` on the next call, and the failure surfaces
|
||||
somewhere unhelpful (a gate that quietly stops ranking prompts, an agent that
|
||||
500s).
|
||||
|
||||
mcpctl's answer is that **it never names a model, and prefers not to name a
|
||||
row**. Consumers ask for a *role*; the registry says who currently fills it.
|
||||
|
||||
## The three layers
|
||||
|
||||
```
|
||||
LiteLLM (llm.ad.itaz.eu) deepseek-v4-flash, -fast, -low, -think, -max
|
||||
▲ served-model ids — change on every model swap
|
||||
│ Pulumi owns this mapping
|
||||
mcpd Llm rows vllm-current (tier: fast)
|
||||
▲ vllm-think (tier: heavy)
|
||||
│ stable identities; `model` follows the deployment
|
||||
mcpctl consumers "give me the fast one"
|
||||
no model id, no row name
|
||||
```
|
||||
|
||||
**Layer 1 → 2 is Pulumi's job.** `deployments/mcpctl/llm-target.ts` in the
|
||||
`kubernetes-deployment` repo declares both rows via the `@mcpctl/pulumi`
|
||||
provider and feeds them the active served-model name. Swap the model, run a
|
||||
targeted `pulumi up`, and the rows follow. See
|
||||
[pulumi-provider-llm-autopoint](../src/pulumi/README.md).
|
||||
|
||||
**Layer 2 → 3 is `src/mcplocal/src/server-llm.ts`.** `tier` is already a
|
||||
first-class field on the `Llm` resource and is set by the same Pulumi resource
|
||||
that sets `model`, so the two cannot drift.
|
||||
|
||||
## Resolution rules
|
||||
|
||||
`resolveServerLlmByTier(client, tier)` returns the row filling a tier, or
|
||||
`null`. A row is eligible when:
|
||||
|
||||
- its `tier` matches exactly;
|
||||
- its `status` is `active` — selecting a row with no live backend just moves
|
||||
the failure to the first inference call;
|
||||
- its `kind` is **not** `virtual`. Virtual rows are backed by some user's
|
||||
`mcplocal` over the SSE control channel, i.e. by a machine and personal
|
||||
credentials nobody chose deliberately. An explicit pin may still name one;
|
||||
automatic resolution must not route a project's traffic through a laptop.
|
||||
|
||||
When a tier holds several rows — `heavy` holds both `vllm-think` and the cloud
|
||||
`anthropic-fallback` — an ordered **well-known name** list breaks the tie:
|
||||
|
||||
| tier | preference order |
|
||||
|------|------------------|
|
||||
| `fast` | `vllm-fast`, `vllm-current` |
|
||||
| `heavy` | `vllm-think`, `vllm-thinking` |
|
||||
|
||||
Names that match nothing are inert. That is deliberate: the list carries both
|
||||
the current names and the symmetric names the rows might be renamed to, so a
|
||||
rename stays a Pulumi-only change and neither deploy order breaks the other.
|
||||
Rows absent from the list are still eligible — they just sort last, then by
|
||||
name, so the choice is stable across calls regardless of mcpd's list order.
|
||||
|
||||
Override the order without a release:
|
||||
|
||||
```bash
|
||||
MCPCTL_LLM_PREFER_FAST=spare-row,vllm-current
|
||||
MCPCTL_LLM_PREFER_HEAVY=vllm-think
|
||||
```
|
||||
|
||||
Resolution never throws. mcpd being unreachable is a normal degraded state for
|
||||
`mcplocal`, and every caller has a fallback path; returning `null` keeps that
|
||||
fallback intact instead of failing session setup.
|
||||
|
||||
## Who uses it
|
||||
|
||||
**Gate prompt-selection** (`project-mcp-endpoint.ts`). Order:
|
||||
|
||||
1. `MCPCTL_GATE_SELECTION_LLM` — a global pin, so selection can sit on a fast
|
||||
no-think Llm while chat keeps a thinking model;
|
||||
2. the project's own `llmProvider` (`none` disables);
|
||||
3. whichever row currently fills the `fast` tier.
|
||||
|
||||
Step 3 is what makes an unpinned project work. Without it, a project with no
|
||||
`llmProvider` fell through to the local personal-token provider — and on a
|
||||
machine whose personal key is stale, that means every gated session silently
|
||||
ran on priority-ordered prompts instead of LLM ranking:
|
||||
|
||||
```
|
||||
[gate] LLM prompt-selection failed: Anthropic HTTP 404: model: claude-opus-4-20250514
|
||||
— falling back to priority-ordered prompts
|
||||
```
|
||||
|
||||
The gate prints a `⚠ Smart prompt-selection unavailable` banner in the
|
||||
`begin_session` response whenever it degrades. Absence of that banner on a
|
||||
project with no pin is what the smoke test asserts.
|
||||
|
||||
Explicit pins still win everywhere — this only changes what happens when
|
||||
nothing is pinned.
|
||||
|
||||
## Why not rename `vllm-current` to `vllm-fast`?
|
||||
|
||||
`name` is immutable in mcpd (agents and projects reference it), so a rename is
|
||||
a data migration plus a Pulumi resource replace, not an edit. Once resolution
|
||||
is tier-based nothing in mcpctl reads the names, so the migration buys nothing.
|
||||
`vllm-fast` would also be actively confusing: the fast row points at
|
||||
`deepseek-v4-flash` (the base route, which carries no forced `extra_body`),
|
||||
while LiteLLM separately serves a `deepseek-v4-fast` route that pins
|
||||
`thinking: false`. Two different things, one name apart.
|
||||
|
||||
## Testing
|
||||
|
||||
- `src/mcplocal/tests/server-llm.test.ts` — selection rules, tiebreak, env
|
||||
override, degraded paths.
|
||||
- `src/mcplocal/tests/smoke/llm-tier.smoke.test.ts` — against the live stack:
|
||||
both tiers resolve, the resolved fast row answers real inference (the drift
|
||||
check — a row pointing at a suspended model fails here), and an unpinned
|
||||
gated project gets LLM-ranked selection.
|
||||
@@ -14,6 +14,7 @@ import type { JSONRPCMessage } from '@modelcontextprotocol/sdk/types.js';
|
||||
import { McpRouter } from '../router.js';
|
||||
import { ResponsePaginator } from '../llm/pagination.js';
|
||||
import { refreshProjectUpstreams, fetchProjectLlmConfig } from '../discovery.js';
|
||||
import { resolveServerLlmByTier } from '../server-llm.js';
|
||||
import { loadProjectLlmOverride } from './config.js';
|
||||
import type { McpdClient } from './mcpd-client.js';
|
||||
import type { ProviderRegistry } from '../providers/registry.js';
|
||||
@@ -149,12 +150,25 @@ export function registerProjectMcpEndpoint(app: FastifyInstance, mcpdClient: Mcp
|
||||
if (resolvedModel) pluginConfig.modelOverride = resolvedModel;
|
||||
// Route gate prompt-selection through a server Llm (mcpd inference proxy)
|
||||
// so cloud/server keys stay at the k8s level; the local personal-token
|
||||
// provider is the fallback. See credential-tiering. A dedicated fast
|
||||
// (no-think) selection Llm can be pinned globally via
|
||||
// MCPCTL_GATE_SELECTION_LLM — it overrides the project's chat llmProvider so
|
||||
// selection stays fast while chat keeps its (thinking) model.
|
||||
const gateSelectionLlm = process.env['MCPCTL_GATE_SELECTION_LLM'] || mcpdConfig.llmProvider;
|
||||
if (gateSelectionLlm) pluginConfig.llmProvider = gateSelectionLlm;
|
||||
// provider is the fallback. See credential-tiering. Resolution order:
|
||||
// 1. MCPCTL_GATE_SELECTION_LLM — a global pin, so selection can stay on a
|
||||
// fast (no-think) Llm while chat keeps its thinking model;
|
||||
// 2. the project's own llmProvider ('none' disables, handled downstream);
|
||||
// 3. whichever row currently fills the 'fast' tier (see server-llm.ts).
|
||||
//
|
||||
// (3) is what keeps this working across a model swap: Pulumi repoints the
|
||||
// row's `model` and mcpctl names neither the model nor the row. Without it
|
||||
// an unpinned project silently fell through to the local heavy provider —
|
||||
// which on a machine with no valid personal key means every gated session
|
||||
// has been running on priority-ordered prompts, not LLM selection.
|
||||
const pinnedSelectionLlm = process.env['MCPCTL_GATE_SELECTION_LLM'] ?? '';
|
||||
let gateSelectionLlm = pinnedSelectionLlm !== '' ? pinnedSelectionLlm : mcpdConfig.llmProvider;
|
||||
if ((gateSelectionLlm === undefined || gateSelectionLlm === '') && !llmDisabled) {
|
||||
gateSelectionLlm = (await resolveServerLlmByTier(requestClient, 'fast')) ?? undefined;
|
||||
}
|
||||
if (gateSelectionLlm !== undefined && gateSelectionLlm !== '') {
|
||||
pluginConfig.llmProvider = gateSelectionLlm;
|
||||
}
|
||||
const basePlugin = createDefaultPlugin(pluginConfig);
|
||||
// Optional favourite-index presentation: curated favourite/<tool> + full
|
||||
// all/<server>/<tool> + a "prefer favourite/" instruction. Composed AFTER
|
||||
|
||||
112
src/mcplocal/src/server-llm.ts
Normal file
112
src/mcplocal/src/server-llm.ts
Normal file
@@ -0,0 +1,112 @@
|
||||
/**
|
||||
* Tier-based resolution of mcpd server `Llm` rows.
|
||||
*
|
||||
* The homelab's served model changes (glm-4.6-reap → deepseek-v4-flash → …).
|
||||
* Pulumi owns that swap: `deployments/mcpctl/llm-target.ts` in the
|
||||
* kubernetes-deployment repo repoints long-lived mcpd `Llm` rows at whatever
|
||||
* LiteLLM currently serves. mcpctl must therefore never name a *model* — and
|
||||
* ideally not name a *row* either, or every swap risks stranding a pin the way
|
||||
* `vllm-think` was stranded on a suspended `glm-4.6-reap`.
|
||||
*
|
||||
* So consumers ask for a ROLE (a tier) and get back whichever row currently
|
||||
* fills it. `tier` is already a first-class field on the Llm resource and is
|
||||
* set by the same Pulumi resource that sets `model`, so the two cannot drift.
|
||||
*
|
||||
* Well-known names are only a TIEBREAK, used when a tier has several rows
|
||||
* (today `heavy` holds both `vllm-think` and the cloud `anthropic-fallback`).
|
||||
* They are an ordered list rather than a single constant so that renaming a row
|
||||
* on the Pulumi side stays a Pulumi-only change: list both the old and the new
|
||||
* name and neither deploy order breaks the other.
|
||||
*/
|
||||
import type { McpdClient } from './http/mcpd-client.js';
|
||||
|
||||
export type LlmTier = 'fast' | 'heavy';
|
||||
|
||||
/**
|
||||
* Preferred row names per tier, most-preferred first. A name that matches
|
||||
* nothing is inert, which is what makes this safe to pre-seed with names that
|
||||
* do not exist yet (`vllm-fast`, `vllm-thinking` are the symmetric names the
|
||||
* rows may be renamed to; `vllm-current`/`vllm-think` are what Pulumi owns
|
||||
* today). Rows absent from this list are still eligible — they just sort last.
|
||||
*/
|
||||
export const WELL_KNOWN_LLM_NAMES: Record<LlmTier, readonly string[]> = {
|
||||
fast: ['vllm-fast', 'vllm-current'],
|
||||
heavy: ['vllm-think', 'vllm-thinking'],
|
||||
};
|
||||
|
||||
/** Env override for the tiebreak order, e.g. `MCPCTL_LLM_PREFER_FAST=a,b`. */
|
||||
const PREFER_ENV: Record<LlmTier, string> = {
|
||||
fast: 'MCPCTL_LLM_PREFER_FAST',
|
||||
heavy: 'MCPCTL_LLM_PREFER_HEAVY',
|
||||
};
|
||||
|
||||
/** Subset of mcpd's LlmView that tier resolution actually reads. */
|
||||
export interface LlmSummary {
|
||||
name: string;
|
||||
tier?: string;
|
||||
kind?: string;
|
||||
status?: string;
|
||||
}
|
||||
|
||||
/** Read the preference order for a tier: env override, else the built-in list. */
|
||||
export function preferredNamesForTier(tier: LlmTier): readonly string[] {
|
||||
const raw = process.env[PREFER_ENV[tier]];
|
||||
if (raw === undefined || raw.trim() === '') return WELL_KNOWN_LLM_NAMES[tier];
|
||||
return raw.split(',').map((s) => s.trim()).filter((s) => s !== '');
|
||||
}
|
||||
|
||||
/**
|
||||
* Pick the row that fills `tier`, or null if none does.
|
||||
*
|
||||
* Eligibility:
|
||||
* - `tier` matches exactly;
|
||||
* - `status` is 'active' — an inactive row has no live backend, and silently
|
||||
* selecting one just moves the failure to the first inference call;
|
||||
* - `kind` is not 'virtual'. Virtual rows are backed by some user's mcplocal
|
||||
* over the SSE control channel, i.e. by a machine and personal credentials
|
||||
* we did not choose. An explicit pin may still name one; automatic
|
||||
* resolution must not route a project's traffic through someone's laptop.
|
||||
*
|
||||
* Ordering: position in `preferred` (unlisted rows sort last), then name, so
|
||||
* the choice is stable across calls and independent of mcpd's list order.
|
||||
*/
|
||||
export function pickLlmForTier(
|
||||
llms: readonly LlmSummary[],
|
||||
tier: LlmTier,
|
||||
preferred: readonly string[] = preferredNamesForTier(tier),
|
||||
): string | null {
|
||||
const eligible = llms.filter(
|
||||
(l) => l.tier === tier && (l.status ?? 'active') === 'active' && l.kind !== 'virtual',
|
||||
);
|
||||
if (eligible.length === 0) return null;
|
||||
|
||||
const rank = (name: string): number => {
|
||||
const i = preferred.indexOf(name);
|
||||
return i === -1 ? Number.MAX_SAFE_INTEGER : i;
|
||||
};
|
||||
eligible.sort((a, b) => rank(a.name) - rank(b.name) || a.name.localeCompare(b.name));
|
||||
return eligible[0]?.name ?? null;
|
||||
}
|
||||
|
||||
/**
|
||||
* Resolve the name of the mcpd `Llm` currently filling `tier`, or null.
|
||||
*
|
||||
* Never throws: mcpd being unreachable is a normal degraded state for
|
||||
* mcplocal, and every caller has a fallback path (the gate falls back to the
|
||||
* local personal-token provider). Returning null keeps that fallback intact
|
||||
* instead of failing the whole session setup.
|
||||
*/
|
||||
export async function resolveServerLlmByTier(
|
||||
mcpdClient: McpdClient,
|
||||
tier: LlmTier,
|
||||
preferred?: readonly string[],
|
||||
): Promise<string | null> {
|
||||
let llms: LlmSummary[];
|
||||
try {
|
||||
llms = await mcpdClient.get<LlmSummary[]>('/api/v1/llms');
|
||||
} catch {
|
||||
return null;
|
||||
}
|
||||
if (!Array.isArray(llms)) return null;
|
||||
return pickLlmForTier(llms, tier, preferred ?? preferredNamesForTier(tier));
|
||||
}
|
||||
122
src/mcplocal/tests/server-llm.test.ts
Normal file
122
src/mcplocal/tests/server-llm.test.ts
Normal file
@@ -0,0 +1,122 @@
|
||||
import { describe, it, expect, afterEach } from 'vitest';
|
||||
import {
|
||||
pickLlmForTier,
|
||||
preferredNamesForTier,
|
||||
resolveServerLlmByTier,
|
||||
WELL_KNOWN_LLM_NAMES,
|
||||
type LlmSummary,
|
||||
} from '../src/server-llm.js';
|
||||
import type { McpdClient } from '../src/http/mcpd-client.js';
|
||||
|
||||
/** Minimal McpdClient stand-in — tier resolution only ever calls `get`. */
|
||||
function stubClient(impl: () => Promise<unknown>): McpdClient {
|
||||
return { get: impl } as unknown as McpdClient;
|
||||
}
|
||||
|
||||
const row = (over: Partial<LlmSummary> & { name: string }): LlmSummary => ({
|
||||
tier: 'fast',
|
||||
kind: 'public',
|
||||
status: 'active',
|
||||
...over,
|
||||
});
|
||||
|
||||
/** The live shape as of the DeepSeek-V4-Flash-0731 swap. */
|
||||
const LIVE: LlmSummary[] = [
|
||||
row({ name: 'anthropic-fallback', tier: 'heavy' }),
|
||||
row({ name: 'vllm-current', tier: 'fast' }),
|
||||
row({ name: 'vllm-think', tier: 'heavy' }),
|
||||
];
|
||||
|
||||
afterEach(() => {
|
||||
delete process.env['MCPCTL_LLM_PREFER_FAST'];
|
||||
delete process.env['MCPCTL_LLM_PREFER_HEAVY'];
|
||||
});
|
||||
|
||||
describe('pickLlmForTier', () => {
|
||||
it('picks the only row filling a tier', () => {
|
||||
expect(pickLlmForTier(LIVE, 'fast')).toBe('vllm-current');
|
||||
});
|
||||
|
||||
it('breaks a multi-row tier on the well-known order, not list order', () => {
|
||||
// 'heavy' holds both the homelab row and the cloud last-resort. The cloud
|
||||
// row sorts first alphabetically, so this fails without the tiebreak.
|
||||
expect(pickLlmForTier(LIVE, 'heavy')).toBe('vllm-think');
|
||||
});
|
||||
|
||||
it('prefers a renamed row over the current one when both exist', () => {
|
||||
// The deploy window during a Pulumi rename: both names are present.
|
||||
const during = [...LIVE, row({ name: 'vllm-fast', tier: 'fast' })];
|
||||
expect(pickLlmForTier(during, 'fast')).toBe('vllm-fast');
|
||||
});
|
||||
|
||||
it('still resolves a tier whose rows are all unknown names', () => {
|
||||
const renamed = [row({ name: 'homelab-quick', tier: 'fast' })];
|
||||
expect(pickLlmForTier(renamed, 'fast')).toBe('homelab-quick');
|
||||
});
|
||||
|
||||
it('sorts unknown names deterministically', () => {
|
||||
const many = [row({ name: 'zeta' }), row({ name: 'alpha' }), row({ name: 'mid' })];
|
||||
expect(pickLlmForTier(many, 'fast')).toBe('alpha');
|
||||
expect(pickLlmForTier([...many].reverse(), 'fast')).toBe('alpha');
|
||||
});
|
||||
|
||||
it('skips inactive rows — a dead backend just defers the failure', () => {
|
||||
const rows = [row({ name: 'vllm-current', status: 'inactive' }), row({ name: 'spare' })];
|
||||
expect(pickLlmForTier(rows, 'fast')).toBe('spare');
|
||||
});
|
||||
|
||||
it('never auto-selects a virtual row', () => {
|
||||
// Virtual rows are backed by some user's mcplocal over SSE. An explicit
|
||||
// pin may name one; automatic resolution must not.
|
||||
const rows = [row({ name: 'vllm-current', kind: 'virtual' })];
|
||||
expect(pickLlmForTier(rows, 'fast')).toBeNull();
|
||||
});
|
||||
|
||||
it('treats a missing status as active (older rows)', () => {
|
||||
expect(pickLlmForTier([{ name: 'legacy', tier: 'fast' }], 'fast')).toBe('legacy');
|
||||
});
|
||||
|
||||
it('returns null when the tier is empty', () => {
|
||||
expect(pickLlmForTier([row({ name: 'x', tier: 'heavy' })], 'fast')).toBeNull();
|
||||
expect(pickLlmForTier([], 'fast')).toBeNull();
|
||||
});
|
||||
});
|
||||
|
||||
describe('preferredNamesForTier', () => {
|
||||
it('defaults to the well-known list', () => {
|
||||
expect(preferredNamesForTier('fast')).toEqual(WELL_KNOWN_LLM_NAMES.fast);
|
||||
expect(preferredNamesForTier('heavy')).toEqual(WELL_KNOWN_LLM_NAMES.heavy);
|
||||
});
|
||||
|
||||
it('honours the env override so a repoint needs no release', () => {
|
||||
process.env['MCPCTL_LLM_PREFER_FAST'] = ' spare , vllm-current ';
|
||||
expect(preferredNamesForTier('fast')).toEqual(['spare', 'vllm-current']);
|
||||
expect(pickLlmForTier([...LIVE, row({ name: 'spare' })], 'fast')).toBe('spare');
|
||||
});
|
||||
|
||||
it('ignores a blank override', () => {
|
||||
process.env['MCPCTL_LLM_PREFER_FAST'] = ' ';
|
||||
expect(preferredNamesForTier('fast')).toEqual(WELL_KNOWN_LLM_NAMES.fast);
|
||||
});
|
||||
});
|
||||
|
||||
describe('resolveServerLlmByTier', () => {
|
||||
it('resolves against mcpd', async () => {
|
||||
const client = stubClient(async () => LIVE);
|
||||
await expect(resolveServerLlmByTier(client, 'fast')).resolves.toBe('vllm-current');
|
||||
await expect(resolveServerLlmByTier(client, 'heavy')).resolves.toBe('vllm-think');
|
||||
});
|
||||
|
||||
it('returns null when mcpd is unreachable instead of throwing', async () => {
|
||||
// Callers have a fallback path; failing session setup would be worse.
|
||||
const client = stubClient(async () => {
|
||||
throw new Error('Cannot connect to mcpd');
|
||||
});
|
||||
await expect(resolveServerLlmByTier(client, 'fast')).resolves.toBeNull();
|
||||
});
|
||||
|
||||
it('returns null on an unexpected payload shape', async () => {
|
||||
const client = stubClient(async () => ({ error: 'forbidden' }));
|
||||
await expect(resolveServerLlmByTier(client, 'fast')).resolves.toBeNull();
|
||||
});
|
||||
});
|
||||
130
src/mcplocal/tests/smoke/llm-tier.smoke.test.ts
Normal file
130
src/mcplocal/tests/smoke/llm-tier.smoke.test.ts
Normal file
@@ -0,0 +1,130 @@
|
||||
/**
|
||||
* Smoke test: tier-based server-Llm resolution against the live stack.
|
||||
*
|
||||
* mcpctl must keep working across a homelab model swap without an mcpctl edit.
|
||||
* Pulumi (kubernetes-deployment `deployments/mcpctl/llm-target.ts`) repoints
|
||||
* long-lived mcpd `Llm` rows at whatever LiteLLM currently serves; mcpctl names
|
||||
* neither the model nor the row, and resolves by tier instead.
|
||||
*
|
||||
* The three things that have to hold for that to be true:
|
||||
* 1. some active, non-virtual row fills the `fast` tier, and it resolves;
|
||||
* 2. that row's `model` is a route LiteLLM actually serves right now — this
|
||||
* is the assertion that catches drift (a suspended model leaves the row
|
||||
* requesting an id the gateway answers with HTTP 400);
|
||||
* 3. a gated project with NO llmProvider pin still gets LLM-ranked prompt
|
||||
* selection, rather than silently degrading to priority order.
|
||||
*
|
||||
* Run with: pnpm test:smoke
|
||||
*/
|
||||
import { describe, it, expect, beforeAll, afterAll } from 'vitest';
|
||||
import { SmokeMcpSession, isMcplocalRunning, loadMcpdAuth, mcpctl } from './mcp-client.js';
|
||||
import { pickLlmForTier, type LlmSummary } from '../../src/server-llm.js';
|
||||
|
||||
const PROJECT_NAME = 'smoke-llm-tier';
|
||||
|
||||
/** Call mcpd directly with the CLI's own credentials. */
|
||||
async function mcpd<T>(path: string, body?: unknown): Promise<T> {
|
||||
const { token, url } = loadMcpdAuth();
|
||||
const res = await fetch(`${url.replace(/\/$/, '')}${path}`, {
|
||||
method: body === undefined ? 'GET' : 'POST',
|
||||
headers: {
|
||||
Authorization: `Bearer ${token}`,
|
||||
Accept: 'application/json',
|
||||
...(body === undefined ? {} : { 'Content-Type': 'application/json' }),
|
||||
},
|
||||
...(body === undefined ? {} : { body: JSON.stringify(body) }),
|
||||
signal: AbortSignal.timeout(60_000),
|
||||
});
|
||||
const text = await res.text();
|
||||
if (!res.ok) throw new Error(`mcpd ${path} → ${String(res.status)}: ${text.slice(0, 300)}`);
|
||||
return JSON.parse(text) as T;
|
||||
}
|
||||
|
||||
describe('Smoke: tier-based server Llm resolution', () => {
|
||||
let ready = false;
|
||||
let llms: LlmSummary[] = [];
|
||||
|
||||
beforeAll(async () => {
|
||||
console.log('\n ━━━ Smoke Test: llm tier resolution ━━━');
|
||||
if (!(await isMcplocalRunning())) {
|
||||
console.log(' ✗ mcplocal not running — skipping\n');
|
||||
return;
|
||||
}
|
||||
try {
|
||||
llms = await mcpd<LlmSummary[]>('/api/v1/llms');
|
||||
} catch (err) {
|
||||
console.log(` ✗ cannot list Llms: ${err instanceof Error ? err.message : String(err)}`);
|
||||
return;
|
||||
}
|
||||
// A gated project with no --llm pin: exactly the case that used to fall
|
||||
// through to the local personal-token provider.
|
||||
try {
|
||||
await mcpctl(`create project ${PROJECT_NAME} --force --server docmost`);
|
||||
ready = true;
|
||||
} catch (err) {
|
||||
console.log(` ⚠ project setup error: ${err instanceof Error ? err.message : String(err)}`);
|
||||
}
|
||||
}, 90_000);
|
||||
|
||||
afterAll(async () => {
|
||||
try { await mcpctl(`delete project ${PROJECT_NAME}`); } catch { /* best effort cleanup */ }
|
||||
console.log('\n ━━━ llm tier smoke complete ━━━\n');
|
||||
});
|
||||
|
||||
it('resolves an active non-virtual row for the fast tier', () => {
|
||||
if (llms.length === 0) return;
|
||||
const picked = pickLlmForTier(llms, 'fast');
|
||||
expect(picked, 'no Llm fills the fast tier — Pulumi mcpctl-reasoning-llm may not have run').not.toBeNull();
|
||||
|
||||
const row = llms.find((l) => l.name === picked);
|
||||
expect(row?.status ?? 'active').toBe('active');
|
||||
expect(row?.kind).not.toBe('virtual');
|
||||
console.log(` ✓ fast tier → ${String(picked)}`);
|
||||
});
|
||||
|
||||
it('resolves an active non-virtual row for the heavy tier', () => {
|
||||
if (llms.length === 0) return;
|
||||
const picked = pickLlmForTier(llms, 'heavy');
|
||||
expect(picked, 'no Llm fills the heavy tier').not.toBeNull();
|
||||
console.log(` ✓ heavy tier → ${String(picked)}`);
|
||||
});
|
||||
|
||||
it('the fast-tier row points at a model the gateway actually serves', async () => {
|
||||
if (llms.length === 0) return;
|
||||
const picked = pickLlmForTier(llms, 'fast');
|
||||
if (picked === null) return; // already failed above
|
||||
|
||||
// Drift check: a row left pointing at a suspended model answers HTTP 400
|
||||
// ("model X not found"), which `mcpd()` surfaces as a thrown error.
|
||||
const resp = await mcpd<{ choices?: Array<{ message?: { content?: string | null } }> }>(
|
||||
`/api/v1/llms/${encodeURIComponent(picked)}/infer`,
|
||||
{ messages: [{ role: 'user', content: 'Reply with the single word: ok' }], max_tokens: 16, stream: false },
|
||||
);
|
||||
expect(resp.choices?.length ?? 0).toBeGreaterThan(0);
|
||||
console.log(` ✓ ${picked} answered live inference`);
|
||||
}, 120_000);
|
||||
|
||||
it('an unpinned gated project gets LLM-ranked selection, not the degraded fallback', async () => {
|
||||
if (!ready) return;
|
||||
const session = new SmokeMcpSession(PROJECT_NAME);
|
||||
try {
|
||||
await session.initialize();
|
||||
const result = await session.callTool(
|
||||
'begin_session',
|
||||
{ description: 'reviewing documentation pages', tags: ['docs', 'wiki', 'search'] },
|
||||
120_000,
|
||||
);
|
||||
const text = result.content.map((c) => c.text ?? '').join('\n');
|
||||
|
||||
// The gate prints this banner whenever it falls back to priority order.
|
||||
// Its absence is the whole point of tier resolution: no pin anywhere, and
|
||||
// selection still ran through a server Llm.
|
||||
expect(text, 'gate degraded — tier resolution did not reach a server Llm').not.toContain(
|
||||
'Smart prompt-selection unavailable',
|
||||
);
|
||||
console.log(' ✓ begin_session used LLM-ranked selection with no llmProvider pin');
|
||||
} finally {
|
||||
await session.close();
|
||||
}
|
||||
}, 180_000);
|
||||
});
|
||||
@@ -32,6 +32,16 @@ function httpRequest(opts: {
|
||||
headers?: Record<string, string>;
|
||||
body?: string;
|
||||
timeout?: number;
|
||||
/**
|
||||
* Resolve as soon as response headers arrive, without waiting for the body
|
||||
* to end. Required for endpoints that never end: an SSE stream's socket only
|
||||
* goes idle when nothing else is talking to mcplocal, so waiting for the
|
||||
* inactivity timeout made this depend on whether other smoke files happened
|
||||
* to be generating traffic concurrently. Resolving on headers is also a
|
||||
* STRONGER assertion — the caller sees the real status instead of inferring
|
||||
* "reachable" from a timeout, which would pass just as happily on a slow 401.
|
||||
*/
|
||||
resolveOnHeaders?: boolean;
|
||||
}): Promise<{ status: number; headers: http.IncomingHttpHeaders; body: string }> {
|
||||
return new Promise((resolve, reject) => {
|
||||
const parsed = new URL(opts.url);
|
||||
@@ -46,6 +56,12 @@ function httpRequest(opts: {
|
||||
timeout: opts.timeout ?? 10_000,
|
||||
},
|
||||
(res) => {
|
||||
if (opts.resolveOnHeaders === true) {
|
||||
resolve({ status: res.statusCode ?? 0, headers: res.headers, body: '' });
|
||||
res.destroy();
|
||||
req.destroy();
|
||||
return;
|
||||
}
|
||||
const chunks: Buffer[] = [];
|
||||
res.on('data', (chunk: Buffer) => chunks.push(chunk));
|
||||
res.on('end', () => {
|
||||
@@ -92,22 +108,19 @@ describe('Smoke: Security — mcplocal unauthenticated endpoints', () => {
|
||||
if (!available) return;
|
||||
|
||||
// /inspect streams ALL MCP traffic (tool calls, arguments, responses)
|
||||
// for ALL projects to any unauthenticated local client
|
||||
// for ALL projects to any unauthenticated local client. The stream never
|
||||
// ends, so take the status off the response headers and hang up.
|
||||
const res = await httpRequest({
|
||||
url: `${MCPLOCAL_URL}/inspect`,
|
||||
method: 'GET',
|
||||
headers: { 'Accept': 'text/event-stream' },
|
||||
timeout: 3_000,
|
||||
}).catch((err) => {
|
||||
// Timeout is expected (SSE keeps connection open) — still means endpoint is accessible
|
||||
if ((err as Error).message.includes('timed out')) {
|
||||
return { status: 200, headers: {} as http.IncomingHttpHeaders, body: '' };
|
||||
}
|
||||
throw err;
|
||||
timeout: 5_000,
|
||||
resolveOnHeaders: true,
|
||||
});
|
||||
|
||||
// Should be accessible without auth (documenting the vulnerability)
|
||||
expect(res.status).toBeLessThan(400);
|
||||
expect(res.headers['content-type']).toContain('text/event-stream');
|
||||
console.log(` ⚠ /inspect accessible without auth (status ${res.status})`);
|
||||
}, 10_000);
|
||||
|
||||
|
||||
Reference in New Issue
Block a user