341 of 356 projects on the shared mcpd were smoke-test leftovers, plus 11
never-expiring `smoke-*` mcptokens sitting on the real `mcpctl-development`
and `sre` projects. Enough to make the project picker unusable.
Two causes, both silent:
- `delete project X --force` — `--force` is valid on `create project` but NOT
on `delete`, so every cleanup call exited non-zero and deleted nothing.
project-llm-ref's afterAll looked correct and had never worked, which is why
there were 84 each of smoke-proj-{ok,orphan,none}-*.
- mcptoken.smoke had no afterAll at all; its cleanup was an `it()` at the end
of the file, so it was skipped whenever an earlier assertion failed.
Cleanup now lives in `afterAll` (runs on failure too) and is no longer gated
on the health probe: the project is created through mcpd, which can be up when
the gateway probe is not, and deleting a project that was never created is a
no-op.
Adds `pnpm smoke:clean` for the case afterAll cannot cover — a killed process
(CI timeout, Ctrl-C, OOM). Dry-run by default, `--yes` to apply. Only touches
`smoke-*` names and protects the shared `smoke-data` / `smoke-aws-docs`
fixture, which five suites depend on and which must survive a concurrent run.
Verified end to end: created a throwaway smoke project, confirmed the dry run
lists without deleting, then `--yes` removed it and left all 15 real projects
untouched.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017BMXdb2qZbPSh8Q7XpTyjB
prime-agent emits session_start from exactly one place — reload() — and never
at startup. A session_start-only handler therefore never ran on a fresh
session, and the '/mcpctl -> already on X' path returns without reloading, so
the indicator stayed blank exactly when it was most wanted.
Also guard against the no-op UI context: the runtime hands extensions
noOpUIContext until the TUI binds the real one, and every setter on it
silently discards. Publishing under it would cache a label that never
rendered.
Now published on session_start (reload), turn_start (earliest reliable point
on a fresh session), and on entering /mcpctl. Repeat calls are a no-op unless
the label changed.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017BMXdb2qZbPSh8Q7XpTyjB
pi's ExtensionSelectorComponent.updateList() renders every option with no
windowing, so 50 rows scrolled the transcript away. prime-agent's selector
windows to ~20 itself and shows a true (20/356) counter; matching that height
makes both hosts behave the same.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017BMXdb2qZbPSh8Q7XpTyjB
Two gaps the pi switcher had already closed.
**The indicator was invisible.** `ctx.ui.setStatus` was the wrong channel:
prime-agent stores extension statuses (`FooterDataProvider.setExtensionStatus`)
but nothing ever reads them back — there is no `getExtensionStatuses()` call
site anywhere in the app, so the value was recorded and never rendered. pi's
`footer.js` does render them, which is why the same code showed a status line
there and nothing here. The active project is now published as a *widget*
(`extensionWidgetsBelow` → `renderWidgets()`), which prime-agent does render.
`setStatus` is still called so pi keeps its footer entry.
**No search.** With 356 projects the picker was a wall of
`smoke-mcptoken-*`. Above 20 projects it now asks for a filter first;
space-separated terms must all match, case-insensitively, against name or
description, so `home auto` finds `homeautomation`. Active project sorts
first, then alphabetical.
Deliberately no client-side truncation here, unlike the pi picker: pi's
`ExtensionSelectorComponent.updateList()` renders every option, so a long list
fills the screen and has to be capped. prime-agent's selector windows the list
itself and shows a true `(20/356)` counter — capping would replace an accurate
total with a misleading one.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017BMXdb2qZbPSh8Q7XpTyjB
mcpd keeps `name` and `description` as columns, and a skill's `content` is
often just the body — `propose-learnings` starts straight at its `#` heading.
The sync wrote `full.content` verbatim, so the SKILL.md landed with no YAML
frontmatter and the host rejected it. Observed live:
[Skill warning]
~/.prime/agent/skills/propose-learnings/SKILL.md
description is required
Claude Code and pi require the same keys, so this affected every target; it
only surfaced now because prime-agent prints the warning at startup.
`ensureSkillFrontmatter` synthesises the block from the columns when it is
absent, and fills in only a missing `name`/`description` when the author
already supplied a header — an existing complete header is never rewritten.
Values are emitted as JSON strings (valid YAML double-quoted scalars) because
descriptions routinely contain `:` and `#`, which break a bare scalar.
Also: `--force` now re-fetches skills whose server content is unchanged.
Without it a skill already on disk could never be repaired by a client-side
fix — the content hash still matched, so it was skipped forever, which is
exactly what happened on the first attempt to repair the file above.
Three existing assertions compared SKILL.md to the raw server content; they
now assert the body survives rather than exact bytes.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017BMXdb2qZbPSh8Q7XpTyjB
Switching projects twice left the agent believing it had lost MCP access
entirely: it kept calling the previous project's tool names, got
"Tool mc_<old>_begin_session not found", and concluded no MCP tools existed.
Root cause is a pi constraint, not a bug in the switch. pi has no way to
unregister a tool — `registerTool` only ever does `extension.tools.set(name)`
— so the previous project's `mc_*` tools stay registered and merely go
inactive. `setActiveTools` correctly drops them from the live set, but the
conversation still contains the old project's tool listing, so the model
keeps calling names that now answer "not found".
Nothing was telling the model the tool set had changed. Now the switch
injects a custom message naming the active project, stating that previously
listed mcpctl tool names are dead, and listing what is actually callable.
Custom messages are converted to user-role messages by `convertToLlm`, so
they do reach the model (unlike `appendEntry`, which is explicitly excluded
from context). `display: false` keeps it out of the transcript — the
notification is what the human reads.
The gate is also called out explicitly, in both the notification and the
injected message. A freshly switched project gets a new mcp-session-id and is
therefore gated again, so "1 tool(s) ready" is correct but reads like a
failure; it now says which begin_session call unlocks the rest.
`toolChangeAnnouncement` is exported and unit-tested rather than left as
wording only reachable through a TUI.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017BMXdb2qZbPSh8Q7XpTyjB
pi's selector is a plain arrow-key list: `ExtensionUIDialogOptions` has no
search field and `ExtensionSelectorComponent` ignores typed characters, so
filtering has to happen before the list is handed over. With 356 projects —
most of them `smoke-proj-none-*` leftovers — arrowing to the one you want is
hopeless.
Above 20 projects the picker now asks for a filter first. Terms are
space-separated and all must match as case-insensitive substrings, so
`home auto` finds `homeautomation`. Blank shows everything, Esc cancels.
The active project sorts first (most likely pick), then alphabetical. A
result set over 50 is capped, and the title says what was dropped — a
silently truncated list reads as "that's all of them".
The ordering and matching are extracted into an exported `filterProjects` so
they are unit-tested rather than eyeballed through a TUI.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017BMXdb2qZbPSh8Q7XpTyjB
The pi extension shipped in `src/pi-ext/` was covered by no tsconfig and no
eslint config, so nothing ever checked it against pi's API. Pointing tsc at
the published @earendil-works/pi-coding-agent types found the command surface
to be inert.
Fixes:
- `/mcpctl` did nothing. `ctx.ui.select` takes `string[]` and returns the
chosen string; it was called with `{value,label}` objects, so the menu
rendered five `[object Object]` rows and `choice === "status"` never
matched any branch. Labels are now plain strings mapped back to actions.
- The headless branch returned a status string from a handler typed
`Promise<void>`; pi drops it. Reports via notify instead.
- "Sync skills" omitted `--agent pi`, writing into ~/.claude/skills — in an
integration whose stated purpose is to not depend on ~/.claude — and said
so in its own success message. It also ran execSync with `stdio: "inherit"`,
painting raw output over pi's TUI, and interpolated the project name into a
shell string. Now execFile with `--agent pi` and captured output.
- Tool results typed `content[].type` as `string`; pi's AgentToolResult wants
the `"text"` literal.
- `callTool` asserted `Promise<unknown>` to `ToolCallResult`.
- Sanitising MCP tool names to `[a-z0-9_]` can collide (`docs.search` vs
`docs-search`). The colliding tool was silently never registered but still
reported active, so its calls were forwarded to the first tool. Names are
now disambiguated and tracked with the MCP tool they forward to.
- `registerWithPi` rewrote settings.json even when nothing changed. Since
parsing strips `//` comments, a no-op run destroyed them.
Guards, so this class of bug can't return:
- `src/pi-ext/tsconfig.json` checks the extension against the real published
pi types (dev dependency, not a shim — a shim drifting from the published
API is the exact failure being guarded). Wired into `pnpm typecheck`.
- eslint now covers `src/pi-ext/*.ts` like every other source file.
- A test fails if the embedded copy in `config/pi-extension.ts` is stale;
editing the sources without regenerating silently shipped old code.
Also: the branch added `config pi` without regenerating shell completions
(the committed-completions test was failing), and the doc advertised
`mcpctl pi sync-skills`, which does not exist. Both corrected, plus a note
on the session-token vs `mcpctl_pat_` bearer difference that would bite
against an authenticated `mcplocal serve`.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017BMXdb2qZbPSh8Q7XpTyjB
- add scripts/generate-pi-extension.ts to embed src/pi-ext sources as
string constants (mirrors prime-agent's embedded switcher pattern).
- config pi now writes the embedded sources by default, so it works from
an /usr/bin install with no source tree (--extension-dir overrides for
dev). Verified installed binary writes files byte-identical to source.
- add installEmbeddedExtension test.
`/mcpctl` could switch projects but there was no way to see which one was
active without running a command. prime-agent exposes the same status-bar
channel the model name uses (`ctx.ui.setStatus`), so the switcher now
publishes `mcpctl:<project>` there.
Wired to `session_start`, which fires on startup *and* on every reload —
including the reload the switch itself triggers — so the footer tracks
settings.json without extra bookkeeping. Cleared when no project is mounted.
Also fixes `notify(..., 'success')`: the API only accepts info|warning|error.
Not a live bug (prime-agent falls through to the same showStatus path as
'info') but it fails a typecheck of the extension against the real
ExtensionAPI, which is how it was found.
The extension ships as a JSON-escaped string and is never compiled by our
build, so it was typechecked out-of-tree against
@earendil-works/pi-coding-agent's published types.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017BMXdb2qZbPSh8Q7XpTyjB
Round 2 fixed the first review but introduced regressions of its own, all
of which only bite against state written by the previously installed build.
`config prime-agent`:
- Mint each credential under a unique `prime-agent-<stamp>` name again.
`McpToken` is unique on (name, projectId) and revoke is a soft delete, so
round 2's fixed `prime-agent` name could only ever be minted once per
project — and the revoke-first ordering destroyed the working credential
before discovering the mint would fail.
- Provision the credential BEFORE touching settings.json. Registering the
new project unmounts the previously active one, so a failed mint must not
be able to leave prime-agent with no working project at all. The command
now aborts with settings.json untouched.
- Retire only the token this auth.json actually held, once its replacement
is stored. Sweeping every `prime-agent*` token for the project would
revoke the credential another install (or a custom --output run) is
using; anything else that looks orphaned is reported, not deleted.
- Validate a pre-existing credential instead of trusting its presence: a
revoked or expired token used to short-circuit provisioning and leave
prime-agent broken while the command reported success. Matched by
tokenPrefix against the project's active tokens, so the secret is never
sent. Fails open when the API can't be consulted.
- Actually write auth.json 0600. `writeFile`'s mode is ignored for an
existing file and prime-agent creates auth.json itself at 0644, so chmod
after writing.
- Recognise the untagged mcpServers entries older CLIs wrote (canonical
proxy URL + an `mcp:<name>` mcpctl PAT in auth.json) so a switch unmounts
them instead of leaving two gateways live. Hand-configured servers have
no such credential and are still preserved. Same rule in the `/mcpctl`
switcher's active-project lookup.
- Add `--skip-marker`, and pass it from the `/mcpctl` switcher: the
extension runs from whatever directory prime-agent was started in, and
was silently re-scoping that repo's `.mcpctl-project`.
`skills sync --agent prime-agent`:
- Record ownership from the skill's own scope, not the syncing project's.
Globals were being pinned to whichever project happened to sync them,
after which every other project refused to update them forever.
- Never adopt legacy, ownership-less state into the current scope. Round 2
did, which deleted the other project's skills on the first sync after
upgrading. Such entries are attributed to the project that last wrote the
state file, and left alone when that isn't the project syncing now.
- Close the overwrite-guard bypass: a sync with no project, or a global
landing on a project-owned name, could still clobber and re-own a
tracked skill.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017BMXdb2qZbPSh8Q7XpTyjB
Addresses the second round of `config prime-agent` review (10 findings).
auth.json (config/prime-agent.ts) — the settings.json data-loss fix had a twin:
- loadPrimeAgentAuth now fails loudly on corrupt JSON instead of swallow-and-
rewrite, so one syntax error can no longer destroy the provider API key and
every other project's credential. hasPrimeAgentAuth shares that guarantee.
- writePrimeAgentAuth writes 0600 (preserving an existing file's mode) instead
of the default umask — bearer tokens are no longer world-readable on first
creation.
state ownership (commands/skills.ts) — the ownership model edge cases:
- orphan-removal guard now normalises a canonical scope (project name, or null
for globals; legacy undefined adopted to current scope) instead of comparing
null against undefined, so global-only syncs and pre-PR state can no longer
leave stale skills on disk forever.
- a same-named skill *tracked* to a different project is preserved (with a
warning) rather than silently overwritten in the shared flat tree.
- mcpServers auto-attach is gated behind !isPrimeAgent with an explicit warning
(prime-agent's HTTP gateway must not mutate shared mcpd project attachments);
this also makes the earlier dropped-attach concern explicit rather than silent.
single-active project + switcher (config/prime-agent.ts, prime-agent-extension.ts):
- registerPrimeAgentMcp tags the project's entry mcpctlManaged:true and removes
other mcpctl-managed entries, so prime-agent has one *active* mcpctl project
while preserving untagged servers (hand-configured sre, websearch, etc).
- the /mcpctl extension now reads that tag as the single source of truth for the
active project, fixing the false short-circuit / no-op switch.
config.ts command:
- an explicit -p now updates a differing up-tree .mcpctl-project marker (scope
no longer silently reverts on the next sync), no-ops when it matches, and
still never scopes $HOME.
- a project left with no usable credential now exits non-zero (the /mcpctl
extension checks child exit status, so it no longer reports a successful
switch after provisioning failed).
- skills sync is treated as best-effort: settings+auth determine switch success,
so a skills error no longer falsely fails the switch.
- token minting now revokes prior active `prime-agent` tokens before creating a
fresh one (no more never-expiring token litter / lost-credential duplication).
Tests (544 green): corrupt auth.json refusal, 0600 mode, mint-failure exit
code, single-active dedup preserving untagged sre, cross-project overwrite
preservation, and global-orphan removal on global-only sync.
Addresses a review of the `config prime-agent` feature and adds the in-app
project switcher.
Safety/correctness fixes (prime-agent's shared, hand-editable ~/.prime/agent
tree must never suffer silent data loss):
- config/prime-agent.ts: loadPrimeAgentSettings now fails loudly on corrupt
JSON instead of swallowing it and rewriting the file (which destroyed every
non-mcpServers setting). A project's mcpServers entry is merged (keeping
user-added fields) rather than replaced wholesale. Added writePrimeAgentAuth
/ hasPrimeAgentAuth helpers for auth provisioning.
- skills sync: unified the near-verbatim prime-agent copy into runSkillsSync
via a `target: 'claude' | 'prime-agent'` option (prime-agent-skills.ts is now
a thin wrapper). Under the prime-agent target it: preserves untracked
pre-existing skill dirs on first sync (no more rm -rf of hand-authored `sre`),
records per-project ownership so configuring a second project never deletes
the first project's skills, skips Claude-only hooks/postInstall, and keeps
the mcpServers auto-attach step.
- config.ts: `config prime-agent` now (a) provisions the bearer credential in
auth.json (--token, existing entry, or auto-mint via POST /api/v1/mcptokens),
(b) writes the .mcpctl-project marker only when none exists up-tree and never
from $HOME, and (c) propagates the skills sync exit code so auth failures are
reported instead of swallowing them.
- skills.ts: `--agent` is validated; an unknown value errors instead of
silently running the Claude sync.
New feature: `config prime-agent` installs a `/mcpctl` project-switcher
extension into ~/.prime/agent/extensions/ (skip with --skip-extension). It lists
mcpctl projects via `mcpctl get projects -o json`, lets you pick one from the
prime-agent TUI, applies the switch through the CLI, and reloads the session.
Regenerated shell completions. Tests: 538 pass (new coverage for settings
corruption, entry merge, auth provisioning, extension install/skip, marker
$HOME handling, untracked/cross-project skill preservation, --agent validation).
Mirror `mcpctl config claude` for prime-agent (which talks to the same mcpctl
proxy MCP gateway over HTTP instead of stdio):
`mcpctl config prime-agent --project X`:
- registers the proxy MCP gateway in ~/.prime/agent/settings.json as
mcpServers.X = { type: "http", url: <gateway>/projects/X/mcp }, merging with
any existing servers (e.g. the bundled `sre` project) and preserving all other
settings
- writes a .mcpctl-project marker so later syncs resolve the project
- syncs the project's skills into ~/.prime/agent/skills/<name>/ as markdown
skills (prime-agent auto-discovers them at session start)
New `mcpctl skills sync --agent prime-agent` target re-syncs the tree later.
- src/cli/src/config/prime-agent.ts: settings.json read/merge/write helpers
- src/cli/src/utils/prime-agent-skills.ts: prime-agent sync (reuses
installSkillAtomic + skills-state; skips Claude-only hooks/postInstall)
- completes config.ts/skills.ts wiring; regenerated shell completions
- tests: commands/prime-agent.test.ts + utils/prime-agent-skills.test.ts
One slow call stalled every later request because the stdin loop awaited each in turn; queued requests were never sent so nothing could time them out and the client saw silence. Dispatch is now concurrent. Regression test verified to fail on the old code.
The bridge's stdin loop awaited every request before reading the next line, so
it handled exactly one at a time. A single slow call therefore stalled every
later request — and because those requests were never even sent, nothing could
time them out. The client saw silence, not an error.
Observed 2026-08-05: two gitea calls through this bridge sat completely mute
until Claude Code aborted them at its own 1800s idle limit, reporting "sent no
response or progress". The upstream was healthy the whole time — a fresh
session answered the same tools in 0.3s, and the gitea MCP server's own log
showed the calls never reached it. They died queued in the bridge.
JSON-RPC ids exist precisely so responses may return out of order, so nothing
here needed a queue. Requests now dispatch concurrently and are tracked in a
set; stdin close awaits them before the session DELETE, otherwise a concurrent
call races the teardown and dies with a 404.
We still serialise until the session id exists: it arrives on the first
response, and firing later requests without it would open a second upstream
session. A client sends `initialize` first and waits for the reply anyway, so
this costs one round trip rather than throughput.
Also makes the per-request timeout configurable via MCPCTL_MCP_TIMEOUT_MS
(default unchanged at 30s) and names it in the timeout error, so a project with
genuinely long tool calls can raise it instead of hitting a hardcoded wall.
Tests pin both halves: a fast request must overtake a slow one (this fails on
the old serial code — verified by reverting), and a failed request must still
produce a JSON-RPC error carrying the original id rather than nothing.
Even "<=8 words" was ignored — the model wrote a full-sentence reasoning, pushing
begin_session to ~7s (thin under the 8s budget). The gate only needs the names, so
request just {"selectedNames":[...]}. Validated live vs vllm-current
(glm-4.6-reap-fast): 1.5s, valid JSON. (reasoning defaults to '' in extractSelection.)
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
With the fast (no-think) route, gate-selection latency is now proportional to the
ANSWER length (a no-think model runs ~10-13 tok/s), not hidden thinking. A verbose
"reasoning" field pushed a selection to ~9s (over the 8s budget). Request compact
JSON with a <=8-word reasoning → ~2s (vs ~9s prose, or ~1.2s with no reasoning).
Validated live against vllm-fast (glm-4.6-reap-fast): 2.0s, valid selection JSON,
zero reasoning tokens.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Lets ops route gate prompt-selection at a fast (no-think) server Llm independent
of the project's chat llmProvider — so selection stays ~1-2s while chat keeps its
thinking model. Overrides the project llmProvider for the gate's server-selection
path; unset → prior behavior (use the project's llmProvider). Pairs with the
litellm qwen3-fast no-think alias (kubernetes-deployment) + a `vllm-fast` mcpd Llm.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Prompt selection is mechanical classification — it gains nothing from chain-of-
thought, and a reasoning model (qwen3-thinking) otherwise burns its whole budget
reasoning (~40s, 6k+ chars) and blows the gate's 8s timeout. The gate's server
selection request now includes thinking-suppression hints, forwarded verbatim by
mcpd's passthrough adapter to litellm/vLLM:
chat_template_kwargs.enable_thinking=false (Qwen3 hard-off)
reasoning_effort=low (o-series / newer vLLM)
Harmless for models that ignore them; if a backend rejects them the selector
falls back to the local provider. Override via MCPCTL_GATE_SELECT_EXTRA_BODY
(JSON), or '' to disable.
NOT yet validated live — the qwen3-thinking vLLM is crashlooping again (0/1,
HTTP 500), and whether litellm forwards chat_template_kwargs to vLLM is unconfirmed
(the /no_think prompt directive was NOT honored). Validate when the backend
recovers. mcpld gate/selector tests green (75).
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Cloud/server keys belong at the k8s/mcpd level; mcplocal handles only the user's
personal tokens. The gate's prompt-selection used mcplocal's LOCAL provider
registry (heavy=anthropic = the user's OAuth token, which is API-gated and 404s),
so selection always degraded.
Now LlmPromptSelector tries sources in priority order: (1) the project's server
Llm via mcpd's inference proxy (POST /api/v1/llms/:name/infer) — cloud/server
keys stay at k8s — then (2) the local personal-token provider as fallback. First
source that returns valid selection JSON wins. Extracted pickCompletionText
(content ?? reasoning_content) + extractSelection helpers.
Wiring: GatePluginConfig.llmProvider (from the project), threaded via
createDefaultPlugin; handleBeginSession builds serverInfer from ctx.postToMcpd.
Verified live: POST .../llms/vllm-current/infer returns valid selection JSON
(qwen3-thinking, 4000-token budget). + selector tests (server preferred, local
fallback on error, server-only). mcplocal suite green (755).
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
The local mcplocal gate resolves heavy=[anthropic]; its OAuth-gated key 404s the
configured opus-4 model, and anthropic.ts (like openai.ts before this series)
JSON.parsed the error body as a completion → empty content → the gate's
misleading "did not contain valid selection JSON". Reject on status >= 400 with
"Anthropic HTTP <status>: <body>" so the true reason (model-not-found / auth)
surfaces as the degradedReason. + parity test.
Note: this makes the error honest; it does not make the local gate SUCCEED — that
needs a working heavy provider (the anthropic OAuth token is fundamentally gated).
The qwen3-thinking gate path (k8s mcplocal) is addressed by the budget+reasoning
fix in the previous commit.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
The presented→canonical resolver was only built during tools/list, so a
favourite/<tool> call arriving before any list this session fell through to the
namespaced router and 404'd (all/ was fine — it prefix-strips). Real MCP clients
list first, but the smoke test's initialize→call flow exposed it. Extract
buildPresentation() and lazily rebuild the resolver in onToolCallBefore when a
favourite//all/ name isn't resolved yet. + unit test (favourite/ call with no
prior list) and realistic listTools() in the smoke routing case.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Two independent, live-reproduced causes of "Smart prompt-selection unavailable
(LLM response did not contain valid selection JSON)" in the gate:
- Cause B (openai.ts request): the response handler JSON.parsed the body with no
res.statusCode check, so a LiteLLM/vLLM HTTP 500 ("Connection error") parsed as
an empty completion and surfaced as generic "bad JSON", masking backend-down.
Now reject on status >= 400 with "OpenAI HTTP <status>: <body>" so the selector
reports the true degradedReason.
- Cause A (parseResponse + selector budget): a reasoning model (qwen3-thinking)
spends its whole 1024-token budget in reasoning_content and emits content=null,
so the {selectedNames} JSON never appears. parseResponse now falls back to
reasoning_content (incl. provider_specific_fields) when content is empty, and
the selector budget is raised 1024 -> 4000 (env MCPCTL_GATE_SELECT_MAX_TOKENS)
so thinking + the JSON answer both fit.
Note: mcpd's chat path got a reasoning_content fallback in PR #67; this is the
separate mcplocal gate-selection provider, which never did. Tests: openai-provider
(500 rejects, reasoning_content capture/preference). mcplocal suite green (749).
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Measured winner from the DGX-Spark bake-off (toolsim.py, 145-tool catalog): a
curated favourite/<tool> shortlist + the full all/<server>/<tool> catalog + a
load-bearing "prefer favourite/ first" instruction nearly halved wander (37→20)
and 2.5x'd first-pick (2→5/8) vs a flat catalog. The instruction is load-bearing;
enriching descriptions did not help.
- New mcplocal plugin `favourite-index.ts`: composes AFTER gate (no-ops while
gated), reshapes the ungated upstream catalog into favourite/ + all/, injects
the instruction (onInitialize), and rewrites presented names back to canonical
server/tool in onToolCallBefore so normal routing + content-pipeline still run.
Gate/agent virtual tools pass through untouched; favourites are upstream-only.
- compose.ts: onInitialize now concatenates plugin instructions (was first-non-null)
so favindex can contribute its banner alongside the gate's.
- Per-project config `Project.favouriteIndex` {enabled, tools[], maxFavourites};
surfaced to the proxy via discovery; wired at project-mcp-endpoint when enabled.
- Usage derivation: mcpd tool-usage ranking over tool_call_trace events
(normalizing presented names → canonical), GET /api/v1/audit/tool-usage, and
`mcpctl favourites suggest|list`.
- CLI: `create project` gains --favourite/--favourite-index/--max-favourites;
favouriteIndex round-trips through get -o yaml | apply -f. Completions regenerated.
- Tests: plugin unit (presentation, rewrite routing, gated no-op, collisions),
compose merge, canonicalizeToolName, buildFavouriteIndex, + a live smoke test.
- Docs: docs/tool-presentation.md.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Record the empirical finding: a Claude subscription OAuth token is gated by
Anthropic (404/429 as a raw API) and can't serve as the server-side fallback,
even though mcpd's adapter now sends it via Bearer. The anthropic-fallback is
wired-but-inert until mcpd Secret anthropic-key holds a real console API key.
Also note the chain is now Pulumi-owned so pulumi up won't wipe it.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude subscription / Claude Code keys are OAuth tokens and authenticate with
Authorization: Bearer, not x-api-key. mcpd's server adapter only sent x-api-key,
so registering an OAuth key as a fallback Llm 401'd ("invalid x-api-key").
Mirror the mcplocal client adapter's isOAuth switch so an OAuth key works as a
server-side anthropic fallback.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Chat is LLM-essential but not model-specific — instead of failing when the
pinned model is down/drifted, it now fails over across an ordered chain and
reports which model actually answered.
- Ordered fallback: an Llm declares `extraConfig.fallbacks: string[]`; the
dispatcher builds primary-pool → fallback-pool(s) candidates and tries them
in order (resolveCandidatesWithFallbacks).
- Fail over on real failures, not just transport: runOneInference now advances
on a non-2xx status (e.g. a 400 from a drifted model) or an empty/invalid
completion, not only thrown transport errors. Streaming fails over
pre-first-chunk (already threw on 4xx).
- Transparency: ChatResult + the SSE `final` frame carry {llm, model,
failedOver}; the CLI prints `model: <llm> (<model>)` each turn and
`⚠ failed over → answered by …` when a fallback was used.
- Exhaustion names the last model + upstream body (not "no choice").
Tests: 3 failover unit tests (primary 400 → fallback answers + model reported;
primary answers → failedOver=false; all fail → clear aggregated error).
mcpd 948 + CLI 508 green; tsc + lint clean. docs/reliability.md updated.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
mcpctl hung on prompt-reading whenever the LLM misbehaved (thinking model =
minutes; drifted model = silent failure). Root cause: the gate's begin_session
prompt-selection called the LLM with no timeout, its fallback only fired on
error and was silent, and it forced the project's vLLM model onto the anthropic
heavy provider (so selection failed silently every time).
- New withTimeout(run, ms, label): Promise.race + AbortSignal (fetch-based
providers cancel). CompletionOptions.signal threaded into anthropic/openai.
- Gate begin_session: LLM selection is time-bounded (MCPCTL_GATE_LLM_TIMEOUT_MS,
8s); on timeout/error it falls back to deterministic tag matching, logs
[gate] loudly, prepends a ⚠ degraded note to the response, and sets
degraded/degradedReason on the audit gate_decision.
- Gate selector no longer forces the project vLLM model — uses the heavy
provider's own model (fixes the always-silent-fail bug).
- Pagination smart-index is time-bounded too (falls back to byte-range pages).
- Chat (LLM-essential) surfaces the upstream status+body (names model+reason)
instead of "Adapter returned no choice".
- docs/reliability.md documents the principle.
Tests: with-timeout unit tests; gate degradation test (error → visible ⚠ +
deterministic prompts, no hang). mcplocal 737 + mcpd 945 green; tsc clean.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Chat directly with a Project (no Agent needed): its Prompts become the system
context, its MCP-server tools are callable, its llmProvider/llmModel drive the
LLM, and (opt-in) the model can read secret values. History is saved inside the
project, attributed per user, resumable, and deletable (RBAC-permitting) — "use
it like Claude, scoped to the project".
Backend (reuses the agent-chat orchestrator):
- ChatThread is now agent-XOR-project (schema + migration + CHECK constraint);
new listThreadsByProject / deleteThread on the repo.
- ChatService: prepareProjectContext (project prompt + Prompts by priority,
llm from llmProvider with llmModel override, project tools), shared
runChatLoop/runChatStreamLoop, project thread CRUD with owner enforcement
(404-not-403 on foreign threads), admin-override delete.
- Gated get_secret virtual tool: offered only with --allow-secrets AND the
caller's view:secrets; resolves via SecretService, never routes to a server.
- routes/project-chat.ts (chat SSE+non-stream, threads create/list/delete);
RBAC run:projects:<name>.
CLI:
- `mcpctl chat --project <name>` (+ --allow-secrets), one-shot/REPL/resume.
- REPL /threads, /resume <id>, /delete <id>; project-aware header + /tools.
- `mcpctl get threads --project <name>`, `mcpctl delete thread <id> --project`.
- completions regenerated (--project completes project names).
Tests: 8 new project-chat unit tests; full mcpd (945) + CLI (508) green;
schema validated against Postgres. Docs: docs/chat.md "Project chat" section.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>