feat(gate): ask reasoning models for a fast, no-think selection response
Prompt selection is mechanical classification — it gains nothing from chain-of- thought, and a reasoning model (qwen3-thinking) otherwise burns its whole budget reasoning (~40s, 6k+ chars) and blows the gate's 8s timeout. The gate's server selection request now includes thinking-suppression hints, forwarded verbatim by mcpd's passthrough adapter to litellm/vLLM: chat_template_kwargs.enable_thinking=false (Qwen3 hard-off) reasoning_effort=low (o-series / newer vLLM) Harmless for models that ignore them; if a backend rejects them the selector falls back to the local provider. Override via MCPCTL_GATE_SELECT_EXTRA_BODY (JSON), or '' to disable. NOT yet validated live — the qwen3-thinking vLLM is crashlooping again (0/1, HTTP 500), and whether litellm forwards chat_template_kwargs to vLLM is unconfirmed (the /no_think prompt directive was NOT honored). Validate when the backend recovers. mcpld gate/selector tests green (75). Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
This commit is contained in:
@@ -21,6 +21,26 @@ import { withTimeout, TimeoutError } from '../../util/with-timeout.js';
|
||||
* begin_session — on timeout we fall back to deterministic tag matching. */
|
||||
const GATE_LLM_TIMEOUT_MS = Number(process.env['MCPCTL_GATE_LLM_TIMEOUT_MS'] ?? '8000');
|
||||
|
||||
/**
|
||||
* Extra request fields for the gate's server-Llm selection call. Prompt
|
||||
* selection is mechanical classification that gains nothing from chain-of-
|
||||
* thought, so we ask reasoning models to skip thinking (they otherwise burn the
|
||||
* whole budget reasoning and blow the gate timeout). Sent verbatim by mcpd's
|
||||
* passthrough adapter → litellm/vLLM. Harmless for models that ignore them, and
|
||||
* if a backend rejects them the selector falls back to the local provider.
|
||||
* - `chat_template_kwargs.enable_thinking:false` — Qwen3 hard-off.
|
||||
* - `reasoning_effort:'low'` — OpenAI o-series / newer vLLM.
|
||||
* Override with a JSON object in MCPCTL_GATE_SELECT_EXTRA_BODY, or '' to disable.
|
||||
*/
|
||||
const GATE_SELECT_EXTRA_BODY: Record<string, unknown> = (() => {
|
||||
const raw = process.env['MCPCTL_GATE_SELECT_EXTRA_BODY'];
|
||||
if (raw === '') return {};
|
||||
if (raw !== undefined) {
|
||||
try { return JSON.parse(raw) as Record<string, unknown>; } catch { /* use default */ }
|
||||
}
|
||||
return { chat_template_kwargs: { enable_thinking: false }, reasoning_effort: 'low' };
|
||||
})();
|
||||
|
||||
export interface GatePluginConfig {
|
||||
gated?: boolean;
|
||||
providerRegistry?: ProviderRegistry | null;
|
||||
@@ -308,6 +328,7 @@ async function handleBeginSession(
|
||||
temperature: o.temperature,
|
||||
max_tokens: o.maxTokens,
|
||||
stream: false,
|
||||
...GATE_SELECT_EXTRA_BODY, // ask reasoning models for a fast, no-think answer
|
||||
});
|
||||
return pickCompletionText(resp);
|
||||
}
|
||||
|
||||
Reference in New Issue
Block a user