# Conflicts: # completions/mcpctl.bash # completions/mcpctl.fish # src/cli/src/commands/create.ts # src/db/src/seed/index.ts
149 lines
5.9 KiB
Markdown
149 lines
5.9 KiB
Markdown
# Web search and documentation lookup
|
||
|
||
Three templates. All self-hosted, none needs an API key or a vendor account, and
|
||
all three are deployed by mcpctl like any other server.
|
||
|
||
| Template | Package / image | Needs |
|
||
|---|---|---|
|
||
| `duckduckgo` | `duckduckgo-mcp-server` (python) | nothing |
|
||
| `searxng` | `mcp-searxng` (node) | a SearXNG instance |
|
||
| `docs-mcp` | `ghcr.io/arabold/docs-mcp-server` | nothing |
|
||
|
||
Search and docs are different jobs, not competing options. A search engine will
|
||
hand you a 2023 blog post with a stale method signature; a docs index cannot
|
||
tell you why a daemon is crash-looping. Attach both to a project that does real
|
||
engineering work.
|
||
|
||
`docs-mcp` is the open-source replacement for Context7 / Ref.tools — same job,
|
||
but the index lives on your infrastructure and can include private repos.
|
||
|
||
## Start here: `duckduckgo`
|
||
|
||
The only one with no infrastructure behind it. It scrapes DuckDuckGo's HTML
|
||
endpoint directly, so there is no engine to run and no key to hold.
|
||
|
||
```bash
|
||
mcpctl create server websearch --from-template duckduckgo --env DDG_SAFE_SEARCH=OFF
|
||
mcpctl get instances | grep websearch # RUNNING / healthy within ~a minute
|
||
```
|
||
|
||
Tools: `search` (`query`, `max_results`, `region`) and `fetch_content` (`url`,
|
||
`start_index`, `max_length`) for pulling a result as markdown.
|
||
|
||
The tradeoff is honest: scraping has no SLA. DuckDuckGo can change its markup or
|
||
rate-limit you, and the server caps itself at 30 searches/min. Set
|
||
`DDG_SEARCH_BACKEND=curl` if bot checks start biting. When it becomes a problem,
|
||
move to `searxng`.
|
||
|
||
## Better results: `searxng`
|
||
|
||
SearXNG aggregates ~25 engines and gives you real filtering, which is what keeps
|
||
search results from flooding a context window:
|
||
|
||
- `min_score` (0.0–1.0) — relevance floor. The most useful knob; start at `0.3`.
|
||
- `time_range` — `day` / `week` / `month` / `year`. Essential for "did this break
|
||
in the last release" questions.
|
||
- `language`, `safesearch`, `pageno`.
|
||
|
||
`web_url_read` then takes `section`, `paragraphRange`, `startChar`/`maxLength`
|
||
and `readHeadings`, so you pull one section of a long page rather than all of it.
|
||
Bounded PDF extraction is included, which covers most vendor documentation.
|
||
|
||
The cost is that SearXNG is a service you have to run. It is plain
|
||
infrastructure, not an MCP server, so mcpctl has nothing to manage it with — it
|
||
belongs in your cluster or in the compose stack:
|
||
|
||
```bash
|
||
cd stack
|
||
cp .env.example .env # set SEARXNG_SECRET
|
||
docker compose --profile websearch up -d
|
||
```
|
||
|
||
Then point the MCP server at it:
|
||
|
||
```bash
|
||
mcpctl create secret searxng-conf --data SEARXNG_URL=http://mcpctl-searxng:8080
|
||
mcpctl create server searxng --from-template searxng --env-from-secret searxng-conf
|
||
```
|
||
|
||
### The one SearXNG gotcha
|
||
|
||
SearXNG ships with `search.formats: [html]`. Every `format=json` request against
|
||
a stock instance returns **403**, and `mcp-searxng` comes back empty with no
|
||
useful error. `stack/searxng/settings.yml` exists only to add `json` to that
|
||
list — it is the one setting with no environment-variable override.
|
||
|
||
This is also why pointing the template at a *public* SearXNG instance usually
|
||
fails: nearly all of them leave the JSON API off.
|
||
|
||
The compose healthcheck probes `format=json` specifically, so a misconfigured
|
||
instance shows up as unhealthy rather than as silently empty search results.
|
||
|
||
## Documentation: `docs-mcp`
|
||
|
||
```bash
|
||
mcpctl create server docs --from-template docs-mcp
|
||
```
|
||
|
||
Tools: `scrape_docs` (index a site, GitHub repo, npm or PyPI package, or local
|
||
files), `search_docs` (query, optionally pinned to a version), and `fetch_url`.
|
||
|
||
Index the things you actually run — Pulumi, the Kubernetes API, Grafana, the
|
||
Terraform provider docs — rather than everything.
|
||
|
||
### Persistence
|
||
|
||
The scraped index is a SQLite file (`better-sqlite3` + `sqlite-vec`) under
|
||
`/data`. There is **no external-database mode** — no `DATABASE_URL`, no
|
||
pgvector — so a Postgres cluster cannot help here. The template declares a
|
||
volume instead:
|
||
|
||
```yaml
|
||
volumes:
|
||
- name: data
|
||
mountPath: /data
|
||
sizeGb: 20
|
||
```
|
||
|
||
The backing PVC is named after the *server* (`mcpctl-docs-data`), not the
|
||
instance, so editing the server or restarting the pod re-attaches to the same
|
||
index rather than starting empty. See "Persistent volumes" in the README.
|
||
|
||
The readiness probe is `list_libraries`: no arguments, read-only, and it reads
|
||
the SQLite store, so a pass proves the volume is mounted and readable rather
|
||
than merely that the process started. It answers "No libraries indexed yet." on
|
||
a fresh instance instead of erroring, which is why `search_docs` (which needs a
|
||
library that only exists after a scrape) cannot serve as the probe.
|
||
|
||
### Embeddings
|
||
|
||
Keyword search works out of the box. Semantic search is noticeably better and
|
||
needs an embedding model. To keep it local, point the server at Ollama:
|
||
|
||
```bash
|
||
ollama pull nomic-embed-text
|
||
mcpctl create server docs --from-template docs-mcp \
|
||
--env DOCS_MCP_EMBEDDING_MODEL=openai:nomic-embed-text \
|
||
--env OPENAI_API_BASE=http://ollama:11434/v1 \
|
||
--env OPENAI_API_KEY=ollama
|
||
```
|
||
|
||
Ollama's OpenAI-compatible endpoint is why the provider prefix is `openai:` —
|
||
the key is a placeholder and never leaves the network.
|
||
|
||
Changing the embedding model invalidates the index: embeddings from different
|
||
models are not comparable, so everything has to be re-scraped.
|
||
|
||
## Alternatives considered
|
||
|
||
- **Tavily / Brave / Exa / Perplexity** — better formatted results, all require
|
||
an API key and send every query to a vendor.
|
||
- **Context7** — the popular docs MCP, but a cloud index of public-docs snippets
|
||
only. `docs-mcp` covers the same ground locally and takes private sources.
|
||
- **agent-search** — bundles SearXNG and a 10-strategy extraction cascade in one
|
||
deploy. Attractive, but at ~67 stars it is too young to build the stack on.
|
||
Worth revisiting.
|
||
- **Writing our own** — the hard parts are extraction heuristics and the
|
||
embedding pipeline, and all three projects above already solved them under
|
||
MIT.
|