# Web search and documentation lookup Three templates. All self-hosted, none needs an API key or a vendor account, and all three are deployed by mcpctl like any other server. | Template | Package / image | Needs | |---|---|---| | `duckduckgo` | `duckduckgo-mcp-server` (python) | nothing | | `searxng` | `mcp-searxng` (node) | a SearXNG instance | | `docs-mcp` | `ghcr.io/arabold/docs-mcp-server` | nothing (see the persistence caveat) | Search and docs are different jobs, not competing options. A search engine will hand you a 2023 blog post with a stale method signature; a docs index cannot tell you why a daemon is crash-looping. Attach both to a project that does real engineering work. `docs-mcp` is the open-source replacement for Context7 / Ref.tools — same job, but the index lives on your infrastructure and can include private repos. ## Start here: `duckduckgo` The only one with no infrastructure behind it. It scrapes DuckDuckGo's HTML endpoint directly, so there is no engine to run and no key to hold. ```bash mcpctl create server websearch --from-template duckduckgo --env DDG_SAFE_SEARCH=OFF mcpctl get instances | grep websearch # RUNNING / healthy within ~a minute ``` Tools: `search` (`query`, `max_results`, `region`) and `fetch_content` (`url`, `start_index`, `max_length`) for pulling a result as markdown. The tradeoff is honest: scraping has no SLA. DuckDuckGo can change its markup or rate-limit you, and the server caps itself at 30 searches/min. Set `DDG_SEARCH_BACKEND=curl` if bot checks start biting. When it becomes a problem, move to `searxng`. ## Better results: `searxng` SearXNG aggregates ~25 engines and gives you real filtering, which is what keeps search results from flooding a context window: - `min_score` (0.0–1.0) — relevance floor. The most useful knob; start at `0.3`. - `time_range` — `day` / `week` / `month` / `year`. Essential for "did this break in the last release" questions. - `language`, `safesearch`, `pageno`. `web_url_read` then takes `section`, `paragraphRange`, `startChar`/`maxLength` and `readHeadings`, so you pull one section of a long page rather than all of it. Bounded PDF extraction is included, which covers most vendor documentation. The cost is that SearXNG is a service you have to run. It is plain infrastructure, not an MCP server, so mcpctl has nothing to manage it with — it belongs in your cluster or in the compose stack: ```bash cd stack cp .env.example .env # set SEARXNG_SECRET docker compose --profile websearch up -d ``` Then point the MCP server at it: ```bash mcpctl create secret searxng-conf --data SEARXNG_URL=http://mcpctl-searxng:8080 mcpctl create server searxng --from-template searxng --env-from-secret searxng-conf ``` ### The one SearXNG gotcha SearXNG ships with `search.formats: [html]`. Every `format=json` request against a stock instance returns **403**, and `mcp-searxng` comes back empty with no useful error. `stack/searxng/settings.yml` exists only to add `json` to that list — it is the one setting with no environment-variable override. This is also why pointing the template at a *public* SearXNG instance usually fails: nearly all of them leave the JSON API off. The compose healthcheck probes `format=json` specifically, so a misconfigured instance shows up as unhealthy rather than as silently empty search results. ## Documentation: `docs-mcp` ```bash mcpctl create server docs --from-template docs-mcp ``` Tools: `scrape_docs` (index a site, GitHub repo, npm or PyPI package, or local files), `search_docs` (query, optionally pinned to a version), and `fetch_url`. Index the things you actually run — Pulumi, the Kubernetes API, Grafana, the Terraform provider docs — rather than everything. ### Persistence The scraped index is a SQLite file (`better-sqlite3` + `sqlite-vec`) under `/data`. There is **no external-database mode** — no `DATABASE_URL`, no pgvector — so a Postgres cluster cannot help here. The template declares a volume instead: ```yaml volumes: - name: data mountPath: /data sizeGb: 20 ``` The backing PVC is named after the *server* (`mcpctl-docs-data`), not the instance, so editing the server or restarting the pod re-attaches to the same index rather than starting empty. See "Persistent volumes" in the README. The template carries no `healthCheck`, because `search_docs` needs a library argument that only exists after a scrape — a synthetic probe would report unhealthy on a fresh instance. ### Embeddings Keyword search works out of the box. Semantic search is noticeably better and needs an embedding model. To keep it local, point the server at Ollama: ```bash ollama pull nomic-embed-text mcpctl create server docs --from-template docs-mcp \ --env DOCS_MCP_EMBEDDING_MODEL=openai:nomic-embed-text \ --env OPENAI_API_BASE=http://ollama:11434/v1 \ --env OPENAI_API_KEY=ollama ``` Ollama's OpenAI-compatible endpoint is why the provider prefix is `openai:` — the key is a placeholder and never leaves the network. Changing the embedding model invalidates the index: embeddings from different models are not comparable, so everything has to be re-scraped. ## Alternatives considered - **Tavily / Brave / Exa / Perplexity** — better formatted results, all require an API key and send every query to a vendor. - **Context7** — the popular docs MCP, but a cloud index of public-docs snippets only. `docs-mcp` covers the same ground locally and takes private sources. - **agent-search** — bundles SearXNG and a 10-strategy extraction cascade in one deploy. Attractive, but at ~67 stars it is too young to build the stack on. Worth revisiting. - **Writing our own** — the hard parts are extraction heuristics and the embedding pipeline, and all three projects above already solved them under MIT.