Two of these four assertions currently FAIL against the live cluster, which
is the point: mcpd there still renders `Secrets: bao-k8s* ✓` from the
rotation field, and its JSON status carries no live/ready pair. Per the
project rule the fix is to deploy, not to relax the assertion.
The status check specifically rejects a bare "name ✓" with no qualifier —
that string is the old rendering and proves the probe was never consulted.
Deliberately does not take the real OpenBao down. Simulating an outage
against shared infrastructure to satisfy a test would be worse than the
bug it covers; the outage paths are unit-tested with an injected clock.
Docs: adds a Reliability section (request hardening, the stale-while-error
table, the cold-cache gap and why persisting values is not the answer,
live-vs-ready) and replaces the stale "Kubernetes ServiceAccount auth is
not shipped yet" note — it shipped in 5152066, and it is what the live
backend has used since June.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_018vybEitX4FykeMatKe5Xki
10 KiB
Secret backends
mcpctl stores the raw data for Secret resources in a pluggable backend.
The default is plaintext — the secret payload lives in Postgres as plain JSON
— which is fine for laptop development but a poor fit for shared clusters. For
production, point at an external KV store and delete secrets from the DB after
migration.
This guide covers the model, the shipped drivers, and how to migrate without downtime.
Model
- A
SecretBackendresource is a single named driver instance (e.g. a pointer at one OpenBao deployment). - Every
Secretrow carries abackendIdFK — the backend that owns its data. - Exactly one
SecretBackendhasisDefault: true. New secrets created through the API/CLI land on that backend. - The
plaintextbackend is seeded at startup and nameddefault. It cannot be deleted — there needs to always be one row where the driver's own credentials can bootstrap from (see below).
CLI
mcpctl get secretbackends # list backends
mcpctl describe secretbackend <name> # inspect config (credentials masked)
mcpctl create secretbackend <name> --type plaintext [--default] [--description ...]
mcpctl create secretbackend <name> --type openbao \
--url http://bao.example:8200 \
--token-secret bao-creds/token \
[--namespace <ns>] [--mount secret] [--path-prefix mcpctl] \
[--default]
mcpctl delete secretbackend <name> # blocked if any secret still points at it
mcpctl migrate secrets --from default --to bao
mcpctl migrate secrets --from default --to bao --names a,b --keep-source
mcpctl migrate secrets --from default --to bao --dry-run
Anything you can do with create secretbackend also works via apply -f:
kind: secretbackend
name: bao
type: openbao
description: "shared cluster OpenBao"
isDefault: true
config:
url: http://bao.svc.cluster.local:8200
tokenSecretRef: { name: bao-creds, key: token }
namespace: platform
Drivers
plaintext
Trivial. Secret.data holds the JSON, externalRef is empty.
- Storage: Postgres column.
- Bootstrap: seeded as
defaultat startup. - Cost: zero setup, zero encryption at rest, full access for any DB reader.
Use for development, CI, or single-tenant self-hosts where the DB itself is treated as sensitive.
openbao
Talks HTTP to an OpenBao (MPL 2.0 Vault fork) KV v2 mount. Also compatible with HashiCorp Vault KV v2 — the wire protocol is the same.
| Config key | Required? | Description |
|---|---|---|
url |
yes | Base URL, e.g. http://bao.svc.cluster.local:8200. |
tokenSecretRef |
yes | { name, key } pointing at a Secret on the plaintext backend that holds the bootstrap token. |
mount |
no | KV v2 mount name. Default secret. |
pathPrefix |
no | Path prefix under the mount. Default mcpctl. Secrets land at <mount>/<pathPrefix>/<secretName>. |
namespace |
no | X-Vault-Namespace header for OpenBao/Vault Enterprise namespaces. |
The driver only stores a reference in Secret.externalRef (mount/path). The
Secret.data column is left empty for openbao-backed rows — you can safely
drop DB-level access to secrets after migration.
Required OpenBao policy
Minimum token policy for a backend that lives at secret/mcpctl/:
path "secret/data/mcpctl/*" {
capabilities = ["create", "read", "update"]
}
path "secret/metadata/mcpctl/*" {
capabilities = ["list", "delete"]
}
path "secret/metadata/mcpctl/" {
capabilities = ["list"]
}
Grant delete on metadata/... only if you need mcpctl to fully remove
secrets — OpenBao soft-deletes until the metadata is gone.
Chicken-and-egg: where does the OpenBao token live?
mcpd reads the OpenBao token from a Secret on the plaintext backend.
That's the whole point of keeping plaintext around — it's the trust root:
- Operator creates a plaintext
Secretholding the bootstrap token. - Operator creates the
openbaobackend, pointing at that secret viatokenSecretRef. - Operator runs
mcpctl migrate secrets --from default --to baoto move all other secrets off plaintext. - After migration, the only sensitive row left on plaintext is the OpenBao token itself. DB access is now equivalent to OpenBao token access (a single key), not equivalent to all API keys in the system.
Kubernetes ServiceAccount auth (no bootstrap token)
auth: kubernetes removes the chicken-and-egg entirely: mcpd exchanges its
projected ServiceAccount JWT for an OpenBao token at
auth/<authMount>/role/<role>, so there is no static credential in the database
at all. The token is cached for its lease and re-minted lazily with a 60s grace
window.
kind: secretbackend
name: bao-k8s
type: openbao
isDefault: true
config:
url: https://bao.example
auth: kubernetes
role: mcpctl
authMount: kubernetes-worker0 # defaults to `kubernetes`
Note that the daily rotator does not apply to these backends — there is no stored token to rotate. That has a consequence for monitoring, see below.
Reliability
Remote backends are network dependencies on the critical path of nearly everything: server env resolution, LLM api keys, chat, git providers, code repos, webhooks. Three mechanisms keep an outage from cascading.
Request hardening
Every call carries a timeout (default 5s) and retries 5xx/429/network
failures with full-jitter exponential backoff (3 attempts). A sealed OpenBao
answers 503, so this covers unseal windows and failovers.
The 403 path is separate and deliberately single-shot: the driver purges its
cached token, re-authenticates and retries once. That is a credential
refresh, not a backend-unavailable condition — looping on it would hide a
genuinely revoked grant.
Value cache with stale-while-error
Resolved values are cached per backend (default TTL 5 minutes, LRU-bounded). Past the TTL the backend is always consulted; if it fails as a transport failure, the last known-good value is served instead of throwing.
| Failure | Behaviour |
|---|---|
| Backend unreachable / timeout / exhausted 5xx | Serve last known-good, mark degraded, log BACKEND_UNREACHABLE once |
| Secret deleted (404) | Evict and throw. Never served stale — that would resurrect a revoked credential |
| 403 after a token refresh | Throw. Revoked grants must stay loud |
| Nothing cached yet | Throw |
The stale window is unbounded on purpose: a cap would mean a long outage eventually takes mcpd down anyway.
plaintext backends are not cached — their read() is an identity function
over the row the caller already supplied.
Cold cache is the known gap. If mcpd restarts while the backend is
unreachable, nothing has a last-known-good value and secret-bearing servers fail
to start. That is deliberate: booting a server with an empty credential is worse
(gitea-mcp once ran for weeks with an empty GITEA_ACCESS_TOKEN, answering
tools/list and reporting healthy while every authenticated call failed). mcpd
mitigates it by warming the cache at boot — one read per referenced secret — so
an outage that starts after startup is fully absorbed.
Health: live vs ready
curl $MCPD/api/v1/secretbackends/<id>/health
{ "live": true, "liveDetail": "active",
"ready": false, "readyDetail": "OpenBao list: HTTP 403 permission denied",
"cache": { "entries": 9, "servingStale": 0 },
"rotation": { "rotatable": false, "lastRotationError": null } }
live— unauthenticatedsys/health. Distinguishes down from sealed from standby.ready— a real read with our credentials.
The two are separate because live && !ready is a distinct, important state: a
re-initialised OpenBao hands back valid-looking tokens that grant nothing.
Collapsing them into one boolean is what let that go unnoticed for four days.
mcpctl status renders the probe directly:
Secrets: bao-k8s* ✓ reachable, default ✓ reachable
Secrets: bao-k8s* ⚠ degraded — serving 7 cached secret(s)
Secrets: bao-k8s* ✗ unreachable: sealed
Secrets: bao-k8s* ✗ auth failed: HTTP 403 permission denied
Secrets: bao-k8s* ? unknown
Historical note. This verdict used to come solely from
tokenMeta.lastRotationError, which only the rotator writes — and the rotator skipsauth: kubernetesbackends. The Secrets line was therefore incapable of going red for a k8s-auth backend, and reported OpenBao healthy while it was unreachable.?(probe failed) renders yellow, never green: not knowing is not health.
Migration — mcpctl migrate secrets
Atomicity is per secret, not per batch. Remote writes can't roll back, so we don't pretend. For each secret the service:
- Reads the plaintext from the source driver.
- Writes it to the destination driver.
- Updates the
Secretrow: flipsbackendId, sets newexternalRef, clearsdata. - Deletes from source (skipped with
--keep-source).
If the command is interrupted between step 2 and 3, the destination has an orphan entry but the source still owns the row. Re-running is idempotent — the service skips secrets that are already on the destination and picks up the rest.
# Dry-run first: see what would move.
mcpctl migrate secrets --from default --to bao --dry-run
# Migrate everything.
mcpctl migrate secrets --from default --to bao
# Migrate a subset only.
mcpctl migrate secrets --from default --to bao --names api-keys,oauth-client
# Leave the source copy in place (useful for A/B validation).
mcpctl migrate secrets --from default --to bao --keep-source
The command prints a per-secret summary (migrated / skipped / failed) and exits non-zero if any secret failed. Ctrl-C during the run is safe — restart when you want, no duplicate writes.
RBAC
resource: secretbackends— gated like any other resource (view,create,edit,delete).role: run, action: migrate-secrets— required to callPOST /api/v1/secrets/migrate.
Describe output masks config values whose keys look like credentials
(token, secret, password, key), so mcpctl describe secretbackend is
safe to paste into tickets.