NewYour coding agent can read the release notes before it upgrades.Set up the MCP server →
Go modules · #688 by repository stars
Last release 3 days ago
05 Oct 2026
Ships on a steady schedule
a new release about every 9 days
Nearly every release is documented
notes for 57 of the last 60 stable releases
Nothing withdrawn
no release was ever pulled
11 months old
927 releases · first in 2025
One column per month.
security(deps): bump pytest to >=9.0.3 (CVE-2025-71176) by @AbirAbbas in https://github.com/Agent-Field/agentfield/pull/488
⚠️ This is a staging/pre-release version for testing. Not recommended for production use.
# Staging binary (use --staging flag)
curl -fsSL https://agentfield.ai/install.sh | bash -s -- --staging
# Python SDK (prerelease - requires --pre flag)
pip install --pre agentfield
# TypeScript SDK
npm install @agentfield/sdk@next
VERSION=v0.1.140-rc.6 curl -fsSL https://agentfield.ai/install.sh | bash
Download the binary for your platform below, make it executable, and move it to your PATH.
agentfield-darwin-amd64agentfield-darwin-arm64agentfield-linux-amd64agentfield-linux-arm64af call rejecting valid input for optional reasoner params by @AbirAbbas in https://github.com/Agent-Field/agentfield/pull/610_current_status issue where status stuck on `S… by @DebanKsahu in https://github.com/Agent-Field/agentfield/pull/673af install by @AbirAbbas in https://github.com/Agent-Field/agentfield/pull/738Note truncated.
Restarts always mint a new execution id, so a consumer polling the id it originally received has no way to discover that the work continued somewhere else. Record the successor on the source row instead.
Adds restarted_as_execution_id to executions and workflow_executions (migration 036, mirrored in the SQLite models so AutoMigrate picks it up), a SetExecutionRestartedAs writer, and the execution-filter predicates the resume sweep needs: updated_at bounds, terminal-only, root-only and status-reason matching.
Also defines the agent_shutdown_cancelled status reason and the cross-SDK "cancelled during graceful shutdown" literal it keys off, so a reasoner cancelled by a draining pod stops being indistinguishable from a user cancel.
Co-Authored-By: Claude Fable 5.1 noreply@anthropic.com
Splits the restart mechanics out of the gin handler into startRestart so the same path can be driven without an HTTP request, keeping the endpoint's status codes, headers, body and queue-full behaviour byte-identical.
Every restart now stamps the forward pointer on the source execution and on the run root it restarted, and mirrors it into the source run's lineage metadata. GET /executions/{id} returns it as restarted_as, but only once the successor actually resolves, so a consumer is never handed a pointer into a run that has since been cleaned up. A healthy execution costs no extra query.
A terminal cancelled callback carrying the SDK graceful-shutdown error is stored as agent_shutdown_cancelled. The SDK's fire-and-forget lifecycle event stream is the path that actually terminalizes a drain-cancelled orchestrator in practice, so it classifies identically.
Co-Authored-By: Claude Fable 5.1 noreply@anthropic.com
A rolling deploy that kills a long-running orchestrator leaves its run terminally dead; the only recovery is the caller re-submitting. With AGENTFIELD_RESUME_INTERRUPTED_RUNS=true the control plane does that itself: an interrupted run root is handed to the existing workflow-scope restart with reuse all-succeeded, so children that already succeeded are replayed rather than dispatched again, and the dead execution is stamped with the forward pointer.
Interruption means agent_restart_orphaned, control_plane_shutdown or agent_shutdown_cancelled — nothing else is ever resumed. Any of the three can terminalize a run, so the handoff is triggered from the orphan reap, the status callback, the lifecycle-event stream, and a bounded sweep at startup for the case where the control plane itself restarted.
Bounded on purpose: run roots only, never an execution that already has a successor, chain depth capped by _MAX_ATTEMPTS, startup capped by _LIMIT and _WINDOW, and no storage query at all when disabled. Inline handoffs are scheduled after _DELAY on a detached context with every guard re-evaluated against a freshly read record, so the replacement pod has registered before the successor is dispatched and two triggers for the same root still produce one successor.
Co-Authored-By: Claude Fable 5.1 noreply@anthropic.com
An execution created while the old pod is gone and the replacement has not yet registered is stamped with the departing instance id, and the deferred reap then fails it. Measured end to end: the handoff's successor ran to completion on the new pod while the control plane had already marked it agent_restart_orphaned.
The reap now carries the instant the replacement registered and skips any row created at or after it — work created after a replacement announces itself cannot belong to the process that left. A zero cutoff behaves exactly as before, so existing callers are unaffected. Wall-clock based and therefore approximate under control-plane clock skew; the stale-execution sweep remains the backstop.
Co-Authored-By: Claude Fable 5.1 noreply@anthropic.com
The Python and TypeScript SDKs both report a reasoner cancelled by their own graceful shutdown as status cancelled with the error "cancelled during graceful shutdown". The Go SDK reported status failed with the raw "context canceled", so a Go agent's drain looked like an ordinary failure and the control plane could not tell it apart from a real error.
Only a cancellation caused by the agent's own shutdown is remapped; a caller cancellation or a timeout keeps its current behaviour.
Co-Authored-By: Claude Fable 5.1 noreply@anthropic.com
Documents the forward pointer on execution reads, the agent_shutdown_cancelled status reason, and the five AGENTFIELD_RESUME_INTERRUPTED_* knobs, in the same tables the existing drain and reap settings live in.
Says plainly that the orchestrator re-executes from its first line — there is no snapshot of interpreter state — and that what makes the restart cheap is replaying the already-succeeded app.call children.
Adds a Limitations section rather than overselling: single control plane, one attempt per interruption, stale-sweep timeouts are not resumed, the startup sweep is one bounded page per boot, the attempt counter is best-effort metadata, and the reap cutoff is wall-clock based.
Co-Authored-By: Claude Fable 5.1 noreply@anthropic.com
Covers the config-merge branches for the new resume settings, the restart endpoint's validation and load-failure paths, and the handoff's guards: each interruption reason and a rejected one, lineage attempt resolution when the store cannot read run metadata or the metadata is absent or malformed, and the early returns when the feature is off or the restart itself fails.
Co-Authored-By: Claude Fable 5.1 noreply@anthropic.com
Co-authored-by: Claude Fable 5.1 noreply@anthropic.com (ee7d65d)
security(deps): bump pytest to >=9.0.3 (CVE-2025-71176) by @AbirAbbas in https://github.com/Agent-Field/agentfield/pull/488
⚠️ This is a staging/pre-release version for testing. Not recommended for production use.
# Staging binary (use --staging flag)
curl -fsSL https://agentfield.ai/install.sh | bash -s -- --staging
# Python SDK (prerelease - requires --pre flag)
pip install --pre agentfield
# TypeScript SDK
npm install @agentfield/sdk@next
VERSION=v0.1.140-rc.5 curl -fsSL https://agentfield.ai/install.sh | bash
Download the binary for your platform below, make it executable, and move it to your PATH.
agentfield-darwin-amd64agentfield-darwin-arm64agentfield-linux-amd64agentfield-linux-arm64af call rejecting valid input for optional reasoner params by @AbirAbbas in https://github.com/Agent-Field/agentfield/pull/610_current_status issue where status stuck on `S… by @DebanKsahu in https://github.com/Agent-Field/agentfield/pull/673af install by @AbirAbbas in https://github.com/Agent-Field/agentfield/pull/738Note truncated.
An admitted async execution answered 202 with status "queued" but was persisted as "running" before a pool worker had picked it up, so GET /api/v1/executions/{id} disagreed with the body the caller had just been handed and reported work as running while it sat in the queue.
Persist the async lane as queued at admission and flip to running in the worker, immediately before the agent call. The flip is a conditional read-modify-write: it only acts on a row that is still queued, so it cannot clobber a cancel or pause that won the race, and it leaves the restart lane (already running) alone. It counts as transitioned only when the write came back persisted, so a storage failure does not emit execution.started or fake a running plan. The worker also stops before dispatching anything that reached a terminal status while it waited.
execution.started now fires at dispatch rather than at admission, and admission emits an execution.updated event carrying the queued status so event consumers still learn about the new run.
queued gains the destinations a running row has, because the workflow projection flip is best-effort and a row still marked queued must be able to reach every state its dispatched counterpart could — otherwise a partial flip strands it. Migration 036 teaches the Postgres workflow_executions CHECK about queued; executions already allowed it.
Admission itself is unchanged: the queue stays bounded and over-capacity requests still get 503 + Retry-After with nothing persisted.
Refs #986
Co-Authored-By: Claude Fable 5.1 noreply@anthropic.com
Two consumers assumed a just-admitted async execution was already running. Pause hardcoded running as its only source status, so pausing a run that had not been dispatched yet now returned 409; the enhanced dashboard built its active-run list from running and waiting records only, so an admitted run vanished from the dashboard until a worker picked it up.
Accept queued as a pausable source — the worker already waits for resume before dispatching a paused execution, so pausing before dispatch just means the agent is not called until resume — and include queued in the dashboard's active-run query.
Refs #986
Co-Authored-By: Claude Fable 5.1 noreply@anthropic.com
Say plainly that the 202 body and GET both report queued until a worker dispatches, that queued is non-terminal so callers poll through it, and that the queue is bounded rather than an unbounded backlog.
Refs #986
Co-Authored-By: Claude Fable 5.1 noreply@anthropic.com
The dispatch loop's post-transition branches had no coverage: an execution that reached a terminal status between the worker's read and its update, and one that was paused in that same window (both when the resume wait fails and when it succeeds and the loop retries). Also cover the error path of the dashboard's new queued-execution query.
Refs #986
Co-Authored-By: Claude Fable 5.1 noreply@anthropic.com
Co-authored-by: Claude Fable 5.1 noreply@anthropic.com (b30dd2f)
security(deps): bump pytest to >=9.0.3 (CVE-2025-71176) by @AbirAbbas in https://github.com/Agent-Field/agentfield/pull/488
⚠️ This is a staging/pre-release version for testing. Not recommended for production use.
# Staging binary (use --staging flag)
curl -fsSL https://agentfield.ai/install.sh | bash -s -- --staging
# Python SDK (prerelease - requires --pre flag)
pip install --pre agentfield
# TypeScript SDK
npm install @agentfield/sdk@next
VERSION=v0.1.140-rc.4 curl -fsSL https://agentfield.ai/install.sh | bash
Download the binary for your platform below, make it executable, and move it to your PATH.
agentfield-darwin-amd64agentfield-darwin-arm64agentfield-linux-amd64agentfield-linux-arm64af call rejecting valid input for optional reasoner params by @AbirAbbas in https://github.com/Agent-Field/agentfield/pull/610_current_status issue where status stuck on `S… by @DebanKsahu in https://github.com/Agent-Field/agentfield/pull/673af install by @AbirAbbas in https://github.com/Agent-Field/agentfield/pull/738Note truncated.
Closes #1059. The stale reaper could time out a running parent in the brief window (~50-600ms) between its child reaching a terminal state and the parent posting its own result. The parent's own updated_at does not move while it waits on the child, and the existing guard only skipped a parent while it had a non-terminal child. The moment the child posted succeeded, the guard vanished and nothing refreshed the parent's clock, so the reaper could mark the parent (and its workflow) timeout, and the parent's real success callback was then rejected with HTTP 409.
Fix (issue's Option 1, query-contained): in MarkStaleExecutions and MarkStaleWorkflowExecutions, also skip a parent when a child reached a terminal state after the cutoff (COALESCE(c.completed_at, c.updated_at)
cutoff). A child that finished long before the cutoff no longer shields the parent, so genuinely stuck parents are still reaped and orphan cleanup keeps working.
Timeout children are excluded from the shield (c.status != 'timeout'): the reaper's own kills set a recent completed_at/updated_at, and letting them shield would stall the bottom-up chain unwind. Applied to both the candidate SELECT and the re-evaluating UPDATE in each function.
Tests: recently-finished child shields the parent; long-finished child does not; reaper-timeout child does not; bottom-up unwind preserved. Same cases for MarkStaleWorkflowExecutions.
PR #1046 protected a workflow whose own clock advances; this covers the distinct case where the parent's clock never advances while it waits.
Note: RetryStaleWorkflowExecutions has the same guard shape and race; I kept this PR to the two functions the issue names and can extend to the retry path in a follow-up if desired.
Review by @AbirAbbas on #1063 surfaced three things:
completed_at is the agent's clock (UpdateExecutionStatusHandler stores req.CompletedAt verbatim), while the reaper cutoff is the control plane's time.Now(). A skewed agent clock or a late-landing callback could make completed_at older than the cutoff at the instant the CP wrote the terminal row, so the shield expired before it was installed and the parent was reaped anyway (the original 409). The shield now reads the LATER of completed_at and updated_at via a new childTerminalRecencyExpr (GREATEST on postgres, MAX(julianday(...)) on sqlite): updated_at is the CP's own write clock, so either clock being after the cutoff protects the parent. On the executions table a terminal row's updated_at is frozen (terminal->terminal is rejected), so this cannot shield indefinitely. Adds Abbas's deterministic TestMarkStaleExecutions_LateChildCallbackStillShieldsParent.
RetryStaleWorkflowExecutions runs before both reapers when max_retries > 0 and carried the identical guard, so the same race reset a live parent to pending mid-flight. Applied the same recent-terminal-child shield to its SELECT and re-evaluating UPDATE, backdated TestRetryStaleWorkflowExecutions_TerminalChildDoesNotShieldParent to the long-finished case, and added TestRetryStaleWorkflowExecutions_RecentlyFinishedChildShieldsParent.
Rewrote the stale MarkStaleExecutions doc comment: it claimed there was no child recency test and the chain unwound one sweep per level; both are now wrong. It documents the recency shield, the timeout-child exclusion, and that unwinding can take a stale window per level.
Co-authored-by: Abir Abbas abirabbas1998@gmail.com (2f41660)
security(deps): bump pytest to >=9.0.3 (CVE-2025-71176) by @AbirAbbas in https://github.com/Agent-Field/agentfield/pull/488
⚠️ This is a staging/pre-release version for testing. Not recommended for production use.
# Staging binary (use --staging flag)
curl -fsSL https://agentfield.ai/install.sh | bash -s -- --staging
# Python SDK (prerelease - requires --pre flag)
pip install --pre agentfield
# TypeScript SDK
npm install @agentfield/sdk@next
VERSION=v0.1.140-rc.3 curl -fsSL https://agentfield.ai/install.sh | bash
Download the binary for your platform below, make it executable, and move it to your PATH.
agentfield-darwin-amd64agentfield-darwin-arm64agentfield-linux-amd64agentfield-linux-arm64af call rejecting valid input for optional reasoner params by @AbirAbbas in https://github.com/Agent-Field/agentfield/pull/610_current_status issue where status stuck on `S… by @DebanKsahu in https://github.com/Agent-Field/agentfield/pull/673af install by @AbirAbbas in https://github.com/Agent-Field/agentfield/pull/738Note truncated.
A single daemon thread drains a bounded FIFO of (stream, line) pairs. Deferral only engages where it can help — the caller is on a running event loop and the destination has a real file descriptor — so synchronous callers and in-memory captures keep writing inline and stay immediately visible.
Under back-pressure the queue discards its oldest pending lines instead of blocking the producer, and the writer emits one log.dropped record per destination naming the count, so loss is never silent. An atexit drain bounded by AGENTFIELD_LOG_QUEUE_FLUSH_SECONDS covers normal shutdown, and an at-fork hook gives the child fresh state.
Refs #985
Co-Authored-By: Claude Opus 5 (1M context) noreply@anthropic.com
_emit_structured_record printed inline on the calling thread, which in an agent node is the event loop, into the _TeeTextIO wrapper over a real pipe. A consumer that stops draining that pipe froze the whole loop for the duration — heartbeats, in-flight reasoners and the control-plane client alike. Size bounding did not help; the stall is the pipe, not the payload.
Both stdout paths now go through the bounded writer, which also keeps their relative order. _DynamicStdoutHandler additionally drops its logging handler lock: Handler.handle() holds that lock across emit(), so a synchronous thread blocked inline on a stalled stdout would otherwise still stall an event-loop caller before it reached the queue. The writer serializes instead.
Refs #985
Co-Authored-By: Claude Opus 5 (1M context) noreply@anthropic.com
_TeeTextIO holds its write lock across the blocking write to the original stream. A fork while another thread holds it leaves the child with a mutex that is locked and has no owner, so every later write in the child deadlocks and its stdout is gone for good.
An after_in_child hook now gives each installed tee a fresh lock, clears the partial line inherited mid-write, re-creates the ring and follower locks and drops the parent's follower queues. There is deliberately no before= hook: acquiring those locks ahead of the fork would make fork() itself wait on a stalled pipe.
The hazard predates this branch, but a dedicated writer thread that can hold the lock for the length of a consumer stall makes it easy to hit.
Refs #985
Co-Authored-By: Claude Opus 5 (1M context) noreply@anthropic.com
AGENTFIELD_LOG_QUEUE, AGENTFIELD_LOG_QUEUE_SIZE and AGENTFIELD_LOG_QUEUE_FLUSH_SECONDS, alongside the existing logging variables: when deferral engages, the drop-oldest policy and its log.dropped marker, that os._exit bypasses the exit flush, and that ordering against a caller's own print() is no longer guaranteed.
Refs #985
Co-Authored-By: Claude Opus 5 (1M context) noreply@anthropic.com
Setting logging.Handler.lock to None relied on handle() calling
self.acquire(), which skips a falsy lock. Python 3.13 changed handle() to
with self.lock:, so None raises TypeError: 'NoneType' object does not
support the context manager protocol — every plain log line on 3.13 died in
the handler, and a test whose synchronous writer thread was killed that way
hung CI rather than failing.
Use a no-op lock object instead: it satisfies both the acquire/release and the context-manager protocols, and keeps the property we want, which is that nothing serializes on the handler while the bounded writer already does.
Refs #985
Co-Authored-By: Claude Opus 5 (1M context) noreply@anthropic.com
These helpers block on a pipe read or a lock on purpose. As non-daemon threads, a failing assertion left them alive and the interpreter waited for them at exit, so a test failure presented as a hung CI job instead of a failure. As daemons the session reports the failure and exits.
Refs #985
Co-Authored-By: Claude Opus 5 (1M context) noreply@anthropic.com
Co-authored-by: Claude Opus 5 (1M context) noreply@anthropic.com (00c8844)
A Pydantic AI Agent runs inside an AgentField reasoner with no adapter, but wiring the observability so one Logfire trace covers both the Pydantic AI run and the app.ai completion is not obvious. This example does it: logfire.configure() once per process, instrument_pydantic_ai() + instrument_litellm() on that provider, and instrument_fastapi() on the Agent itself (it is a FastAPI subclass) so the inbound reasoner request is the root span.
The AgentField run identifiers ride along as OpenTelemetry baggage, which Logfire copies onto every descendant span, so a trace joins back to a run without threading anything through either library's API. Pydantic AI's own token usage and cost are folded into the per-execution cost tracker so they reach AgentField's usage accounting alongside app.ai's.
Co-Authored-By: Claude Opus 5 (1M context) noreply@anthropic.com
Issue #997 asked for Pydantic AI and Logfire integration. The Logfire half of it needs no AgentField code — what was missing is a written-down recipe for the combination, and the honest boundary around it.
This documents running a Pydantic AI agent inside a reasoner with one Logfire trace per execution covering both the Pydantic AI spans and the app.ai completions, correlated to the AgentField run through OpenTelemetry baggage; how to fold Pydantic AI's tokens into AgentField's per-execution usage; how the recipe differs from AGENTFIELD_LITELLM_CALLBACKS=logfire; and what it does not give you — Pydantic AI's durable-execution adapters are not wired into AgentField's run DAG, pause/resume or replay.
Co-Authored-By: Claude Opus 5 (1M context) noreply@anthropic.com
Co-authored-by: Claude Opus 5 (1M context) noreply@anthropic.com (d310f75)
security(deps): bump pytest to >=9.0.3 (CVE-2025-71176) by @AbirAbbas in https://github.com/Agent-Field/agentfield/pull/488
⚠️ This is a staging/pre-release version for testing. Not recommended for production use.
# Staging binary (use --staging flag)
curl -fsSL https://agentfield.ai/install.sh | bash -s -- --staging
# Python SDK (prerelease - requires --pre flag)
pip install --pre agentfield
# TypeScript SDK
npm install @agentfield/sdk@next
VERSION=v0.1.140-rc.2 curl -fsSL https://agentfield.ai/install.sh | bash
Download the binary for your platform below, make it executable, and move it to your PATH.
agentfield-darwin-amd64agentfield-darwin-arm64agentfield-linux-amd64agentfield-linux-arm64af call rejecting valid input for optional reasoner params by @AbirAbbas in https://github.com/Agent-Field/agentfield/pull/610_current_status issue where status stuck on `S… by @DebanKsahu in https://github.com/Agent-Field/agentfield/pull/673af install by @AbirAbbas in https://github.com/Agent-Field/agentfield/pull/738Note truncated.
The Go OpenCode provider folded the harness system prompt into the positional
opencode run prompt and passed no agent, so a run depended on whatever agent
and permissions the ambient OpenCode config happened to define. Python moved off
that in #1023; this brings Go to the same behaviour.
Each run now selects the fixed agentfield-harness agent with --agent and
supplies it through OPENCODE_CONFIG_CONTENT: system prompt, model,
reasoningEffort, mode primary, a fixed steps budget, and a headless permission
baseline that denies question, task and the agentfield* skills so an
AgentField-launched worker cannot dispatch back into the control plane.
The overlay is deep-merged into the caller's per-call value, or the ambient one when there is no per-call value, so a deployment's mcp servers, plugins, providers and other agents survive and the OpenRouter attribution overlay and the harness agent coexist. A caller value that is not a JSON object fails the run before the concurrency slot is taken and before the child is launched.
AGENTFIELD_OPENCODE_INLINE_SYSTEM_PROMPT restores the inline prompt transport
and strips the agent's configured prompt, keeping the agent selection and
permissions, so a caller with a very long system prompt can roll back without
pinning an older SDK. tools and permission_mode stay untranslated: with the
wildcard allow in place a tool mapping would only write allow on top of allow.
steps is a named constant with an AGENTFIELD_OPENCODE_STEPS override and is
never fed from max_turns.
Refs #960
Co-Authored-By: Claude Opus 5 (1M context) noreply@anthropic.com
Mirrors the Go and Python OpenCode providers: each run selects the fixed
agentfield-harness agent with --agent and defines it through a per-run
OPENCODE_CONFIG_CONTENT overlay (system prompt, model, reasoningEffort, mode
primary, fixed steps, and the headless permission baseline that denies
question, task and the agentfield* skills). The task is the only
positional prompt.
The overlay is deep-merged into the caller's per-call value or the ambient one, so deployment-owned mcp servers, plugins and agents survive and the OpenRouter attribution overlay is no longer the only thing that can occupy the variable. Object key order is preserved deliberately, with the wildcard first and AgentField's denials last, because OpenCode applies the last matching rule. A caller value that is not a JSON object throws before runCli is called.
AGENTFIELD_OPENCODE_INLINE_SYSTEM_PROMPT restores the inline prompt transport,
and tools / permission_mode remain accepted but untranslated.
This also fixes a key-name bug the overlay would otherwise inherit: the
provider read options.system_prompt, but HarnessRunner forwards HarnessOptions
verbatim, so a system prompt set through the public TypeScript API arrived as
systemPrompt and never reached opencode at all. It now accepts both spellings,
as the aforge provider already does.
Refs #960
Co-Authored-By: Claude Opus 5 (1M context) noreply@anthropic.com
The OpenCode agent overlay hard-coded steps: 500 as a bare literal. Give it a
name and an AGENTFIELD_OPENCODE_STEPS override (per-call environment first, then
ambient; non-numeric, zero and negative values fall back to the default), so all
three SDKs expose the same knob. The default is unchanged and max_turns is
still never serialized as steps.
Refs #960
Co-Authored-By: Claude Opus 5 (1M context) noreply@anthropic.com
The standalone-runs section was written when only Python had the per-run agent overlay, and the provider-parity table still said OpenCode receives only the model, directory and prompt. Both are now true of Go and TypeScript too.
Also documents AGENTFIELD_OPENCODE_STEPS and states plainly that tools and
permission_mode are accepted and ignored, rather than leaving readers to infer
they are wired to OpenCode permissions.
Refs #960
Co-Authored-By: Claude Opus 5 (1M context) noreply@anthropic.com
OpenCode applies the last matching permission rule, so the Python and
TypeScript providers deliberately emit AgentField's skill/question/task denials
after any caller rules. Go relied on encoding/json sorting map keys, which is
only accidentally correct: a caller supplying
{"agent":{"agentfield-harness":{"permission":{"skill":{"agentfield-*":"allow"}}}}}
serialized as {"agentfield*":"deny","agentfield-*":"allow"} — * sorts before
- — so the caller's allow was the last match and the recursion guard was off.
Serialize the permission object through a small ordered JSON type instead, in
the same order Python uses: wildcard, caller rules, then AgentField's denials,
with agentfield* last inside skill. Both the merged and the generated-only
paths now go through it, so one mechanism governs the order.
Also sizes the deep-merge map from the base alone; summing both lengths is what CodeQL's allocation-size-overflow rule flags.
The two new tests assert on the serialized JSON rather than a decoded map, because a decoded map cannot express order; both fail against the previous implementation.
Refs #960
Co-Authored-By: Claude Opus 5 (1M context) noreply@anthropic.com
HarnessRunner forwards HarnessOptions verbatim, so a caller sets
permissionMode and maxTurns; the test only passed the snake_case aliases, so
"tools and permission_mode add nothing to the overlay" was not actually checked
against the keys the public API sends. Pass both spellings.
Also scopes the Windows stdin sentence in the harness docs: Python and Go send the prompt over stdin there, the TypeScript adapter always uses the positional argument. That difference is pre-existing and stays out of this change.
Refs #960
Co-Authored-By: Claude Opus 5 (1M context) noreply@anthropic.com
Co-authored-by: Claude Opus 5 (1M context) noreply@anthropic.com (f7bae4b)
Your coding agent can read these notes before it upgrades. Set up the MCP server →