NewYour coding agent can read the release notes before it upgrades.Set up the MCP server →
Go modules · #633 by repository stars
Last release today
01 Oct 2026
Ships on a steady schedule
a new release about every 9 days
Nearly every release is documented
notes for 57 of the last 60 stable releases
Nothing withdrawn
no release was ever pulled
11 months old
923 releases · first in 2025
One column per month.
> ⚠️ This is a staging/pre-release version for testing. Not recommended for production use.
⚠️ This is a staging/pre-release version for testing. Not recommended for production use.
# Staging binary (use --staging flag)
curl -fsSL https://agentfield.ai/install.sh | bash -s -- --staging
# Python SDK (prerelease - requires --pre flag)
pip install --pre agentfield
# TypeScript SDK
npm install @agentfield/sdk@next
VERSION=v0.1.85-rc.12 curl -fsSL https://agentfield.ai/install.sh | bash
Download the binary for your platform below, make it executable, and move it to your PATH.
agentfield-darwin-amd64agentfield-darwin-arm64agentfield-linux-amd64agentfield-linux-arm64Full Changelog: https://github.com/Agent-Field/agentfield/compare/v0.1.85-rc.11...v0.1.85-rc.12
Keep generated output behind the intended trust boundary while preserving the normal safe workflow.
Add regression coverage for the unsafe flow and the expected safe behavior.
Add coverage for generated tool target handling
test: avoid parallel port contention in package runner tests
Stabilize control-plane coverage tests
Co-authored-by: Santosh kumar 29346072+santoshkumarradha@users.noreply.github.com Co-authored-by: Abir Abbas abirabbas1998@gmail.com (e96c3fd)
> ⚠️ This is a staging/pre-release version for testing. Not recommended for production use.
⚠️ This is a staging/pre-release version for testing. Not recommended for production use.
# Staging binary (use --staging flag)
curl -fsSL https://agentfield.ai/install.sh | bash -s -- --staging
# Python SDK (prerelease - requires --pre flag)
pip install --pre agentfield
# TypeScript SDK
npm install @agentfield/sdk@next
VERSION=v0.1.85-rc.11 curl -fsSL https://agentfield.ai/install.sh | bash
Download the binary for your platform below, make it executable, and move it to your PATH.
agentfield-darwin-amd64agentfield-darwin-arm64agentfield-linux-amd64agentfield-linux-arm64Full Changelog: https://github.com/Agent-Field/agentfield/compare/v0.1.85-rc.10...v0.1.85-rc.11
Addresses two informational findings from PR review.
agent_ai.ai(): reject timeout <= 0 with ValueError instead of letting asyncio.wait_for raise a confusing immediate TimeoutError. Defensive guard against user error; cheap one-line check.
vision.generate_image_openrouter: add comment explaining that the falsy check for image_config in the strip-and-retry branch is intentional. image_config={} (user explicitly opting into provider defaults per the existing docstring) produces an identical wire call whether stripped or not, so skipping the useless extra attempt is the right behavior. Comment prevents a future "fix" that would just burn cycles.
One new test: test_ai_rejects_non_positive_timeout covers both timeout=0 and timeout=-1.0. (fc445e9)
Fix(sdk): per-call ai() timeout + OpenRouter image retry/strip fallback
Two SDK gaps surfaced by the reel-af example project under a real URL-to-vertical-reel workload. Both forced consumer-side workarounds that this PR removes.
vision.generate_image_openrouter — retry on routing 404 OpenRouter's "No endpoints found that support the requested output modalities" surfaces in two flavours: 1. Deterministic 404 when image_config (e.g. aspect_ratio=9:16) hits a model whose upstream replicas don't expose that param. 2. Intermittent 1-3% 404 under load when routing momentarily lands on a replica without image modality. Now wraps the litellm.acompletion call in a 3-attempt backoff (1s, 2s) and, if image_config was set and all retries failed, makes one final attempt with image_config stripped, logging a warning. Other exceptions still propagate immediately. The existing asyncio.wait_for timeout still wraps every attempt. Worst-case wait: 3s without image_config, 7s with.
agent_ai.ai() — per-call timeout override ai() now accepts timeout: Optional[float] = None. When set, it overrides async_config.llm_call_timeout for that single call, propagating to both litellm_params["timeout"] (httpx socket-level) and the asyncio.wait_for safety net (2x). Previously the only knob was the agent-wide config, which forced every call in a mixed fast/slow pipeline to use the slowest expected timeout. Wired through all three call paths: direct, tool-loop _make_call, and non-tool _make_litellm_call.
Tests
> ⚠️ This is a staging/pre-release version for testing. Not recommended for production use.
⚠️ This is a staging/pre-release version for testing. Not recommended for production use.
# Staging binary (use --staging flag)
curl -fsSL https://agentfield.ai/install.sh | bash -s -- --staging
# Python SDK (prerelease - requires --pre flag)
pip install --pre agentfield
# TypeScript SDK
npm install @agentfield/sdk@next
VERSION=v0.1.85-rc.10 curl -fsSL https://agentfield.ai/install.sh | bash
Download the binary for your platform below, make it executable, and move it to your PATH.
agentfield-darwin-amd64agentfield-darwin-arm64agentfield-linux-amd64agentfield-linux-arm64Full Changelog: https://github.com/Agent-Field/agentfield/compare/v0.1.85-rc.9...v0.1.85-rc.10
.github/ contained two PR-template files differing only by case:
Git tracks both as separate blobs because it's case-sensitive, but case-insensitive filesystems (macOS/Windows, default) can only materialize one of them. Symptoms on those platforms:
git status reports a permanent phantom "modified" file that
cannot be cleaned; git checkout just flips which casing is dirty.Linux/CI is unaffected (both files materialize fine), which is why this slipped past review when #368 landed.
Keep the lowercase, detailed template; remove the uppercase stub.
Co-authored-by: Claude Opus 4.7 (1M context) noreply@anthropic.com (3c750ac)
Replace the plain bulleted Learn More list with a card grid mirroring the Built With AgentField pattern: five blog posts (The AI Backend, IAM for AI Backends, and the three-part harness orchestration series) as clickable image cards with title + abstract + CTA. Docs links move to a Documentation subsection under the same heading to avoid duplicate-H2 anchors.
Commits hero images locally under assets/blog/ and adds the three new harness posts to assets/utm-links.csv (existing blog/IAM UTM ids reused).
Co-Authored-By: Claude Opus 4.7 (1M context) noreply@anthropic.com
Move The AI Backend and IAM for AI Backends to the bottom of the grid so the three-part harness orchestration series leads.
Co-Authored-By: Claude Opus 4.7 (1M context) noreply@anthropic.com
Add a one-line description above the blog card grid to soften the transition from heading to table and clarify the section's purpose.
Co-Authored-By: Claude Opus 4.7 (1M context) noreply@anthropic.com
Co-authored-by: OG oktaygoktas@users.noreply.github.com Co-authored-by: Claude Opus 4.7 (1M context) noreply@anthropic.com Co-authored-by: Santosh kumar 29346072+santoshkumarradha@users.noreply.github.com (155c2f3)
When root_error_category is something other than the canonical seven
slugs (e.g. a diagnostic message like "Agent Restart Orphaned: Tier2-Test
Re-Registered With New Instance (was 38ffec8279...)"), the previous
fallback echoed the raw string verbatim as the badge label. In a fixed-
width status cell that meant multi-line wrapping that overflowed into
adjacent rows.
The fallback now produces a fixed "Unknown error" label whenever the raw
category exceeds 24 chars, and surfaces the original string via a new
tooltip field on ExecutionErrorCategoryMeta (description for known
categories). Consumers can attach this to title= so hover still gives
full detail.
Adds four unit tests pinning the new behavior — nullish handling, known canonical category, short unknown category, and long free-form category.
Co-Authored-By: Claude Opus 4.7 (1M context) noreply@anthropic.com
Both the Runs table status cell (w-44 = 11rem) and the StepDetail error box render the error-category badge with a fixed h-5 height. When the label runs long, the previous flex-wrap layout caused the badge to wrap across multiple lines and visually overflow into the row below — see screenshot on PR #N.
Make the badge a single-line truncated pill in both places:
Combined with the parent commit's label cap, even a pathological 100+
char root_error_category now renders as a single ellipsised pill that
fits inside the 11rem column.
Co-Authored-By: Claude Opus 4.7 (1M context) noreply@anthropic.com
Co-authored-by: Claude Opus 4.7 (1M context) noreply@anthropic.com (208205e)
Nothing published for this version
Fix: handle missing function.arguments in execute_tool_call_loop
When an LLM returns a tool call without the arguments field, json.loads(None) raises TypeError which wasn't caught. Now checks for None first and reports the error back to the LLM so it can retry, instead of crashing the loop.
Closes #353 (f21289a)
Chore(deps): bump go.opentelemetry.io/otel/exporters/otlp/otlptrace/otlptracehttp from 1.32.0 to 1.43.0 in /control-plane in the go_modules group acro
Bumps the go_modules group with 1 update in the /control-plane directory: go.opentelemetry.io/otel/exporters/otlp/otlptrace/otlptracehttp.
Updates go.opentelemetry.io/otel/exporters/otlp/otlptrace/otlptracehttp from 1.32.0 to 1.43.0
updated-dependencies:
Signed-off-by: dependabot[bot] support@github.com
otel v1.42+ requires Go 1.25, but the control-plane module and CI/Docker
still target Go 1.24. Pin the otel family (otel, sdk, trace, metric,
otlptrace, otlptracehttp) to v1.41.0 — the highest release line that still
declares go 1.24.0 in its go.mod — and revert the go directive back to
1.24.0 so go.mod stays consistent with the rest of the repo.
Accepts the otel v1.43.0 upgrade to fix two Dependabot security alerts:
otel v1.43.0 requires Go 1.25, so bump the toolchain across the repo: control-plane/go.mod, all four Dockerfiles, and the three GitHub workflows (control-plane, coverage, release). SDK (sdk/go) is left on Go 1.21 since it does not depend on otel.
Signed-off-by: dependabot[bot] support@github.com Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com> Co-authored-by: Santosh santosh@agentfield.ai (656c46a)
Fix: block execution of agents in pending_approval lifecycle state
The ExecuteReasonerHandler and ExecuteSkillHandler did not check for pending_approval status, allowing agents awaiting tag approval to be invoked. The permission middleware had this check but only runs when DID authorization is enabled — leaving a gap in the default path.
Adds pending_approval guards to both handlers and excludes pending_approval agents from the versioned-agent fallback selection.
Co-authored-by: Claude Opus 4.6 (1M context) noreply@anthropic.com (cd375ec)
Fix: implement runs search filter and reject empty agent node IDs
The workflow-runs list endpoint silently ignored the search query
parameter sent by the frontend, so typing in the search bar had no
visible effect (TC-04-f). Add search support: parse the param in the
handler, thread it through ExecutionFilter, and apply LIKE matching
against run_id, agent_node_id, and reasoner_id in QueryRunSummaries.
Agents could also register with an empty-string node ID, which
rendered as a blank row in the UI (TC-05-b). Add a validate:"required,min=1"
tag on AgentNode.ID to reject empty IDs at registration, filter out
existing empty-ID records in GetNodesSummary, and show "(unknown)" as
a frontend fallback.
Co-Authored-By: Claude Opus 4.6 (1M context) noreply@anthropic.com
RegisterServerlessAgentHandler was missing the validate.Struct() call that RegisterNodeHandler has, allowing empty/whitespace-only node IDs to bypass validation through the serverless registration path.
Co-Authored-By: Claude Opus 4.6 (1M context) noreply@anthropic.com
Adds TestQueryRunSummariesSearchFilter exercising the new LIKE filter across run_id, agent_node_id and reasoner_id, plus a no-match case.
Co-authored-by: Claude Opus 4.6 (1M context) noreply@anthropic.com Co-authored-by: Santosh santosh@agentfield.ai (9138475)
Fix(web-ui): remove orphan test files blocking the release pipeline
PR #350 (Chore/UI audit phase1 quick wins) deleted HealthBadge.tsx and NodeDetailPage.tsx but left behind two test files that import them:
control-plane/web/client/src/test/components/GeneralComponents.test.tsx control-plane/web/client/src/test/pages/NodeDetailPage.test.tsx
Both fail at TypeScript compile with:
TS2307: Cannot find module '@/components/HealthBadge' TS2307: Cannot find module '@/pages/NodeDetailPage'
The tsc -b && vite build step inside the Release workflow (goreleaser
pre-build hook) fails on these errors, which means every release since
v0.1.65-rc.12 has failed — including the commits that cut rc.13 and rc.14
tags. No binaries have been published past rc.12. The staging install
has been stuck on rc.12 for everyone running curl … | bash -s -- --staging.
The tests cannot run at all (imports fail at TS compile time) so they are providing zero coverage. Removing them is safe and restores the release pipeline.
This unblocks v0.1.65-rc.15 which will include:
Co-Authored-By: Claude Opus 4.6 (1M context) noreply@anthropic.com (0be47a3)
Chore(skill): bump agentfield-multi-reasoner-builder to v0.3.0
The previous commit (1f0c8872) rewrote the skill content (meta-philosophy rewrite, async curl, post-mortem hardenings, memory surface, cross-boundary data rules, mandatory live smoke test) but forgot to bump the SkillVersion constant in catalog.go, so the binary kept reporting v0.2.0 even with the new embedded content.
Bumps agentfield-multi-reasoner-builder Version from 0.2.0 -> 0.3.0 and refreshes the Description to match the rewritten skill. The version-store symlink logic in skillkit/install.go will now recognize the new content as a new version on re-install:
~/.agentfield/skills/agentfield-multi-reasoner-builder/ ├── current → ./0.3.0/ (atomic swap on install) ├── 0.2.0/ (previous version, kept as rollback) └── 0.3.0/ (new version with rewritten content)
Users with an existing install get the new content via: af skill update # preferred — re-installs into all # currently-installed targets at # the binary's embedded version OR af skill install --all --force # explicit full re-install
Users installing fresh from the next staging release pull 0.3.0 automatically.
Co-Authored-By: Claude Opus 4.6 (1M context) noreply@anthropic.com (6a41ddf)
The SSE handler in memory_events.go defers header flushing until the first matching event is written, and uses the deprecated CloseNotify() for client…
chore: remove internal test coverage audit doc
test(control-plane): add tests for services, middleware, and config overlay
Adds white-box unit tests for previously-untested control plane files:
services/
server/
server/middleware/
Also adds .plandb.db to .gitignore.
reasoners.go (~700 LOC, previously untested):
memory_events.go (WS + SSE memory subscriptions):
agent/registration_integration_test.go
agent/verification_test.go (LocalVerifier)
agent/memory_backend_test.go (ControlPlaneMemoryBackend)
harness/provider_error_integration_test.go
Python SDK has good happy-path coverage; these add failure-mode tests:
test_did_manager_error_paths.py
test_vc_generator_error_paths.py
test_tool_calling_error_paths.py
test_agent_graceful_shutdown.py
Five subtests are intentionally skipped with 'source bug:' markers documenting real defects discovered while writing the tests (Agent.stop() unimplemented, graceful_shutdown does not track in-flight tasks, etc.). These are targets for follow-up fixes in the implementation, not test bugs.
The TS SDK had only ~6 real test files for ~50 source files. This adds behavior tests for the most-critical surfaces:
Core client
Features
42 tests across 8 suites, all green via vitest.
The SSE handler in memory_events.go defers header flushing until the first matching event is written, and uses the deprecated CloseNotify() for client disconnect detection. Both behaviors interact poorly with httptest in CI: http.Client.Do blocks until the handler writes, and the test never completes within the CI test deadline.
The other tests in this file (WS happy path, invalid-pattern cleanup, backpressure disconnect, upgrade rejection) already cover the same code paths, so skipping just this one is a clean win.
Tracked source fix: #358
The earlier deadlock was fixed in 7c81c534 by running the request in a goroutine and publishing after the subscription registers. The follow-up skip in c8992cd9 was redundant — the restructured test passes locally and in CI. Source-side flush refactor still tracked in #358.
Add coverage reporting workflow and fix local test entrypoints
test(web-ui): expand client coverage
test(sdk): raise python and typescript coverage
test(sdk-go): raise go coverage above 80
test(control-plane): cover storage vector and config paths
test(web-ui): cover vc service flows
fix(ci): stabilize coverage branch test runs
test(web-ui): expand api service coverage
test(control-plane): cover dashboard helper paths
test(control-plane): cover execution helper paths
test(web-ui): extract workflow dag utilities
test(web-ui): extract runs page utilities
test(web-ui): expand service coverage wave
test(control-plane): cover storage helper utilities
test(coverage): cover settings runs and workflow helpers
test(coverage): cover comparison page and execution records helpers
test(control-plane): cover execution log handlers
test(coverage): cover agent pages and identity handlers
test(web-ui): cover node detail page flows
test(web-ui): rebase client tests onto main
test(web-ui): align node detail mocks with current types (cf922f9)
Chore/UI audit phase1 quick wins
Replaces the toast-only notification system with a dual-mode center:
Cross-cutting UX pass addressing multiple issues from rapid review:
Backend
Unified status primitives (web)
icon: LucideIcon and
motion: "none" | "live". Single source of truth for glyph and motion
per canonical status.Runs table liveness
root_execution_status ?? status. Paused/cancelled rows stop
ticking immediately even when aggregate stays running.Run detail page
Dashboard
Notification center — compact tree, semantic icons
to={/runs/${runId}}
directly (no href string manipulation) so basename=/ui prepends
properly. Sonner toast action builds the URL from VITE_BASE_PATH.Top bar + layout
Dependencies
useFocusManagement ran on every route change and explicitly blurred the currently focused element "to prevent blue outline" — which is precisely the WCAG 2.4.7 Focus Visible indicator. It also moved focus to document.body (non-focusable) on initial load and popstate, killing route-change announcements for screen readers.
Deleting the hook and its call site in App.tsx restores:
No replacement scroll behavior added — browser default scroll restoration on SPA navigation is fine.
Refs: UI audit C1 (CRITICAL).
Co-Authored-By: Claude Opus 4.6 (1M context) noreply@anthropic.com
Adds a WCAG 2.4.1 "Bypass Blocks" skip link as the first child of SidebarInset. The anchor is visually hidden until focused (Tab from the URL bar), then appears in the top-left corner and jumps focus to #main-content when activated.
Previously a keyboard user had to Tab through ~14 sidebar + header items on every navigation before reaching page content. This fix makes that a single Tab + Enter.
Used an inline <a> rather than importing SkipLink from AccessibilityEnhancements.tsx because that file is dead code (unreachable from App.tsx) and will be deleted in a later pass.
Refs: UI audit C2 (CRITICAL).
Co-Authored-By: Claude Opus 4.6 (1M context) noreply@anthropic.com
Two changes in CommandPalette.tsx:
Placeholder was "Search pages, runs, agents..." but the component only navigates to 10 hardcoded items — it has no search over runs or agents. Changed to "Jump to page..." to match reality. A real async search over runs/agents/reasoners is deferred to a later pass.
Action label drift: "Show failed runs" and "Show running executions" coexisted in the same file, both navigating to /runs?status=*. The product has committed to "runs" as the canonical noun (/executions redirects to /runs in App.tsx). Changed "Show running executions" to "Show active runs".
Refs: UI audit C3 (CRITICAL).
Co-Authored-By: Claude Opus 4.6 (1M context) noreply@anthropic.com
The routeNames map in AppLayout.tsx kept entries for /nodes, /reasoners, /executions, and /workflows — all of which are legacy paths that App.tsx redirects to canonical routes (/agents, /runs). Because routeNames still contained them and longestSectionPath() picks the longest matching prefix, visiting /executions/xyz (e.g. from an old bookmark) would briefly render "Executions > Xyz" in the breadcrumb for ~100ms before the Navigate redirect fired.
Stale vocabulary flashing into the UI on every legacy link is a wayfinding bug. Delete the four dead keys; the redirects in App.tsx still handle the URLs.
Refs: UI audit H3 (HIGH).
Co-Authored-By: Claude Opus 4.6 (1M context) noreply@anthropic.com
refactor(ui): delete orphan EnhancedDashboardPage and /dashboard/legacy route
feat(nav): promote Playground to sidebar position 2
Developer persona is priority 1 — Playground should sit directly under Dashboard, above operator-centric items like Runs and Agent nodes. Addresses audit finding H1.
Rows are interactive containers, not links — they have no href and navigate via onClick. Removing role="link" prevents screen readers from announcing them as links. tabIndex=0 is preserved so keyboard users can still focus rows and activate via Enter/Space. Addresses audit finding M8.
size-7 (28px) was below the 36×36 minimum. size-9 brings it to 36×36, matching other icon buttons in the app. Placeholder span is updated to match so column alignment is preserved. Addresses audit finding M11.
Animation durations currently drift between 150ms / 200ms / 300ms across components. Exposing a 'base' token lets callers write duration-base instead of picking a number from the air. Addresses audit finding M14.
Weight 300 is not referenced anywhere — this drops one font file from the critical path. Full self-hosting via @fontsource/inter is tracked separately. Addresses part of audit finding M19.
The 30s auto-reset masked real bugs (error appears, disappears, user can't reproduce) and created flaky retry loops in tests. The manual "Try again" button still works. Addresses audit finding H13.
Three webhook/DLQ actions (remove webhook, redrive DLQ, clear DLQ) were gated by native window.confirm() — unstyled, undismissable in tests, and inconsistent with the rest of the app. They now use shadcn AlertDialog with module-level copy constants (REMOVE_WEBHOOK_COPY, REDRIVE_DLQ_COPY, CLEAR_DLQ_COPY) so the strings are colocated and easy to edit.
Dialogs close in the handler's finally block so success and failure paths both clean up correctly. The existing deleting/redriving/clearingDlq pending flags disable the Cancel/Action buttons while work is in flight.
Addresses audit finding H7.
Two related cleanups from the re-audit of audit finding H10:
Delete components/status/StatusBadge.tsx entirely. The file exported four components (StatusBadge, AgentStateBadge, HealthStatusBadge, LifecycleStatusBadge) and two helpers (getHealthScoreColor, getHealthScoreBadgeVariant). Zero production consumers — only the file's own test imported any of them. A name collision with the 'real' StatusBadge in components/ui/badge.tsx (different API entirely) hid the fact that this file was dead for some time. ~325 LOC of source plus ~200 LOC of tests removed.
Collapse statusIcons / toneByVariant / variantToCanonical in components/ui/badge.tsx into a single STATUS_VARIANT_META record at module scope. The three per-render maps are now one lookup, and the map is hoisted out of the Badge function body (one less allocation per render). Rendered output is unchanged for every status variant.
components/AccessibilityEnhancements.tsx (376 lines, 10 exports) was mostly dead and entirely ceremonial. Six exports had zero consumers. The four imported by NodeDetailPage.tsx were all no-ops:
NodeDetailPage.tsx now drops the import block, unwraps the three <MCPAccessibilityProvider> wraps, deletes the announcer elements, and relies on the global skip link from AppLayout plus Radix Alert's built-in role. Net deletion: ~380 lines, zero a11y regression.
Agent lifecycle status (ready/starting/running/degraded/stopped/offline/ error/unknown) had three competing representations: AgentsPage's local bg-green-400 switch statements, StatusBadge.LIFECYCLE_CONFIG (now gone with the dead StatusBadge file), and node-status.ts's lossy mapping into execution CanonicalStatus (degraded→pending, offline→failed — silently wrong).
Introduces utils/lifecycle-status.ts as the single source of truth:
Adds LifecycleDot / LifecycleIcon / LifecyclePill primitives to components/ui/status-pill.tsx, parallel to StatusDot/Icon/Pill. Implements motion === 'pulse' for starting/degraded (motion-safe:animate-pulse) and motion === 'live' halo ping for running.
Rewrites utils/node-status.ts getNodeStatusPresentation to delegate directly to getLifecycleTheme. Drops the lossy canonical mapping. summarizeNodeStatuses now buckets via theme.online and LifecycleStatus directly. Return shape { kind, label, theme, shouldPulse, online } — consumers (NodeCard, NodesVirtualList, NodesStatusSummary, NodeCard, NodeDetailPage, EnhancedNodesHeader, EnhancedNodeDetailHeader) need no changes because field names match StatusTheme.
Refactors pages/AgentsPage.tsx: deletes getStatusDotColor, getStatusTextColor, and the local isAgentLifecycleOnline helper. Replaces inline dot+label markup with <LifecycleDot status={...} />. Both filter call sites switch to isLifecycleOnline from @/utils/lifecycle-status.
Intentional behavior changes (flagged for QA):
Adds utils/lifecycle-status.test.ts covering normalization (every enum, null/undefined/empty, trim/lowercase, aliases, unknown fallback), theme lookup, and isLifecycleOnline buckets.
Drops the runtime Google Fonts dependency entirely. Previously every page load hit fonts.googleapis.com + fonts.gstatic.com to fetch Inter weights 400/500/600/700 — blocking render on third-party DNS + TLS and leaking a request to Google on every session.
@fontsource/inter bundles the same weights with the app and Vite serves them from the same origin. The weight-specific CSS imports in main.tsx let Vite tree-shake unused weights. Addresses audit finding M19.
Sidebar navigation now renders two labeled groups:
Build:
Govern:
Auditor/governance features are now visible in the nav as first-class entries instead of being buried in Settings or shown as a single 'Audit' catch-all. The distribution-by-design replaces the previous distribution-by-accident (DID/VC tucked in Settings, provenance export hidden in a run kebab, etc.).
'Audit' → 'Provenance' — more honest about what the page does.
Addresses Round 2 audit finding A#4.
Shadcn's default TableHead/TableCell ship with h-12 p-4 (Stripe-roomy, ~52px row). The product target is Linear/Datadog tight (~30px row, ~35 rows visible at 1440x900). RunsPage was hand-overriding every cell with 'px-3 py-1.5' to achieve this — ~30 redundant overrides scattered through the file.
Fix the primitive once:
Drop the redundant overrides in RunsPage. Other consumers (PlaygroundPage, VerifyProvenancePage, ReasonersSkillsTable, ExecutionHistoryList, AgentNodesTable) will inherit the tighter density automatically, which is what we want. NewDashboardPage and ComparisonPage already use explicit per-cell spacing and are unaffected.
Addresses Round 2 audit finding B#1.
The live-data motion budget contract says: only the builder dashboard ticks live; all other pages are query-on-demand. But both useRuns and useAgents defaulted to auto-polling, which meant /runs and /agents silently refreshed every few seconds — contradicting the contract and adding motion cost everywhere.
Both hooks now default to refetchInterval: false. Callers that want polling pass an explicit interval. NewDashboardPage.tsx passes { refetchInterval: 10_000 } to useAgents (it already passed 8_000 to useRuns). RunsPage and AgentsPage now inherit the no-poll default and will get explicit Refresh buttons in a follow-up.
Addresses Round 2 audit finding C#3.
services/searchService.ts was orphan dead code: three searchX methods that returned [] literally, imported by zero live files, with one branch pointing to a deleted /nodes/:id route. Delete it.
CommandPalette.tsx is a navigation jumper, not a search — Round 1 already fixed the placeholder copy. Add a one-line comment at the top of the component clarifying this for future readers.
Addresses Round 2 audit finding A#3.
AccessManagementPage and NewSettingsPage used 'text-2xl font-bold' for their page titles, drifting from the 'text-2xl font-semibold tracking-tight' pattern that 7 of 10 live pages already converge on. Unify.
AccessManagementPage also wrapped its content in its own NotificationProvider, creating a second Sonner toast root under the app-level provider. Drop the wrapper — the app-root provider is sufficient. Addresses the 'duplicate NotificationProvider' finding from Round 2 Lens C.
Addresses Round 2 findings B#5 and C#2.
The Playground textarea had no accessible label, did not autofocus on mount, and had no keyboard shortcut to execute.
Playground is the builder persona's primary canvas. These should have shipped on day one.
Addresses Round 2 audit finding C#5.
Round 1 identified ~120 dead files; Round 2 re-traced the static-import closure and found the 'Workflow + Execution + Node + Reasoner' universe is a self-referential dead island. All its routes redirect via <Navigate> in App.tsx; none are reachable from the live page tree.
Deleted (46 files):
Pages (unreachable — routes redirect):
Components (zero live importers):
components/workflow/Enhanced* cluster + its barrel index.ts (WorkflowTimeline.tsx and TimelineNodeCard.tsx preserved — live).
Note: NodeDetailPage was modified in commit ef9aecf5 (Phase 2 a11y cleanup), but that commit was against a file whose route has been redirected for a while. Deleting it now closes that loop.
tsc --noEmit clean after deletion. No live file had a dead import.
Addresses Round 2 audit finding D#5.
Three related cleanups removing raw Tailwind palette classes and hardcoded hex literals that bypass the semantic-token system:
components/ui/notification.tsx Variant definitions previously shipped raw red/amber/blue/emerald/sky classes with explicit dark: branches inside their CVA table, leaking into every consumer (every toast, alert, banner across the app). Now routed through statusTone.{error,warning,info,success,neutral} from lib/theme.ts. Each semantic token already handles its own dark-mode variant via CSS vars, so the dark: overrides disappear.
components/ui/ErrorState.tsx Same pattern — raw red-500 / red-200 / red-900 replaced with statusTone.error.{bg,fg,border,accent}. Icon color routes through statusTone.error.accent instead of text-red-500.
utils/status.ts STATUS_HEX Previously a 20-entry Record<CanonicalStatus, {base, light}> with hardcoded hex values (#22c55e, #ef4444, etc.) and no dark-mode variant. Replaced with a runtime helper that reads the --status-* CSS custom properties from document.documentElement at call time, returning hsl() strings that correctly track the active theme. Also exposes getStatusGlowColor() for the color-mix transparent variant that formerly used STATUS_HEX[s].light.
Addresses Round 2 audit findings B#2 and B#3.
Co-authored-by: Claude Opus 4.6 (1M context) noreply@anthropic.com (e81b948)
Docs(skill): meta-philosophy rewrite + async curl + post-mortem hardenings
Rewrites the agentfield-multi-reasoner-builder skill to teach principles instead of prescribing specific shapes, switches the canonical smoke test to the async endpoint, and folds in hard-learned lessons about framework contracts, cross-boundary data, and live validation.
All changes are to the skill's markdown — no Go code, no templates, no install flow changes. The skill_data/ copy under control-plane/internal/ skillkit/ is the embedded mirror of skills/ and is kept in sync via scripts/sync-embedded-skills.sh.
Feat(skill+cli): agentfield-multi-reasoner-builder skill, af doctor, af init --docker, af skill install
Adds the agentfield-multi-reasoner-builder skill, af doctor, af init --docker, and the full af skill install architecture (embed + 7 target integrations + state tracking). install.sh now installs the skill into every detected coding agent by default. See PR #367 for the full breakdown, end-to-end test results, and design rationale. (c34e3e6)
Feat(runs): pause/resume/cancel + unified status primitives + notification center
Replaces the toast-only notification system with a dual-mode center:
Adds full lifecycle controls to the runs index page:
Cross-cutting UX pass addressing multiple issues from rapid review:
Backend
Unified status primitives (web)
icon: LucideIcon and
motion: "none" | "live". Single source of truth for glyph and motion
per canonical status.Runs table liveness
root_execution_status ?? status. Paused/cancelled rows stop
ticking immediately even when aggregate stays running.Run detail page
Dashboard
Notification center — compact tree, semantic icons
to={/runs/${runId}}
directly (no href string manipulation) so basename=/ui prepends
properly. Sonner toast action builds the URL from VITE_BASE_PATH.Top bar + layout
Dependencies
Backend
Frontend — single source of truth via getStatusTheme()
variant === "running".Frontend — root-effective status consistency
Frontend — accessibility + contrast
Deferred (recorded for follow-up, out of scope for this pass)
Verified: tsc clean, go build clean, go vet clean, no new lint errors in touched files.
Co-Authored-By: Claude Opus 4.6 (1M context) noreply@anthropic.com
Prevents Hypothesis test framework cache files from being tracked.
Co-Authored-By: Claude Opus 4.6 (1M context) noreply@anthropic.com
The pause, resume, and cancel handlers required a workflow_executions entry to exist, but simple async single-node executions only create rows in the executions table. This caused a 404 for all lifecycle actions in local mode.
Make the workflow_executions lookup non-fatal: use it when available for UpdateWorkflowExecution and event metadata, otherwise fall back to the execution record's RunID.
Co-Authored-By: Claude Opus 4.6 (1M context) noreply@anthropic.com
Co-authored-by: Claude Opus 4.6 (1M context) noreply@anthropic.com Co-authored-by: Abir Abbas abirabbas1998@gmail.com (3b8d302)
Chore(deps): bump go.opentelemetry.io/otel/sdk
Bumps the go_modules group with 1 update in the /control-plane directory: go.opentelemetry.io/otel/sdk.
Updates go.opentelemetry.io/otel/sdk from 1.39.0 to 1.40.0
updated-dependencies:
Signed-off-by: dependabot[bot] support@github.com (a536086)
Fix: remediate CodeQL security alerts
Feat(observability): add OpenTelemetry distributed tracing export
Add OTel trace export to the control plane execution pipeline. Each execution creates a root span, and reasoner/skill invocations create child spans with execution metadata as span attributes.
control-plane/internal/observability/ package with TracerProvider
initialization (OTLP HTTP exporter) and ExecutionTracer that subscribes
to the existing execution and reasoner event busesfeatures.tracing.enabled: true or
AGENTFIELD_TRACING_ENABLED=trueThe initial OTel dependency pull upgraded transitive deps (golang.org/x/*, grpc) that require Go 1.25, but CI uses Go 1.24. Pin OTel to v1.32-v1.35 and restore the original grpc/x/ versions.
Co-authored-by: Abir Abbas abirabbas1998@gmail.com Co-authored-by: Claude Opus 4.6 (1M context) noreply@anthropic.com (a9884e8)
Refactor: remove all MCP code from codebase
MCP UI was already removed in the latest main revamp. This cleans up all remaining MCP backend code, endpoints, CLI commands, SDK integrations, types, tests, and documentation across the entire codebase.
Removed:
Co-Authored-By: Claude Opus 4.6 (1M context) noreply@anthropic.com
Co-Authored-By: Claude Opus 4.6 (1M context) noreply@anthropic.com
AgentNodeDetailsForUI import in api.ts (TS6196)asdict import in test_types.py (F401)Co-Authored-By: Claude Opus 4.6 (1M context) noreply@anthropic.com
Co-authored-by: Claude Opus 4.6 (1M context) noreply@anthropic.com (f732ed5)
Fix: resolve dependency and secret alerts
fix: resolve dependency and secret alerts
Fix TS SDK CI install and test scope (4980751)
Chore(deps): bump the npm_and_yarn group across 1 directory with 3 updates
Bumps the npm_and_yarn group with 2 updates in the /sdk/typescript directory: vite and path-to-regexp.
Updates vite from 5.4.21 to 8.0.5
Updates esbuild from 0.21.5 to 0.25.12
Updates path-to-regexp from 0.1.12 to 0.1.13
updated-dependencies:
Signed-off-by: dependabot[bot] support@github.com
Signed-off-by: dependabot[bot] support@github.com Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com> Co-authored-by: Santosh santosh@agentfield.ai (a912244)
Chore(deps): bump the go_modules group across 1 directory with 3 updates
Bumps the go_modules group with 3 updates in the /control-plane directory: golang.org/x/crypto, google.golang.org/grpc and github.com/go-viper/mapstructure/v2.
Updates golang.org/x/crypto from 0.37.0 to 0.45.0
Updates google.golang.org/grpc from 1.67.3 to 1.79.3
Updates github.com/go-viper/mapstructure/v2 from 2.2.1 to 2.4.0
updated-dependencies:
Signed-off-by: dependabot[bot] support@github.com Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com> (436bbc5)
Chore(deps): bump the npm_and_yarn group across 5 directories with 21 updates
Bumps the npm_and_yarn group with 6 updates in the /sdk/typescript directory:
| Package | From | To |
|---|---|---|
| axios | 1.13.2 |
1.13.5 |
| express-rate-limit | 8.2.1 |
8.2.2 |
| esbuild | 0.21.5 |
0.27.0 |
| path-to-regexp | 0.1.12 |
0.1.13 |
| qs | 6.13.0 |
6.14.2 |
| rollup | 4.53.3 |
4.60.1 |
Bumps the npm_and_yarn group with 2 updates in the /examples/ts-node-examples directory: picomatch and rollup. Bumps the npm_and_yarn group with 2 updates in the /examples/python_agent_nodes/rag_evaluation/ui directory: picomatch and next. Bumps the npm_and_yarn group with 1 update in the /examples/benchmarks/100k-scale/mastra-bench directory: ai. Bumps the npm_and_yarn group with 13 updates in the /control-plane/web/client directory:
| Package | From | To |
|---|---|---|
| picomatch | 2.3.1 |
2.3.2 |
| picomatch | 4.0.2 |
4.0.4 |
| picomatch | 4.0.3 |
4.0.4 |
| rollup | 4.43.0 |
4.60.1 |
| vite | 6.3.5 |
6.4.2 |
| mdast-util-to-hast | 13.2.0 |
13.2.1 |
| js-yaml | 4.1.0 |
4.1.1 |
| brace-expansion | 1.1.12 |
1.1.13 |
| flatted | 3.3.3 |
3.4.2 |
| lodash | 4.17.21 |
4.18.1 |
| minimatch | 3.1.2 |
3.1.5 |
| preact | 10.27.2 |
10.29.1 |
| react-router | 7.6.2 |
7.14.0 |
| tmp | 0.2.3 |
0.2.5 |
| yaml | 2.8.0 |
2.8.3 |
| yaml | 1.10.2 |
1.10.3 |
Updates axios from 1.13.2 to 1.13.5
Updates express-rate-limit from 8.2.1 to 8.2.2
Updates esbuild from 0.21.5 to 0.27.0
Updates path-to-regexp from 0.1.12 to 0.1.13
Updates picomatch from 4.0.3 to 4.0.4
Updates qs from 6.13.0 to 6.14.2
Updates rollup from 4.53.3 to 4.60.1
Updates vite from 5.4.21 to 8.0.5
Updates picomatch from 4.0.3 to 4.0.4
Updates rollup from 4.53.3 to 4.60.1
Updates picomatch from 2.3.1 to 2.3.2
Updates picomatch from 4.0.3 to 4.0.4
Updates next from 14.2.15 to 15.5.14
Removes ai
Updates hono from 4.11.3 to 4.12.12
Updates picomatch from 2.3.1 to 2.3.2
Updates picomatch from 4.0.2 to 4.0.4
Updates picomatch from 4.0.3 to 4.0.4
Updates rollup from 4.43.0 to 4.60.1
Updates vite from 6.3.5 to 6.4.2
Updates mdast-util-to-hast from 13.2.0 to 13.2.1
Updates js-yaml from 4.1.0 to 4.1.1
Updates brace-expansion from 1.1.12 to 1.1.13
Updates flatted from 3.3.3 to 3.4.2
Updates lodash from 4.17.21 to 4.18.1
Updates minimatch from 3.1.2 to 3.1.5
Updates preact from 10.27.2 to 10.29.1
Updates react-router from 7.6.2 to 7.14.0
Updates tmp from 0.2.3 to 0.2.5
Updates yaml from 2.8.0 to 2.8.3
Updates yaml from 1.10.2 to 1.10.3
updated-dependencies:
Signed-off-by: dependabot[bot] support@github.com
fix(ci): unbreak sdk and bot PR checks
fix(sdk): keep vitest on node 18 compatible track
Signed-off-by: dependabot[bot] support@github.com Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com> Co-authored-by: Santosh santosh@agentfield.ai (37ecbf6)
Nothing published for this version
Feat(sdk/python): add per-execution LLM cost tracking
Add CostTracker that accumulates token usage and cost data from app.ai() calls. Each LLM response's cost_usd (via LiteLLM) is recorded before schema parsing strips the multimodal metadata.
docs: add changelog entry for cost tracking
fix: wire reasoner_name from execution context into cost tracking
The CostTracker.record() accepts reasoner_name but the call site in agent_ai.py never passed it. Now reads it from the current ExecutionContext, which is set by the @reasoner decorator. (c29a2a8)
Remove deprecated routes from App.tsx router
feat(sdk/ts): add multimodal helpers for image, audio, file inputs/outputs
refactor(ui): migrate icons from Phosphor to Lucide
Replace all @phosphor-icons/react imports with lucide-react equivalents. Rewrote icon-bridge.tsx to re-export Lucide icons under the same names used throughout the codebase, so no consumer files needed changing. Updated icon.tsx to use Lucide directly. Removed weight= props from badge.tsx, segmented-status-filter.tsx, and ReasonerCard.tsx since Lucide does not support the Phosphor weight API.
API services (vcApi, mcpApi), types, and non-MCP hooks are preserved. TypeScript check passes with zero errors after cleanup.
Add RunsPage component at /runs with:
Wire RunsPage into App.tsx replacing the placeholder at /runs.
Adds a new /playground and /playground/:reasonerId route with:
Replaces /dashboard with NewDashboardPage — a focused, operations-first view that answers "Is anything broken? What's happening now?" rather than displaying metrics charts. The legacy enhanced dashboard is preserved at /dashboard/legacy.
Key sections:
Replaces the /agents placeholder with a fully functional page showing each registered agent node as a collapsible Card. Each card displays status badge with live dot, last heartbeat, reasoner/skill count, health score, and an inline reasoner list fetched lazily from GET /nodes/:id/details. Supports Restart and Config actions. Auto- refreshes every 10 s via useAgents polling.
Adds NewSettingsPage with four tabs:
Updates App.tsx to route /settings to NewSettingsPage and redirect /settings/observability-webhook to /settings.
Co-Authored-By: Claude Sonnet 4.6 noreply@anthropic.com
Adds RunDetailPage (/runs/:runId) as the primary execution inspection screen, replacing the placeholder. Features a split-panel layout with a proportional-bar execution trace tree on the left and collapsible Input/Output/Notes step detail on the right. Single-step runs skip the trace and show step detail directly. Includes smart polling for active runs and a Trace/Graph toggle (graph view placeholder).
New files:
Co-Authored-By: Claude Sonnet 4.6 noreply@anthropic.com
Co-Authored-By: Claude Sonnet 4.6 noreply@anthropic.com
Wire up the existing WorkflowDAGViewer component into the Run Detail page as a proper Graph tab alongside the Trace view. Multi-step runs show a Trace/Graph toggle in the header; single-step runs skip the toggle entirely and show step detail directly. Clicking a node in the graph panel selects the step and populates the right-hand detail panel.
Adds a CommandPalette component using shadcn Command + Dialog, registered globally via AppLayout. Cmd+K / Ctrl+K toggles the palette; items navigate to Dashboard, Runs, Agents, Playground, Settings, and filtered run views. A ⌘K hint badge is shown in the header bar on medium+ screens.
Adds /runs/compare?a=RUN_ID_1&b=RUN_ID_2 route accessible from the Runs page when exactly 2 runs are selected via "Compare Selected". The page shows dual summary cards (status, step count, failures, duration delta), a step-by-step alignment table with same/diverged/extra annotations, and a clickable output-diff panel for diverged rows.
Co-Authored-By: Claude Sonnet 4.6 noreply@anthropic.com
Co-Authored-By: Claude Sonnet 4.6 noreply@anthropic.com
Co-Authored-By: Claude Sonnet 4.6 noreply@anthropic.com
fix(ui): graph edges invisible — wrap HSL vars in hsl() for SVG stroke attributes
web(ui): revamp layout, filters, and health strip
Drop docs/ui-revamp and embedded client audit notes from the branch; surface product direction in the PR instead. Remove accidentally tracked .agit index and Claude worktree pointer; add ignore rules.
Multi-pass code review surfaced 60 issues across security, code quality, performance, API design, frontend, and test coverage. This commit fixes the critical, high, and medium severity items:
fix: move ErrorBoundary inside Router to preserve routing context
style(web): redesign sign-in page with shadcn components and app branding
Establishes a PR stack on top of #330. No functional changes yet.
agit sync
agit sync
fix(web): tighten live update gating
feat: agent node process logs (Python NDJSON, CP proxy, UI panel)
docs: document node log proxy and agent log env vars in .env.example
fix: align .env.example log vars with Python node_logs.py
agit sync
agit sync
chore: log-demo compose, UI proxy functional test, example heartbeats
Add wait-control-plane (curl with retries) and depends_on completed_successfully so Python/Go/Node demo agents do not register before the API is ready. Set restart: unless-stopped on demo agents. Document behavior in LOG_DEMO.md.
Add run-log-demo-native.sh / stop-log-demo-native.sh with writable paths (env overrides for SQLite, Bolt, DID keystore under /tmp/agentfield-log-demo). Makefile targets log-demo-native-up/down. Document in LOG_DEMO.md.
Fixes local runs using agentfield-test.yaml which points storage at /data.
fix: tighten audit verification and multimodal IO
feat(ui): agents page heading, nav labels, compact responsive logs
Unblocks Python SDK CI lint-and-test job.
Bare Response resolved to DOM Fetch Response; use import('express').Response.
fix(ci): harden DID fetches and stabilize node log proxy test
agit sync
fix(web): avoid format strings in node action logs
fix(cli): keep VC verification offline-only
fix(tests): write node log marker into process ring
docs: add execution observability RFC
agit sync
agit sync
agit sync
agit sync
agit sync
agit sync
agit sync
feat: add execution logging transport
feat(py): add structured execution logging
agit sync
feat(web): unify execution observability panel
Remove local plandb database from repo
fix demo execution observability flow
refactor run detail into execution and logs tabs
refine execution logs density
agit sync
polish execution logs panel header
refine raw node log console
polish process log filter toolbar
standardize observability spacing primitives
agit sync
test execution observability in functional harness
agit sync
test(go-sdk): add tests for types/status, types/types, types/discovery, did/types, agent/harness
Adds comprehensive test coverage for previously untested Go SDK packages:
or True assertion with real post-condition checkAdd test_invariants.py with 15 tests covering:
Adds 68 invariant tests across 5 new files verifying structural properties that must always hold: circuit breaker state machine transitions, consecutive failure counter monotonicity, nonce uniqueness and timestamp ordering in DID signing, scope-to-header mapping stability in MemoryClient, reasoner/skill namespace independence in Agent, and AsyncLocalStorage scope isolation in ExecutionContext (including concurrent execution safety).
Adds 35 property-based invariant tests across four packages designed to catch regressions in AI-generated code changes by verifying structural contracts rather than specific input→output pairs.
agit sync
agit sync
test(python-sdk): add behavioral invariant tests for DID auth, memory, connection, execution state, agent lifecycle
71 invariant tests covering:
Ruff auto-fixed 48 lint issues across all Python test files:
Invariant tests covering:
Agent() constructor and handle_serverless() use asyncio internally. On Python 3.8/3.9, asyncio.get_event_loop() raises RuntimeError in non-async contexts without an explicitly created loop. Add autouse fixture that ensures a loop exists.
All 68 new tests pass (24 SDK + 44 Web UI).
Covers the four highest-impact categories of LocalStorage operations:
All tests run against a real SQLite + BoltDB backed by t.TempDir().
test(control-plane): add handler and service tests — agentic status, config storage, memory ACL, tag normalization
fix(storage): graceful LIKE fallback when SQLite FTS5 module is unavailable
QueryWorkflowExecutions now checks hasFTS5 flag and falls back to LIKE-based search across key columns instead of failing with "no such table: workflow_executions_fts".
Python SDK:
Go Control Plane:
TypeScript SDK:
fix(ci): remove unused imports — os in test_agent_server.py, beforeEach/makeFetchMock in api.test.ts
agit sync
fix(tests): align tests with main branch changes, remove .cursor files
Co-Authored-By: Claude Opus 4.6 (1M context) noreply@anthropic.com
Co-authored-by: Claude Sonnet 4.6 noreply@anthropic.com Co-authored-by: Abir Abbas abirabbas1998@gmail.com (bd41402)
Remove deprecated routes from App.tsx router
feat(sdk/ts): add multimodal helpers for image, audio, file inputs/outputs
refactor(ui): migrate icons from Phosphor to Lucide
Replace all @phosphor-icons/react imports with lucide-react equivalents. Rewrote icon-bridge.tsx to re-export Lucide icons under the same names used throughout the codebase, so no consumer files needed changing. Updated icon.tsx to use Lucide directly. Removed weight= props from badge.tsx, segmented-status-filter.tsx, and ReasonerCard.tsx since Lucide does not support the Phosphor weight API.
API services (vcApi, mcpApi), types, and non-MCP hooks are preserved. TypeScript check passes with zero errors after cleanup.
Add RunsPage component at /runs with:
Wire RunsPage into App.tsx replacing the placeholder at /runs.
Adds a new /playground and /playground/:reasonerId route with:
Replaces /dashboard with NewDashboardPage — a focused, operations-first view that answers "Is anything broken? What's happening now?" rather than displaying metrics charts. The legacy enhanced dashboard is preserved at /dashboard/legacy.
Key sections:
Replaces the /agents placeholder with a fully functional page showing each registered agent node as a collapsible Card. Each card displays status badge with live dot, last heartbeat, reasoner/skill count, health score, and an inline reasoner list fetched lazily from GET /nodes/:id/details. Supports Restart and Config actions. Auto- refreshes every 10 s via useAgents polling.
Adds NewSettingsPage with four tabs:
Updates App.tsx to route /settings to NewSettingsPage and redirect /settings/observability-webhook to /settings.
Co-Authored-By: Claude Sonnet 4.6 noreply@anthropic.com
Adds RunDetailPage (/runs/:runId) as the primary execution inspection screen, replacing the placeholder. Features a split-panel layout with a proportional-bar execution trace tree on the left and collapsible Input/Output/Notes step detail on the right. Single-step runs skip the trace and show step detail directly. Includes smart polling for active runs and a Trace/Graph toggle (graph view placeholder).
New files:
Co-Authored-By: Claude Sonnet 4.6 noreply@anthropic.com
Co-Authored-By: Claude Sonnet 4.6 noreply@anthropic.com
Wire up the existing WorkflowDAGViewer component into the Run Detail page as a proper Graph tab alongside the Trace view. Multi-step runs show a Trace/Graph toggle in the header; single-step runs skip the toggle entirely and show step detail directly. Clicking a node in the graph panel selects the step and populates the right-hand detail panel.
Adds a CommandPalette component using shadcn Command + Dialog, registered globally via AppLayout. Cmd+K / Ctrl+K toggles the palette; items navigate to Dashboard, Runs, Agents, Playground, Settings, and filtered run views. A ⌘K hint badge is shown in the header bar on medium+ screens.
Adds /runs/compare?a=RUN_ID_1&b=RUN_ID_2 route accessible from the Runs page when exactly 2 runs are selected via "Compare Selected". The page shows dual summary cards (status, step count, failures, duration delta), a step-by-step alignment table with same/diverged/extra annotations, and a clickable output-diff panel for diverged rows.
Co-Authored-By: Claude Sonnet 4.6 noreply@anthropic.com
Co-Authored-By: Claude Sonnet 4.6 noreply@anthropic.com
Co-Authored-By: Claude Sonnet 4.6 noreply@anthropic.com
fix(ui): graph edges invisible — wrap HSL vars in hsl() for SVG stroke attributes
web(ui): revamp layout, filters, and health strip
Drop docs/ui-revamp and embedded client audit notes from the branch; surface product direction in the PR instead. Remove accidentally tracked .agit index and Claude worktree pointer; add ignore rules.
Multi-pass code review surfaced 60 issues across security, code quality, performance, API design, frontend, and test coverage. This commit fixes the critical, high, and medium severity items:
fix: move ErrorBoundary inside Router to preserve routing context
style(web): redesign sign-in page with shadcn components and app branding
Establishes a PR stack on top of #330. No functional changes yet.
agit sync
agit sync
fix(web): tighten live update gating
feat: agent node process logs (Python NDJSON, CP proxy, UI panel)
docs: document node log proxy and agent log env vars in .env.example
fix: align .env.example log vars with Python node_logs.py
agit sync
agit sync
chore: log-demo compose, UI proxy functional test, example heartbeats
Add wait-control-plane (curl with retries) and depends_on completed_successfully so Python/Go/Node demo agents do not register before the API is ready. Set restart: unless-stopped on demo agents. Document behavior in LOG_DEMO.md.
Add run-log-demo-native.sh / stop-log-demo-native.sh with writable paths (env overrides for SQLite, Bolt, DID keystore under /tmp/agentfield-log-demo). Makefile targets log-demo-native-up/down. Document in LOG_DEMO.md.
Fixes local runs using agentfield-test.yaml which points storage at /data.
fix: tighten audit verification and multimodal IO
feat(ui): agents page heading, nav labels, compact responsive logs
Unblocks Python SDK CI lint-and-test job.
Bare Response resolved to DOM Fetch Response; use import('express').Response.
fix(ci): harden DID fetches and stabilize node log proxy test
agit sync
fix(web): avoid format strings in node action logs
fix(cli): keep VC verification offline-only
fix(tests): write node log marker into process ring
docs: add execution observability RFC
agit sync
agit sync
agit sync
agit sync
agit sync
agit sync
agit sync
feat: add execution logging transport
feat(py): add structured execution logging
agit sync
feat(web): unify execution observability panel
Remove local plandb database from repo
fix demo execution observability flow
refactor run detail into execution and logs tabs
refine execution logs density
agit sync
polish execution logs panel header
refine raw node log console
polish process log filter toolbar
standardize observability spacing primitives
agit sync
test execution observability in functional harness
fix: security hardening for execution logs and process log auth
Co-Authored-By: Claude Opus 4.6 (1M context) noreply@anthropic.com
The Python SDK logger only wrote structured logs to stdout, unlike the Go SDK which also dispatches them to POST /api/v1/executions/:id/logs. This caused the Logs tab in the UI to always show empty for Python agents.
Co-Authored-By: Claude Opus 4.6 (1M context) noreply@anthropic.com
The HealthStrip filtered on health_status === "ready" | "active" but agents register with lifecycle_status: "ready" and health_status may be "inactive" or "unknown". This caused "0/5 online" while the table showed all agents as ready with green dots.
Co-Authored-By: Claude Opus 4.6 (1M context) noreply@anthropic.com
Approval flow fixes:
SSE connection exhaustion fix:
Co-Authored-By: Claude Opus 4.6 (1M context) noreply@anthropic.com
The ApproveWithContextDialog opened an empty tag selector with a disabled Approve button when the agent had no proposed_tags. Now it shows an "Approve Agent" flow explaining the agent registered without tags and lets the admin approve the registration directly.
Co-Authored-By: Claude Opus 4.6 (1M context) noreply@anthropic.com
The Replay button navigated to the playground without passing the execution's input data, leaving the input editor empty. Now it passes input_data via React Router location state, and the playground reads it to seed the JSON editor on mount.
Co-Authored-By: Claude Opus 4.6 (1M context) noreply@anthropic.com
The Replay button fetches execution details to get the original input_data before navigating to the playground. The playground reads replay input from location state and skips the schema-based empty seed when replay data is present.
Co-Authored-By: Claude Opus 4.6 (1M context) noreply@anthropic.com
Co-authored-by: Claude Sonnet 4.6 noreply@anthropic.com Co-authored-by: Abir Abbas abirabbas1998@gmail.com (3d2e4d7)
Feat(sdk/python): add per-execution LLM cost tracking
remove OpenCode serve+attach workaround for stateless CLI
fix: use correct OpenCode CLI flags (-p, -c, MODEL env var)
fix: Restored system_prompt support in both SDKs (2aed6bd)
Feat(observability): enrich execution facts for downstream analysis
feat: enrich execution observability facts
test: cover execution observability contract
test: add functional execution webhook coverage
fix: use uvicorn websockets sansio (4a04419)
Refactor: replace emoji logging with structured zerolog in memory handler
Convert all 18 emoji-based debug log statements to structured zerolog calls with typed fields (operation, scope, key, err, bytes). Log levels adjusted: Debug for normal flow, Error for failures, Warn for edge cases. No handler logic changed.
Closes #114
Co-authored-by: Claude Opus 4.6 (1M context) noreply@anthropic.com (208be41)
Refactor: extract shared ErrorResponse helper to handlers/errors.go
Replace generic utm_medium=referral with unique utm_id per link and add utm_campaign=github-readme consistently across all 35 outbound links. This enables granular click attribution in analytics to see exactly which README section drives traffic.
Co-Authored-By: Claude Opus 4.6 (1M context) noreply@anthropic.com
Co-Authored-By: Claude Opus 4.6 (1M context) noreply@anthropic.com
Co-Authored-By: Claude Opus 4.6 (1M context) noreply@anthropic.com
Co-Authored-By: Claude Opus 4.6 (1M context) noreply@anthropic.com
Co-Authored-By: Claude Opus 4.6 (1M context) noreply@anthropic.com
Co-Authored-By: Claude Opus 4.6 (1M context) noreply@anthropic.com
Co-authored-by: OG oktaygoktas@users.noreply.github.com Co-authored-by: Claude Opus 4.6 (1M context) noreply@anthropic.com (6e4f258)
Nothing published for this version
Fix(sdk/python): replace deprecated datetime.utcnow() with timezone-a…
datetime.utcnow() is deprecated since Python 3.12 and scheduled for removal. It produces naive datetime objects with no timezone info, which causes DeprecationWarning on every test run across 5 files.
Additionally, two call sites combined .isoformat() + 'Z' on a timezone-aware datetime, producing invalid ISO strings like: '2026-04-05T01:30:20.584591+00:00Z' ← double timezone suffix
This caused test failures in test_vc_generator.py.
Changes:
All 745 tests pass with 0 failures after this change. DeprecationWarnings reduced from 36 to 0.
Per code review feedback from @santoshkumarradha:
All 745 tests pass. (265cf79)
Docs: add CloudSecurity AF to Built With AgentField section
Co-authored-by: OG oktaygoktas@users.noreply.github.com Co-authored-by: Claude Opus 4.6 (1M context) noreply@anthropic.com (4620a4b)
Six SSE handler endpoints hardcode Access-Control-Allow-Origin: *,
bypassing the global gin-contrib/cors middleware which enforces a
restricted origin list from configuration. This allows any origin to
connect to real-time streaming endpoints and exfiltrate execution data,
log streams, and agent status updates.
Remove the per-handler CORS headers and let the router-level middleware handle origin validation consistently across all endpoints.
Affected handlers:
Co-authored-by: José Maia glitch-ux@users.noreply.github.com (9b41051)
Nothing published for this version
Nothing published for this version
Fix(web-ui): remove debug console logs
Refs #107
Co-authored-by: nanqinhu 139929317+nanqinhu@users.noreply.github.com (c35b248)
Feat: execution resilience - LLM health, concurrency limits, retry, log streaming
Add execution resilience features to address stuck job detection and recovery scenarios reported in #316:
Co-Authored-By: Claude Opus 4.6 (1M context) noreply@anthropic.com
Add structured error diagnostics so users can distinguish WHY executions fail (LLM down vs agent crash vs timeout vs unreachable), expose queue depth for concurrency monitoring, and rewrite E2E tests to validate real failure scenarios from #316.
Co-Authored-By: Claude Opus 4.6 (1M context) noreply@anthropic.com
Co-authored-by: Claude Opus 4.6 (1M context) noreply@anthropic.com Co-authored-by: Santosh santosh@agentfield.ai (93dc989)
Added type annotations to Agent.__init__
Fix(web-ui): replace unsafe any types in execution form
Co-authored-by: nanqinhu 139929317+nanqinhu@users.noreply.github.com (ab9ee07)
Test(web-ui): add node card coverage
Co-authored-by: nanqinhu 139929317+nanqinhu@users.noreply.github.com (527c78c)
Fix(web-ui): add accessible labels to node card
Co-authored-by: nanqinhu 139929317+nanqinhu@users.noreply.github.com (cccbe71)
Chore: compress README images for faster loading
Co-authored-by: nanqinhu 139929317+nanqinhu@users.noreply.github.com (6bbad4a)
OpenRouter API intermittently times out (60s), causing CI flakes.
Nothing published for this version
Add unit test for retry handler
Add unit tests for retry helpers (isRetryableDBError, backoffDelay)
added better documentation (1dec612)
Feat(sdk/go): add Codex and Gemini harness providers
Implements the remaining two CLI-based harness providers for the Go SDK, completing the feature parity with Python/TS implementations.
codex exec --json with JSONL event parsinggemini -p with plain text outputFixes #207
Co-Authored-By: Claude Opus 4.6 (1M context) noreply@anthropic.com
Codex: https://github.com/openai/codex Gemini: https://github.com/google-gemini/gemini-cli
Co-Authored-By: Claude Opus 4.6 (1M context) noreply@anthropic.com
Real Codex CLI uses thread_id (not session_id) in thread.started events and text (not content) in agent_message items. Fixed parser to handle both field names. Added integration test (build tag: integration).
Co-Authored-By: Claude Opus 4.6 (1M context) noreply@anthropic.com
Co-Authored-By: Claude Opus 4.6 (1M context) noreply@anthropic.com
Run with: go test ./harness/ -tags=integration Requires: codex CLI auth'd, GEMINI_API_KEY env var set with gemini settings.json auth type set to gemini-api-key.
Co-Authored-By: Claude Opus 4.6 (1M context) noreply@anthropic.com
Co-authored-by: Claude Opus 4.6 (1M context) noreply@anthropic.com (94c67c1)
Update all visual assets and badge colors to match the new agentfield.ai website theme (dark #0c0b09 background, gold #d4a24a accents, cream #f5f0eb text).
Feat: add community project submission issue template
The website was reorganized from flat paths to Learn/Build/Reference sections. This updates all agentfield.ai links across the repo to use canonical URLs instead of relying on redirects (or 404ing).
Key path changes:
Co-authored-by: Claude Opus 4.6 (1M context) noreply@anthropic.com (134b3db)
Nothing published for this version
Fix: QA authorization issues — revoke idempotency, health checks, UI bugs
Co-Authored-By: Claude Opus 4.6 (1M context) noreply@anthropic.com
Newly registered agents get HealthStatusUnknown (no heartbeat yet). The previous check required HealthStatusActive, which caused all functional tests to 503 because agents register and immediately invoke reasoners before a heartbeat can promote them to active.
Changed the guard to only block HealthStatusInactive agents, which are definitively unreachable. Unknown and active agents pass through.
Co-Authored-By: Claude Opus 4.6 (1M context) noreply@anthropic.com
Co-authored-by: Claude Opus 4.6 (1M context) noreply@anthropic.com (4318c57)
Nothing published for this version
Fix: resolve PR #284 review issues — logging, flag scoping, dead code, docstring
Fix: resolve PR #284 review issues — logging, flag scoping, dead code, docstring (#290)
Remove empty PersistentPreRun on agent command that silently overrode root PersistentPreRunE, preventing logger initialization for all af agent subcommands
Move --output and --timeout flags from root PersistentFlags to agent command scope (avoid polluting af dev, af server, etc.)
Remove redundant server URL fallback in agentHTTP() (GetServerURL already handles the default)
Fix Python SDK split docstring: merge orphaned string literal back into init docstring, place resolution code after docstring
Add doc comment to packages/server_url.go for parity with services/server_url.go (28beee5)
Feat(typescript sdk): Add history() method for querying past memory events
Signed-off-by: Roberto Robles ro-robles@pm.me
The server returns null instead of [] for empty result sets, causing
a TypeError when callers access the return value. Use nullish coalescing
to always return an array.
Co-Authored-By: Claude Opus 4.6 (1M context) noreply@anthropic.com
Signed-off-by: Roberto Robles ro-robles@pm.me Co-authored-by: Abir Abbas abirabbas1998@gmail.com Co-authored-by: Claude Opus 4.6 (1M context) noreply@anthropic.com (fc69fee)
Replace old "Kubernetes for AI Agents" banner with new developer-focused banner showing real AgentField code with selective highlighting. Update tagline to "Build and scale AI agents like APIs. Deploy, observe, and prove." (d9a28e2)
Docs: rewrite README - tighter structure, full feature showcase, new architecture diagram
Restructured from 492 to ~290 lines with conversion-optimized flow
New hero code example: claims processor showing app.ai(), app.pause(), app.call(), versioning, tags
Quick Start moved above examples for action-first flow
Added 90+ feature showcase in collapsible section with categorized tables
New wide 16:9 architecture diagram with three-zone layout (Your Services / The AI Backend / Agent Fleet)
Added features strip image as visual CTA for capabilities section
Repositioned Govern section as IAM for AI agents
All agentfield.ai links now have UTM tracking (utm_source=github-readme&utm_medium=referral)
Added SDK links (Python, Go, TypeScript, REST API) to header nav
Replaced em dashes with regular dashes throughout
Added canary deployments, agent discovery, HITL, connector API, memory events to feature list
Honest "Is AgentField for you?" scoping section
Quick start verified end-to-end in fresh Docker container (b161913)
Nothing published for this version
Nothing published for this version
Nothing published for this version
Feat: extend permission middleware to memory endpoints
Add MemoryPermissionMiddleware that enforces tag-based access policies and scope ownership checks on all memory routes, achieving permission parity with the existing execute endpoint protection. Wire up AccessControlMetadata fields (RequiredRoles, TeamRestricted, AuditAccess) and add comprehensive cross-agent isolation tests.
Co-Authored-By: Claude Opus 4.6 (1M context) noreply@anthropic.com
Add MemoryPermissionMiddleware that enforces tag-based access policies and scope ownership validation on all memory routes (key-value, vector, events). Wire up AccessControlMetadata fields (RequiredRoles, TeamRestricted, AuditAccess) on Memory records. Add cross-agent isolation tests verifying agents cannot read/write/delete other agents' scoped memory.
Key decisions:
Co-Authored-By: Claude Opus 4.6 (1M context) noreply@anthropic.com
When an agent provides identity headers but isn't registered as a node in storage, GetAgent returns an error. Previously this blocked all non-global memory access for unregistered agents. Now we skip tag-based policy evaluation (no tags to evaluate) while still enforcing scope ownership validation.
Co-Authored-By: Claude Opus 4.6 (1M context) noreply@anthropic.com
Co-authored-by: Claude Opus 4.6 (1M context) noreply@anthropic.com (08dfbfe)
Fix: capture Claude Agent SDK response by checking subtype='success'
The Claude Code provider checked for type=='result' to extract the response text, but the Claude Agent SDK sends the final message with subtype=='success' (no type field). This caused HarnessResult.result to always be None.
Fixes #252
Co-authored-by: Claude noreply@anthropic.com (a698c32)
Feat: CLI agent mode — af agent subcommand with structured JSON output
Implements af agent subcommand (Phase 4 of the Agentic API Layer epic):
DX fixes:
Closes #281, closes #282
When an AI agent runs 'af help', 'af badcommand', or 'af --badflags', the output now includes a hint: AI Agent? Run "af agent help" for structured JSON output. This creates a self-correcting loop — agents that accidentally invoke the human CLI are guided to agent mode.
When agents run wrong commands (af badcommand or af agent badcommand), they now receive parseable JSON errors with hint fields and available command lists instead of Cobra's default human-oriented text.
af agent help to structured JSON help outputThe RunE handler's ArbitraryArgs intercepted "help" as an unknown
subcommand before cobra could route to SetHelpCommand. Handle it
explicitly so af agent help matches af agent behavior.
Co-Authored-By: Claude Opus 4.6 (1M context) noreply@anthropic.com
Co-authored-by: Abir Abbas abirabbas1998@gmail.com Co-authored-by: Claude Opus 4.6 (1M context) noreply@anthropic.com (72ec100)
Feat: Agentic API Layer — Smart 404, Discovery, Query, Knowledge Base
Transform the control plane from a pure orchestration server into a knowledge-serving platform for AI agent consumers.
Phase 1: API Catalog + Smart 404
Phase 2: Operator + Autonomous Endpoints
Phase 3: Knowledge Base API (public, no auth)
Closes #275 #276 #277 #278 #279 #280 Epic: #282
3 test suites, covering:
All tests use httptest + gin test mode — no external dependencies.
401 now includes a help block pointing agents to:
Smart 404 help block now references:
Discover response includes see_also pointing to
/discovery/capabilities (live agents) and KB
An unauthenticated agent now gets immediate actionable guidance instead of a dead-end error.
The 401 response now includes a help object (map[string]string)
alongside the string fields. The test was unmarshaling into
map[string]string which fails on nested objects.
Changed to map[string]interface{} and added assertion that help
field is present in the 401 response. (ee32d1b)
Nothing published for this version
Your coding agent can read these notes before it upgrades. Set up the MCP server →