NewYour coding agent can read the release notes before it upgrades.Set up the MCP server →
PyPI · #1619 most downloaded on PyPI
Framework for large language model evaluations
Last release today
17 Sep 2026
Ships on a steady schedule
a new release about every 1 weeks
Nearly every release is documented
notes for 60 of the last 60 stable releases
2 versions withdrawn
withdrawn after publishing
2 years old
243 releases · first in 2024
Eval Logs: Ensure that condense_events() is called when re-writing eval logs.
condense_events() is called when re-writing eval logs.NOANSWER instead of INCORRECT when the pattern scorer fails to match.Timelines: Improved detection of forked_at for branches from user or system messages.
forked_at for branches from user or system messages.One column per month.
OpenAI: Add cyber_policy to "content_filter" stop reason
cyber_policy to "content_filter" stop reasonBranchEvent and timeline_branch() to delineate timeline branches.TimelineBranch into TimelineSpan via forked_at property.Google: Update to google-genai v1.69.0 to address type changes (async_http_client can now be None for Vertex with Google Auth).
google-genai v1.69.0 to address type changes (async_http_client can now be None for Vertex with Google Auth).read_approval_policies() function for reading approval policies from a config file.metadata field to Approval which is in turn forwarded to ApprovalEvent.parse_tool_info() to improve performance when there are many tools defined.required field to get_model() for ensuring that model roles are specified.model_roles() function to get model roles for the active task.forked_at detection for forking on non-assistant messages.src/inspect_ai/_view/www/ into ts-mono/apps/inspect/ monorepo (pnpm + Vite + Jest). Built assets are copied to src/inspect_ai/_view/dist/ via a Vite plugin. No user-facing changes.Google: Remove deprecated gemini-3-pro-preview from computer use model check and replace with gemini-3.1-pro-preview in tests and docs.
gemini-3-pro-preview from computer use model check and replace with gemini-3.1-pro-preview in tests and docs.completion_mode for CPT/base models, sending completions-style request payloads with logprobs and prompt_logprobs support.read_timeout and connect_timeout model args.do_sample model arg for overriding default sampling behavior.client_timeout to OpenAICompatibleAPI and VLLMAPI.- (e.g. -0.07") by using the = form for the --text argument.embed_viewer=True, and keep listing.json updated as logs are created.run_multiple silently swallowing task finalisation errors and returning success=True with no results.tool_calls and source when combining assistant messages.math.inf so it never blocks.Google: Fix intermittent FAILED_PRECONDITION error when using native code execution by omitting function calling system instruction hint.
exec_remote to prevent OOM from unbounded subprocess output.timeout and timeout_retry for exec_remote() requests.bind()/listen() in sandbox tools server.approval() context manager and approval arguments to execute_tools() and react() agent.ValueError when sandbox provider is not found.OpenAI: Store readable reasoning text in summary when both text and encrypted reasoning content are provided.
summary when both text and encrypted reasoning content are provided.<think> tags that appear in assistant messages.OpenAI: Support for image output for multimodal modals.
reasoning_effort for Nova models.CompactionEdit and CompactionNative strategies.as_tool() agent wrapper.samples and reductions fields lazily in EvalLog returned by eval().edit_eval_log().tags parameter to Task, merged with eval-level tags at eval time.condense_events/expand_events API pair and EventsData TypedDict for deduplicating and restoring repeated model event inputs and call messages.--max-dataset-memory option to limit the size of datasets held in memory during execution. When exceeded, samples are paged to disk.viewer/ subdirectory, fixing permission issues when serving logs.NotADirectoryError when locating sandbox tools binary so S3 download/build fallbacks run on Python < 3.13.staticx incompatibility with setuptools 82+ (removed pkg_resources).OpenAI: Don't serialize unspecified fields in ResponseCustomToolCallParam.
ResponseCustomToolCallParam.Anthropic: Preserve OAuth beta header when per-request betas are set.
agent_result for timeline spans.exec_remote() now auto-injects sandbox tools CLI if needed.on_sample_event() hook.upload_fileobj for automatic multipart uploads.absolute_file_path (matches local fs behavior).inspect view embed command and --embed-viewer eval-set option to embed a log viewer into a log directory.OpenAI: Fix mypy issues with OpenAI SDK 2.26 (now required).
Anthropic: Handle updated Anthropic compaction not supported error message.
frozendict as fallback; update jsonpath-ng dependency.<summary> from <details> tags in tool views.anonymous and region_name parameters to support credential-free access to public S3 buckets..eval file sizes.Anthropic: Use text_editor_20250728 for all Claude 4.x models per Anthropic docs.
text_editor_20250728 for all Claude 4.x models per Anthropic docs.agent_span_id property to tool events for associating them with their associated agent.Improved naming and typeinfo for exec_remote() output stream events.
exec_remote() output stream events.Events: Rename EventNode to EventTreeNode and SpanNode to EventTreeSpan (old type names will still work at runtime with a deprecation warning).
anthropic package is now required).samples_df() (now 50x faster).config_deserialize().exec_remote() method for async execution of long-running commands.type field to CompactionEvent to record compaction type.cost_limit() context manager for scoped application of cost limits.EventNode to EventTreeNode and SpanNode to EventTreeSpan (old type names will still work at runtime with a deprecation warning).thinking_config) under generation_config in the REST schema.EXPIRED or PARTIALLY_SUCCEEDED state.Bugfix: Fix shutdown hang by draining nest_asyncio event loop.
Inspect View: Fix a regression which affected the display of samples within VSCode.
Added stable_message_ids() function for yielding stable ids based on model content (but always unique within a given conversation).
stable_message_ids() function for yielding stable ids based on model content (but always unique within a given conversation).CompactionEvent for model compactions.<think> tags when included in user messages.--reasoning-history CLI argument (don't parse as boolean).human_cli() with no answer now correctly completes task.Web Search: Fallback to Google CSE provider only when Google CSE environment variables are defined (the CSE service has been deprecated by Google).
<think> tags in assistant message loading (now all done directly by model providers).INSPECT_USE_ZSTD environment variable.Early Stopping: Check for early stopping after sample semaphore is acquired rather than before.
json.dumps for message cache keys (incompatible with BaseModel types).Eval Logs: Improve load time by using JSON in duplicate message cache rather than frozendict.
frozendict.trim_message() to use the same behavior).Google: Provide JSON schema directly rather than converting it to Google Schema type.
ContentReasoning as <think> with attributes to prevent bridge clients from doing a more lossy <think> tag conversion.count_tokens() method.count_tokens() method.yyyy-mm-dd hh:mm:ss format (keep local time zone).Anthropic: Only re-order reasoning blocks for Claude 3 (as we use interleaved thinking for Claude 4).
samples_df().Google: Add streaming model arg to opt-in to streaming generation.
streaming model arg to opt-in to streaming generation.image_input (data URI) in field spec for multimodal tasksComposeConfig directly to Docker sandbox provider.supported_fields parameter from parse_compose_yaml() (packages handle their own validation).Sandbox: parse_compose_yaml() for parsing Docker Compose files into typed configuration for sandbox providers.
parse_compose_yaml() for parsing Docker Compose files into typed configuration for sandbox providers.strict=True. This is required by HF Inference Providers and Fireworks (and possibly others).task status and task stop.Agent Bridge: Consolidate bridged tools implementation into the existing sandbox model proxy service (eliminate Python requirement for using bridged t
{} as value for additionalProperties in tool schema.background parameter as this is OpenAI service-specific.minimal and medium reasoning effort levels for Gemini 3 Flash.max_tokens is greater than 16000.combined_from metadata field when combining consecutive user or assistant messages for call to generate.kwargs to JSON readers (built-in reader and jsonlines reader).async_connection()builtins module rather than __builtins__ when parsing tool function types.Compaction: Compacting message histories for long-running agents that exceed the context window.
count_tokens() method for estimating token usage for messages.ModelInfo for retrieving information about models (e.g. organization, context window, reasoning, release date, etc.)reasoning_details (map onto standard reasoning fields for viewer).IO[bytes] via read_eval_log().shell_command tool calls.skill() tool to make agent skills available to models.
Eval Set: Correct log reuse behavior when epochs and limit change.
RuntimeError.Anthropic: Treat reasoning text as a summary (true for all models after Sonnet 3.7).
reasoning_details field to forward native reasoning replay to models.summary in serialization for agent bridge.@agent functions with no return type decoration.retry_refusals option to retry on stop_reason == "content_filter".choices in EvalSampleSummary.inspect view bundle to publish hf/ prefixed targets to Hugging Face Spaces.Log File column of the samples list.type field to enable lists of types.Eval Set: Defer reading eval samples until they are actually needed (prevents memory overload for large logs being retried).
anthropic/azure).-M streaming=true).Early Stopping API for ending tasks early based on previously scored samples.
az://).cleanup() function at the end of the sample (after scoring) rather than after solvers.disable_retry) for waiting time tracking.--verbosity generation option.reasoning_effort option.temperature, top_p, and logprobs for GPT 5.x models with reasoning disabled.-M stream=false to disable).stream option (disabled by default, use -M stream=true to enable).model option is now used only as a fallback if the request model is not for "inspect" or "inspect/*".metadata field to new eval for eval-retry.Agent Bridge: Don't print serialization warnings when going from Pydantic -> JSON (as we use beta types that can cause warnings even though serializat
Update Plan tool for tracking steps and progress across longer horizon tasks.
--effort) for trading off between response thoroughness and token efficiency.web_fetch tool as part of web_search() implementation (matching capability of other providers that have native web search).caller field for server tool uses (required by package version 0.75, which is now the minimum version).-M conversation_api=False).retry_delay).BridgedToolsSpec and bridged_tools parameter to sandbox_agent_bridge() for exposing host-side Inspect tools to sandboxed agents via MCP protocol.EvalLog objects directly to dataframe functions (samples_df(), evals_df(), messages_df(), events_df()).mcp package version 1.23.0.Memory tool: Added memory() tool and bound it to native definitions for providers that support it (currently only Anthropic).
xai_sdk, which is now required).output_limit as a cap enforced with a circular buffer (rather than a limit that results in killing the process and raising).evals_in_eval example for running Inspect evaluations inside other evaluations.datetime's via DTZ lint rule.nest_asyncio, which is fundamentally incompatible with Python 3.14, to nest_asyncio2, which has explicit 3.14 compatibility.read_eval_log() with JSON log files.Anthropic: Enable interleaved-thinking by default for Claude 4 models.
max_tokens handling to prevent exceeding model max tokens when reasoning tokens are specified.RegistryInfo and registry_info() to the public API.prompt_cache_retention is correctly forwarded by agent bridge to responses API.Inspect View: Truncate display of large sample summary fields to improve performance.
Bugfix: Fix Google provider serialization of thought signatures on replay.
Added cache configuration to GenerateConfig (formerly was only available as a parameter to generate()).
cache configuration to GenerateConfig (formerly was only available as a parameter to generate()).on_continue can now return a new AgentState.APIConnectionError.reasoning_effort="none" (now available with gpt-5.1).None as value of arguments when parsing tool calls.OpenAI: Show reasoning summaries by default (auto-detect whether current account is capable of reasoning summaries and fallback as required).
logprobs and top_logprobs in Responses API (note that logprobs are not supported for reasoning models).xai_sdk package (rather than using OpenAI compatible endpoint).web_search() tool.reasoning_enabled model arg to optionally disable reasoning for hybrid models.eval_set_idcontent of type str in Responses API agent bridge.Eval Set: Task identifiers can now vary on GenerateConfig and solver (which enables sweeping over these variables).
GenerateConfig and solver (which enables sweeping over these variables).ModelEvent.call by default, which prevents O(N) memory footprint for reading transcripts.ResponseOutputText in v2.7.0 of openai package.reasoning_effort parameter (only supported by grok-3-mini and only low and highvalues are supported).Google: Correct capture and playback of thought_signature in ContentReasoning blocks.
thought_signature in ContentReasoning blocks.@scanner functions as scorers.output during execution (only condense call).logs in data frame functions.Google: Distribute citations from web search to individual ContentText parts (rather than concatenating into a single part).
Tests: Skip git revision detection and realtime logging during pytest runs to improve test performance.
OpenAI: Handle Message input types that have no "type" field in responses API.
Message input types that have no "type" field in responses API.override_api_key() hook.dict as well as list for underlying data.edit_score() silently editing only first epoch in multi-epoch evaluations (now requires explicit epoch parameter).Added model API for Hugging Face Inference Providers.
messages_to_openai_responses() function.gpt-5-pro by default.parallel_tool_calls option for tool choice.logprobs and top_logprobs.model_length stop reason based on additional error pattern.content can never be empty.edit_score() and recompute_metrics() functions for modifying evaluation scores with provenance tracking and metric recomputation.attempt_timeout to GenerateConfig (governs timeout for individual attempts and still retries if timeout is exceeded).ModelEvent immediately to prevent O(N) memory usage for long message histories.inspect view server using uvicorn / fastapi.Union for TypeAlias (required by Python 3.14).umask that led to permission problems with bash_session tool.list and dict parameters (rather use None and initialize on use).SubtaskEvent.input values that aren't of the required dict type.OpenAI: Support for tool calls returning images (requires v2.0 of openai package, which is now required).
openai package, which is now required).run() function.inspect_ai.event module with event related tyeps and functions.rich > 13.3.3 save for 14.0.0 (which had an infinite recursion bug affecting stack traces with exception groups).resolve_attachments in score command when stream=True.fail_on_error=False.Google: Manage Google client lifetime to scope of call to generate().
generate().inspect score command (and inspect_score function) when scoring log files on S3.OpenAI: Capture reasoning summaries even when there is encrypted reasoning content.
Bugfix: Update various tests to react to Google's deprecation of old models.
retry_refusals loop.detail along with images.--num-choices is specified.SampleScores column group for extracting score answer, explanation, and metadata.inspect-ai package installation type detection code.update_plan and shell tool inputs.EvalSet when optional values are missing.Sandbox tools: bash_session, text_editor, and sandbox MCP servers no longer require a separate pipx install (they are now automatically injected into
retry_refusals option for automatically retrying refusals a set number of times.convert_eval_logs().convert_eval_logs().model_graded_qa(), model_graded_fact()) now look for the "grader" model-role by default.on_sample_scoring() and on_model_cache_usage() hooks.on_run_end() even when the eval is cancelled.None to indicate that they did not score the sample. Such samples are excluded from reductions and metrics.Sequence and Mapping types for metrics on scorer decorator.async_score in eval score.Anthropic: Support for images with mime type image/bmp.
store=False from bridge client and don't insist on id being included with reasoning (as it is not returned in store=False mode).attachments:/ appearing in content when viewing running samples.OpenAI: Correct serialization of web search tool calls (prevent 400 errors).
Agent Bridge: Option to force the sandbox agent bridge to use a specific model.
filter option to enable bridge to filter model generate calls.responses_api model arg.eval_set_id to log file (unique id for eval set across invocations for the same log_dir).log_format as the log which is being retried.EvalSetStart and EvalSetEnd hook methods.inspect score now supports streaming via the --stream argument.Agent Bridge: Don't use concurrency() for agent bridge interactions (not required for long-running proxy server or cheap polling requests).
concurrency() for agent bridge interactions (not required for long-running proxy server or cheap polling requests).concurrency parameter to exec() to advise whether the execution should be subject to local process concurrency limits.Bugfix: Work around OpenAI breaking change that renamed "find" web search action to "find_in_page" (bump required version of openai package to v1.104.…
as_solver()).openai v1.104.1 (which is now the minimum required version).ThinkChunk types in mistralai v1.9.10.--reasoning-effort parameter (works w/ gpt-oss models).str_to_float() fails.openai package to v1.104.0).Bugfix: Preserve sample list state (e.g. scroll position, selection) across sample open/close.
Agent Bridge: OpenAI Responses API and Anthropic API are now supported alongside the OpenAI Completions API for both in-process and sandbox-based agen
AgentState changes via inspecting model traffic running over the bridge.messages_df().GenerateConfig values for models override bridged agent config.handoff(): Use content_only() filter by default for handoff output and improve detection of new content from handed off to agents.ContentToolUse ("web_search" or "mcp_call")internal field from ChatMessageBase (no longer used).responses_store model arg for explicitly enabling or disabling the responses API.enum typed fields.thought_signature for thought parts.MCPServerConfig and Stdio and HTTP variants).instance option for multiple services of the same type in a single container.polling_interval option for controlling polling interval from sandbox to scaffold (defaults to 2 seconds, overridden to 0.2 seconds for Docker sandbox).completion).--no-epochs-reducer CLI flag for specifying no reducers.match more lenient when numeric matches container markdown formatting.visible option for concurrency() contexts to control display in status bar.sample_init, sandbox, store, and state events.Scoring: Refactor inspect score to call same underlying code as score().
inspect score to call same underlying code as score().Agent Bridge: New context-manager based agent_bridge() that replaces the deprecated bridge() function.
agent_bridge() that replaces the deprecated bridge() function.sandbox_agent_bridge() to integrate with CLI based agents running inside sandboxes.user_prompt() function for getting the last user message from a list of messages.messages_to_openai() and messages_from_openai() functions for converting to and from OpenAI-style message dicts.response_schema option for providing a JSON schema for model output.--log-dir-allow-dirty option.--continue-on-fail option for eval() and eval_set().copy option to score_async() (defaults to True) to control whether the log is deep copied before scoring.inspect_tool_support.eval-retry by delaying their creation until after task creation.CitationBase so that it no longer leads to an invalid schema.background to OpenAI Responses if specified.tool_choice to Anthropic thinking models.Your coding agent can read these notes before it upgrades. Set up the MCP server →