NewYour coding agent can read the release notes before it upgrades.Set up the MCP server →
PyPI · #1619 most downloaded on PyPI
Framework for large language model evaluations
Last release 13 days ago
04 Sep 2026
Ships on a steady schedule
a new release about every 1 weeks
Nearly every release is documented
notes for 60 of the last 60 stable releases
2 versions withdrawn
withdrawn after publishing
2 years old
241 releases · first in 2024
MCP: Support for mcp_server_http() (which replaces the deprecated SSE server mode).
timeout generation config option through to API Client.GOOGLE_VERTEX_BASE_URL.background, safety_identifier and prompt_cache_key custom model args (bump required version of openai package to v1.98).client_timeout to 900s when flex processing is enabled.reasoning_effort option to reasoning dict.mcp_server_http() (which replaces the deprecated SSE server mode).authorization to provide OAuth Bearer token for HTTP based servers.list[str] rather than str without crashing.message_limit to 50 when both message_limit and token_limit are None.with pytest.raises, add test for env vars.with pytest.raises, add test for env vars.MockLLM.ANSWER: None of the above. Allow answers to end with full stop (.).inspect_tool_support to use a Unix socket rather than a tcp port for intra-container RPC.background() task is now scoped to the sample lifetime in the presence of retry_on_error.waiting_time from within coroutines spawned from the main sample coroutine.inspect-tool-support reference container to support executing tool code with non-root accounts.reasoning_effort and reasoning_tokens for OpenRouter provider.bridge() no longer causes a recursion error when running a large number of samples with openai modelsmodel_roles are available within task initialization code.One column per month.
Bugfix: Conform to breaking changes in openai package (1.99.2).
reasoning_effort.openai package (1.99.2).sample_shuffle is None (rather than 0) when not specified on the command line.Analysis functions are out of beta (inspect_ai.analysis.beta is deprecated in favor of inspect_ai.analysis).
inspect_ai.analysis.beta is deprecated in favor of inspect_ai.analysis).store for scorers run on existing log files.Remove support for vertex provider as the google-cloud-aiplatform package has deprecated its support for Vertex generative models. Vertex can still be…
vertex provider as the google-cloud-aiplatform package has deprecated its support for Vertex generative models. Vertex can still be used via the native google and anthropic providers.emulate_tools model arg) to OpenAI API compatible providers.--sample-shuffle eval option to control sample shuffling (takes an optional seed for determinism).Added Fireworks AI model provider.
user and http_client custom model arguments.is_mistral model arg for mistral compatible tool calling.hidden_states model arg to get model activations.--max-connections, --max-retries, and --timeout now provide defaults for all models rather than only the main model being evaluated.max_tool_output.files field.log_viewer() data preparation function.full option to samples_df() for reading full sample metadata.EvalConfig column defs to EvalConfiguration._repr_ for EvalLog (print JSON representation of log header).metadata_as() typesafe metadata accessor to ChatMessageBase.model_info() operation.Added display_name property to Task (e.g. for plotting).
display_name property to Task (e.g. for plotting).task_info() operation for data frame preparation.Analysis: model_info() and frontier() operations for data frame preparation.
model_info() and frontier() operations for data frame preparation.ThinkChunk and AudioChunk in package v1.9.3 (which is now the minimum required version).<think> and <internal> tags from tool messages to prevent leakage in multi-agent scenarios where an inner assistant message can be coerced into a tool message.BaseModel types in tool call schemas.OpenAI: Move model classification functions into ModelAPI class so that subclasses can override them.
ModelAPI class so that subclasses can override them.prepare() function for doing common data preparation tasks and log_viewer() operation for adding log viewer URLs to data frames.uuid field (or event_id in analysis data frames).inspect log convertBatch processing API support for OpenAI and Anthropic models.
HookedTransformer models with Inspect.thought=True on content when replaying ContentReasoning back to the model.uuid field and metadata field to Event.message_id field to ToolEvent for corresponding ChatMessageTool.uuid in read_eval_log_sample().keep_in_messages option to AgentSubmit to preserve calls to submit() in message history.Value type to use covariant types (Mapping and Sequence).display parameter to score() to control display type.scored_samples and unscored_samples fields to indicate how many samples were scored and how many were not. The viewer will display these values if there are unscored samples.--resolve-attachments option to inspect log dump.EvalSample (rather than only the summary) to on_sample_end() hook.inspect view bundle.claude-3-7-sonnet-latest.AgentState output field.run_in_background allowing it to properly function outside the context of a task.None out TaskLogger's SampleBufferDatabase after cleaning it up to avoid crashing on subsequent logging attempts.Bugfix: Conform to breaking changes in mistralai package (1.9.1).
--sample_id filter for datasets.score_headline_stderr field in standard evals column definitions.task_name without package namespace by default.order field in messages_df() and events_df().run_samples option to disable running samples (resulting in a log file with status "started" and no samples).--display=log (improved task info formatting, ability to disable rich logging)mistralai package (1.9.1).Inspect View: Fix issue with tab switching when running in VS Code.
Bugfix: Return inner exception from run_sample.
run_sample.MCP: Conform to breaking changes in latest mcp package (1.10.0).
web_browser() and bash_session() to indicate that you must pass an instance explicitly to get distinct processes.Bugfix: Don't raise error on Anthropic cited_text not being a str.
str.Bugfix: Shield critical shutdown code from cancel scope.
OpenAI: Use prefix matching when detecting compatible models for web_search().
web_search().executed_tools field as model output metadata.str returned from on_continue to the model (formerly this was only done if there were no tool calls).background() function for executing work in the background of the current sample.
sample.limit (as opposed to ones bound to context managers or agents).text_editor tool; don't use native tool definition for Haiku 3.5 or Opus 3.0.ContentReasoning blocks for Gemini (regressed in v0.3.104).seed of 0.Web Search: Added provider for Anthropic's internal web search tool.
metadata field for arbitrary additional metadata.ContentData for model specific content blocks.Citation suite of types and included citations in ContentText (supported for OpenAI and Anthropic models).task_args now includes defaulted args (formerly it only included explicitly passed args).retry_connections now defaults to 1.0 (resulting in no reduction in connections across passes).
OpenAI: Work around OpenAI Responses API issue by filtering out leading consecutive reasoning blocks.- with _ when looking up provider environment variables..devcontainer) configuration.trim_messages() now removes any trailing assistant message after compaction.responses_api=False when pass as an OpenAI model config arg.logs.json summary file during bundling.Eval set: Do not read full eval logs into memory at task completion.
OpenAI: Use responses API for codex models.
Eval set: Default max_tasks to the greater of 4 and the number of models being evaluated.
max_tasks to the greater of 4 and the number of models being evaluated.time_limit() and working_limit() context managers for scoped application of time limits.
docker compose concurrency to 2 * os.cpu_count() by default (override with INSPECT_DOCKER_CLI_CONCURRENCY).on_continue message to the model if the model made no tool calls.Enum types in tool arguments.plain mode (no outline, don't expand tables to console width).Exported view() function for running Inspect View from Python.
view() function for running Inspect View from Python.eval() or eval_set().google-genai to 1.16.1 (which includes support for reasoning summaries and is now compatible with the trio async backend).Google: Disable reasoning when reasoning_tokens is set to 0.
reasoning_tokens is set to 0.React agent: Use of submit() tool is now optional.
submit() tool is now optional.is_agent() typeguard function for checking whether an object is an Agent.temperature, top_p, and top_k).tools or tool_choice in requests when emulating tool calling (avoiding a 400 error).<tool_calls> plural from Llama models (as it sometimes uses this instead of <tool_call>).--sample-id is not found in the target dataset (raise error if there are no matches at all).EvalLog and Exception in ColumnError.inspect view bundle, support linking to individual transcript events or messages.Dataframes: events_df() function, improved message reading, log filtering, don't re-sort passed logs
events_df() function, improved message reading, log filtering, don't re-sort passed logsmcp package.submit().Dataframe functions for reading dataframes from log files.
max_tokens option to control tokens used for generate().working_limit after solvers have completed executing (matching behavior of other sample limits).user parameter on to sandboxes if is not None (eases compatibility with older sandbox providers).type in the error message body is "overloaded_error".request() method in v1.78.0 of openai package (now the minimum required version).mcp package (now the minimum required version).input_text and user_prompt properties now read the last rather than first user message.metadata from TaskState rather than from Sample.thought argument being passed to think() tool.text_editor() tool for Claude Sonnet 3.5.span() function for grouping transcript events.
inspect log convert now always fully re-writes log files even of the same format (so that e.g. sample summaries always exist in the converted logs).answer_only and answer_delimiter to control how submitted answers are reflected in the assistant message content.bash() and python() tools.Scoped Limits for enforcing token and message limits using a context manager.
bash_session() tool to provide richer interface to model and to support interactive sessions (e.g. logging in to a remote server).vllm package (improved concurrency and resource utilization).streaming model argument to control whether streaming API is used (by default, streams when using extended thinking).--sample-id option can now include task prefixes (e.g. --sample-id=popularity:10,security:5)).--no-log-realtime option for disabling realtime event logging (live viewing of logs is disabled when this is specified)._resources directories from package (reduces pressure on path lengths for Windows).OpenAI: In responses API, don't pass back assistant output that wasn't part of the output included in the server response (e.g. output generated from
submit() tool).Support for using tools from Model Context Protocol providers.
responses_store model argument to control whether the store option is enabled (it is enabled by default for reasoning models to support reasoning playback).python) for responses API.google-genai (which is now required).reasoning_tokens option for Gemini 2.5 models.reasoning_effort option and capturing reasoning content.reasoning_effort and reasoning_tokens to reasoning field.ToolSource for dynamic tools inputs (can be used in calls to model.generate() and execute_tools())submit() tool.user parameter for running the human agent cli as a given user.model_graded_qa() and model_graded_fact().match() scorer.sample_metadata_as() method to SampleScore.user parameter to connection() method for getting connection details for a given user.datetime, date, time, and Set)Path.--traceback-locals CLI option to print values of local variables in tracebacks.CUDA_VISIBLE_DEVICES environment variable for vLLM provider<think> tag in conversation view if there is no reasoning content.Inspect View: Collapse user messages after 15 lines by default.
Model Roles for creating aliases to models used in a task (e.g. "grader", "red_team", "blue_team", etc.)
on_continue message, including using a dynamic name for the submit tool.metadata field to bridge input for backward compatibility with solver-based bridge.default argument to get_model() to explicitly specify a fallback model if the specified model isn't found.history argument (rather than TaskState) to better handle agent conversation state.2025-03-01-preview as default API version if none explicitly specified.trim_messages() function for pruning messages to fit within model context windows.registry_create() function for dynamic creation of registry objects (e.g. @task, @solver, etc.).chdir option from @task (tasks can no longer change their working directory during execution).INSPECT_EVAL_LOG_FILE_PATTERN environment variable for setting the eval log file pattern.openai/azure/model-name).__future__ as needed.dict keys that are numeric in store diffs.Tools: Restore formerly required (but now deprecated) type field to ToolCall.
type field to ToolCall.reasoning_tokens in total_tokens (they are already included).Eval: Fix an error when attempting to display realtime metrics for an evaluation.
Open AI: Treat UnprocessableEntityError as bad request so we can include the request payload in the error message.
UnprocessableEntityError as bad request so we can include the request payload in the error message.Remove support for goodfire model provider (dependency conflicts).
goodfire model provider (dependency conflicts).description without name.Bugfix: Suppress link click behavior in vscode links.
Model API: New execute_tools() function (replaces deprecated call_tools() function) which handles agent handoffs that occur during tool calling.
submit_append option to append the submit tool output to the completion rather than replacing the completion (note that the new react() agent appends by default).call_tools() function) which handles agent handoffs that occur during tool calling.generate_loop() method for calling generate with a tool use loop.Model (works only with providers that don't require an async close).tool_choice="none" (added in v0.49.0, which is now required).logprobs to pass 1 rather than True (protocol change).bash_session() and web_browser() now create a distinct sandbox process each time they are instantiated.openai/computer-use-preview)task_with() and tool_with() no longer copy the input task or tool (rather, they modify it in place and return it).name (save registry_name separately).instance option for store_as() for using multiple instances of a StoreModel within a sample.StoreModel within another StoreModel.sandbox_default() context manager for temporarily changing the default sandbox.write_file() function now gracefully handles larger input file sizes (was failing on files > 2MB).user and assistant solver events.score() function within Jupyter notebooks.eval() when a cancellation occurs.api_version model argument for OpenAI on Azure.None passed to tool call by model for optional parameters..compose.yml when not in working directory.Bugfix: Correct handling of backward compatibility for inspect-web-browser-tool image.
max_tasks is greater than total tasksRequirements: Temporarily upper-bound rich to < 14.0.0 to workaround issue.
rich to < 14.0.0 to workaround issue.Google: Compatibility with httpx client in google-genai >= 1.8.0 (which is now required).
google-genai >= 1.8.0 (which is now required).mistralai >= v1.6.0 (which is now required).Google: Compatibility with v1.7 of google-genai package (create client per-generate request)
OpenAI: Ensure that assistant messages always have the msg_ prefix in responses API.
msg_ prefix in responses API.New think() tool that provides models with the ability to include an additional thinking step.
openai/azure/gpt-4o-mini or mistral/azure/Mistral-Large-2411).--metadata option to eval for associating metadata with eval runs.think.bash_session() tool for creating a stateful bash shell that retains its state across calls from the model.
reasoning_effort parameter to o1-preview (as it is not supported).Model API: Specifying a default model (e.g. --model) is no longer required (as some evals have no model or use get_model() for model access).
--model) is no longer required (as some evals have no model or use get_model() for model access).model, and model is no longer a required axis for parallel tasks.id for ChatMessage when deserialising (id is now str | None and is only populated when messages are directly created).reasoning_tokens for standard thinking blocks (redacted thinking not counted).APIError status codes for retry.--env option for defining environment variables for the duration of the inspect process.inspect view and viewing the log in a browser.store_as()) from log.solver list to eval() (decorate chain function with @solver).Bugfix: Exclude chat message id from cache key (fixes regression in model output caching).
id from cache key (fixes regression in model output caching).The ModelAPI class now has a should_retry() method that replaces the deprecated is_rate_limit() method.
ModelAPI class now has a should_retry() method that replaces the deprecated is_rate_limit() method.generate().inspect trace http command which will show all HTTP requests for a run.max_retries and timeout configuration options. These options now exclusively control Inspect's outer retry handler; model providers use their default behaviour for the inner request, which is typically 2-4 retries and a service-appropriate timeout.tool_choice option.ChatMessage now includes an id field (defaults to auto-generated uuid).mistralai is now required).@task(chdir=True) to preserve the previous behavior.parallel_tool_calls correctly for OpenAI models served through Azure.Computer: Updated tool definition to match improvements in Claude Sonnet 3.7.
Anthropic: Support for extended thinking features of Claude Sonnet 3.7 (minimum version of anthropic package bumped to 0.47.1).
anthropic package bumped to 0.47.1).ContentReasoning type for representing model reasoning blocks.reasoning_tokens for setting maximum reasoning tokens (currently only supported by Claude Sonnet 3.7)reasoning_history can now be specified as "none", "all", "last", or "auto" (which yields a provider specific recommended default).reasoning_tokens when reported.None for assistant content (can happen when there is a refusal).working_start attribute to events and completed and working_time to model, tool, and subtask events.task quit command for giving up on tasks.TimeoutError for running shell commands in the computer tool container.working_limit option for specifying a maximum working time (e.g. model generation, tool calls, etc.) for samples.
SandboxEvent to transcript for recording sandbox execution and I/O.as_type() function for checked downcasting of SandboxEnvironmentstate.completed=True when entering scoring (basic_agent() no longer sets completed so can be used in longer compositions of solvers).uuid property to TaskState and EvalSample (globally unique identifier for sample run).cleanup to tasks for executing a function at the end of each sample run.bridge() is now compatible with the use of a custom OPENAI_BASE_URL.mistralai package to 1.5 (required for working_limit).fsspec version constraints to avoid pip errors when installing alongside datasets.dict as a parameter.sample_id for --no-sandbox-cleanup.Google provider updated to use the Google Gen AI SDK, which is now the recommended API for Gemini 2.0 models.
compose up when no services startup successfully.--no-sandbox-cleanupTimeoutError for subprocess timeouts (ensure kill/cleanup of timed out process).Task display: Improve spacing/layout of final task display.
Memoize calls to get_model() so that model instances with the same parameters are cached and re-used (pass memoize=False to disable).
get_model() so that model instances with the same parameters are cached and re-used (pass memoize=False to disable).Model class for optional scoped usage of model clients.assistant_message() solver.prompt_template(), and system/user/assistant solvers.Docker: Correct compose file generation for Dockerfiles w/ custom stem or extension.
Compatibility with textual 2.0 (which had several breaking changes).
null values.Metrics now take list[SampleScore] rather than list[Score] (previous signature is deprecated but still works with a warning).
cluster parameter for the stderr() metric.inspect score command and score() function).list[SampleScore] rather than list[Score] (previous signature is deprecated but still works with a warning).var() metric.sandbox argument for running in non-default sandboxes.ScoreEvent (with intermediate=True) when the score() function is called.source field to InfoEvent and use it for events logged by the human agent..Dockerfile extension.container_name (incompatible with epochs > 1).compose up timeout when there are healthcheck entries for services.log_dir is writeable at startup.None).Add shuffle_choices to dataset and dataset loading functions. Deprecate shuffle parameter to the multiple_choice solver.
shuffle_choices to dataset and dataset loading functions. Deprecate shuffle parameter to the multiple_choice solver.stop_words param to the f1 scorer. stop_words will be removed from the target and answer during normalization.inspect_ai.tool.beta into inspect_ai.tool).tee for write_file operations.type parameter of answer() to pattern to address registry serialisation error.Various improvements for reasoning models including extracting reasoning content from assistant messages.
reasoning_effort, max_tokens, temperature, and parallel_tool_calls correctly for o3 models.content_filter stop reason.model_length StopReason.Your coding agent can read these notes before it upgrades. Set up the MCP server →