NewYour coding agent can read the release notes before it upgrades.Set up the MCP server →
PyPI · #1619 most downloaded on PyPI
Framework for large language model evaluations
Last release 2 days ago
19 Sep 2026
Ships on a steady schedule
a new release about every 1 weeks
Nearly every release is documented
notes for 60 of the last 60 stable releases
2 versions withdrawn
withdrawn after publishing
2 years old
244 releases · first in 2024
Metrics now take list[SampleScore] rather than list[Score] (previous signature is deprecated but still works with a warning).
cluster parameter for the stderr() metric.inspect score command and score() function).list[SampleScore] rather than list[Score] (previous signature is deprecated but still works with a warning).var() metric.sandbox argument for running in non-default sandboxes.ScoreEvent (with intermediate=True) when the score() function is called.source field to InfoEvent and use it for events logged by the human agent..Dockerfile extension.container_name (incompatible with epochs > 1).compose up timeout when there are healthcheck entries for services.log_dir is writeable at startup.None).One column per month.
Add shuffle_choices to dataset and dataset loading functions. Deprecate shuffle parameter to the multiple_choice solver.
shuffle_choices to dataset and dataset loading functions. Deprecate shuffle parameter to the multiple_choice solver.stop_words param to the f1 scorer. stop_words will be removed from the target and answer during normalization.inspect_ai.tool.beta into inspect_ai.tool).tee for write_file operations.type parameter of answer() to pattern to address registry serialisation error.Various improvements for reasoning models including extracting reasoning content from assistant messages.
reasoning_effort, max_tokens, temperature, and parallel_tool_calls correctly for o3 models.content_filter stop reason.model_length StopReason.Computer: Enable viewing computer tool's remote mouse cursor via VNC.
SampleLimitExceededError with current state so that messages, etc. are preserved when limits are hit.incorrect_message can now optionally be an async function.suffix from eval-set CLI args.Exception from sandboxenv_init (allow cancelled to propagate)Agent Bridge for integrating external agent frameworks with Inspect.
@wraps to functions wrapped by Inspect decorators to preserve type information.refusal field from assistant message when provided.openai/azure)anthropic/bedrock, anthropic/vertex).extra_body model arg (for adding additional JSON properties to the request)tools to state so that tools added in init are preserved.Beta version of computer() tool which models with a computer desktop environment.
user_message() solver for appending parameterised user messages.prompt_template(), system_message() and user_message() solver now also include the sample store in substitution parameters.state.completed for limit enforcement).SampleLimitExceededError.model_input option that determines how tool call result content is played back to the model.ToolDef with **kwargs: Any (just pass them through).fail_on_error option)plain mode with periodic updates on progress, metrics, etc.Literal["a", "b", "c"])) in tool function declarations.Print model conversations to terminal with --display=conversation (was formerly --trace, which is now deprecated).
timeout_retry option (defaulting to True) to exec() function.type and optional container properties to SandboxConnection.task_with() function for creating task variants.--filter argument to trace CLI commands for filtering on trace log message content.--display=conversation (was formerly --trace, which is now deprecated).bootstrap_std metric to bootstrap_stderr (deprecate bootstrap_std)Tracing API for custom trace logging.
StopReason)max_completion_tokens option for o1 full.error and web_at fields are present in response.Store from existing dictionary.metadata_as and store_as typed accessors for sample metadata and store.None are now supported.Human Agent solver for human baselining of computing tasks.
Sample store and metadata using Pydantic models.Task level (eval level approval policies take precedence).ContentText and ContentImage.SandboxConnection that contains login information from sandboxes.display_type() function for detecting the current display type (e.g. "full", "rich", etc.)eval() running in multiple processes at once (trace file per-process)docker build and docker pull commands.store.get() not auto-inserting default value.Bedrock: redact authentication model args from eval logs.
temperature is used with o1 models (as it is not supported).Tracing for diagnosing runs with unterminated action (e.g. model calls, docker commands, etc.).
--no-score-display option to disable realtime scoring metrics.logprobs.OpenAI: Support for o1 including native tool calling and reasoning_effort generation option.
reasoning_effort generation option.setup step that always runs even if solver is replaced.model_args passed through to session.Client.jpeg images.max_sandboxes option for (per-provider) maximum number of running sandboxes.max_tasks as a lower bound for max_samples.eval format logs.--sample-id integer comparisons.None)Eval: --sample-id option for evaluating specific sample id(s).
--sample-id option for evaluating specific sample id(s).emulate_tools model arg to force tool emulation (emulation is enabled by default for Llama models).max_tool_output parameter to override default max tool output from generate config.--trace mode.Bugfix: Task display fails to load when no scorers are defined for a task.
Tools: Improved typing/schema support (unions, optional params, enums).
append argument to use_tools() for adding (rather than replacing) the currently available tools.BaseModel types for sandbox config (formerly only a file path could be passed).Enter key).input_panel() API for adding custom panels to the fullscreen task display..eval log reading performance for remote filesystem (eagerly fetch log to local buffer).token_usage property to TaskState which has current total tokens used across all calls to generate() (same value that is used for enforcing token limits).time field to ModelOutput that records total time spent within call to ModelAPI generate().tokenizer_call_args dict to specify custom args during tokenization, such as max_length and truncation.None for content.hf_dataset now explicitly requires the split argument (previously, it would crash when not specified).Logging: Only call CreateBucket on Amazon S3 when the bucket does not already exist.
Realtime display of sample transcripts (including ability to cancel running samples).
EvalLog now includes a location property indicating where it was read from.max_samples across sandbox and non-sandbox evals (both now apply max_samples per task, formerly evals with sandboxes applied max_samples globally).--login option so that e.g. .bashrc is read before executing the command.logprobs and other new 1.5 (002 series) options.cached argument to control whether to use a previously cached version of the dataset if available (defaults to True).revision option to load a specific branch or commit SHA (when using revision datasets are always revalidated on Hugging Face, i.e. cached is ignored).Basic agent: Ensure that the scorer is only run once when max_attempts = 1.
eval is now the default log format (use --log-format=json to use old format).
--log-format=json to use old format).--no-log-images).eval format log files.inspect log convert to convert a single log file.exec() output limit from 1 MiB to 10 MiB.time_limit option for specifying a maximum execution time for samples.
.eval log files.INSPECT_DISABLE_MODEL_API into generate() (as opposed to get_model()).eval files as logs (don't apply file name pattern restrictions as we do with .json).metadata field to ModelOutput and provide various fields for the Groq provider.Revert change to single epoch reducer behavior (regressed some scoring scenarios).
New binary log format which yields substantial size and speed improvements (JSON format log files are still fully supported and utilities for converti
--model-config, --task-config, and --solver-config CLI arguments for specifying model, task, and solver args using a JSON or YAML config file.casefold() for case-insensitive compare in includes(), match(), exact(), and f1() scorers.strict tool calling (sporadically supported across models and we already internally validate).read_eval_log_sample() for JSON log files.SandboxEnvironment.exec() output streams to 1 MiB. Limit SandboxEnvironment.read_file() to 100 MiB.INSPECT_DISABLE_MODEL_API environment variable for disabling all Model APIs save for mockllm.tool_call_id param to ModelOutput.for_tool_call().file_dataset() function.Open log files in binary mode when reading headers (fixes ijson deprecation warning).
--tags option to eval for tagging evaluation runs.eval().include_history option to model graded scorers to optionally include the full chat history in the presented question.delimiter option to csv_dataset() (defaults to ",")list and dict of registry objects when logging plan.model_usage field to EvalSample to record token usage by model for each sample.…used per sample ( message_limit replaces deprecated max_messages).
as_dict() utility method to Scoretoken_limit and message_limit) for capping the number of tokens or messages used per sample ( message_limit replaces deprecated max_messages).metadata field to Task and record in log EvalSpec.--log-level-transcript option for separate control of log entries recorded in the eval log filefail_on_error option for eval_retry() and inspect eval-retry.self_critique() and model_graded_qa().--port and --host on Inspect View.Add interactive option to web_browser() for disabling interactive tools (clicking, typing, and submitting forms).
interactive option to web_browser() for disabling interactive tools (clicking, typing, and submitting forms).basic_agent(), defer to task max_messages if none is specified for the agent (default to 50 is the task does not specify max_messages).content parameter to ModelOutput.for_tool_call().sample_reductions when returning eval logs with header_only=True.The sample transcript will now display the target for scoring in the Score Event (for newly run evaluations).
max_messages on TaskState.max_messages option for basic_agent() (defaulting to 50) and use it rather than any task max_messages defined.Rename web_browser_tools() to web_browser(), and don't export individual web browsing tools.
web_browser_tools() to web_browser(), and don't export individual web browsing tools.parallel option to @tool decorator and specify parallel=False for web browsing tools.</tool_call> end sequence for Llama models.Move evals into inspect_evals package.
Web Browser tool which provides a headless Chromium browser that supports navigation, history, and mouse/keyboard interactions.
auto_id option for dataset readers to assign an auto-incrementing ID to records.Catch o1-preview "invalid_prompt" exception and convert to normal content_filter refusal.
metadata field in the first epoch@task.Support for max_tokens on OpenAI o1 models (map to max_completion_tokens).
max_tokens on OpenAI o1 models (map to max_completion_tokens).inspect viewepochs is less than 1StopReason: Added "model_length" for exceeding token window and renamed "length" to "max_tokens".
fork().--no-ansi or INSPECT_NO_ANSImultiple_choice() and export MultipleChoiceTemplate enumerationx-default to be referred to by their declared service name.max_messages for use of basic_agent() (as without it, the agent could end up in an infinite loop).answer, explanation, and metadata only if equal across all epochs.Fix issue w/ subtasks not getting a fresh store() (regression from introduction of fork() in v0.3.30)
fork() in v0.3.30)web_search() provider environment variables.Deprecated Plan in favor of Solver (with chain() function to compose multiple solvers).
Plan in favor of Solver (with chain() function to compose multiple solvers).max_tool_output generation option (defaults to 16KB).header_only log reading (switch from json-stream to ijson).eval-set (run a single eval then stop).epochs in task status (formerly wasn't included for multiple task display)info() strings as markdownAdded fork() function to fork a TaskState and evaluate it against multiple solvers in parallel.
fork() function to fork a TaskState and evaluate it against multiple solvers in parallel.answer, explanation, and metadata.inspect info log-typescache argument to basic_agent() for specifying cache policy for the agent.cache field to ModelEvent to track cache reads and writes.Added --plan and -P arguments to eval and eval-set commands for replacing the task default plan with another one.
--plan and -P arguments to eval and eval-set commands for replacing the task default plan with another one.eval() or eval_set() with a plan argument.--log-images).epoch_reducer from being used in eval-retry.epoch from overriding task level epoch.basic_agent() that provides a ReAct tool loop with support for retries and encouraging the model to continue if its gives up or gets stuck.
system_message() now supports custom parameters and interpolation of metadata values from Sample.generate() solver now accepts arbitrary generation config params.use_tools() now accepts a variadic list of Tool in addition to literal list[Tool].bash() and python() tools now have a user parameter for choosing an alternate user to run code as.bash() and python() tools now always return stderr and stdout no matter the exit status.debug_errors option to eval() to raise task errors (rather than logging them) so they can be debugged.eval() from a script or notebook.eval_set() issue with cleaning up failed logs on S3.Fix missing timestamp issue with running eval_set() with an S3-backed log directory.
eval_set() with an S3-backed log directory.f1() and exact() scorers.exact() scorer.Eval Sets for running groups of tasks with automatic retries.
f1() (precision and recall in text matching) and exact() (whether normalized text matches exactly).metrics now override built in scorer metrics (previously they were merged). This enables improved re-use of existing scorers where they only change required is a different set of metrics.write_log_dir_manifest() to write a log header manifest for a log directory.store() and @subtask from solver to utils module; relocate transcript() from solver to log module.cwd that are relative paths as relative to sample working directory.--runapi flag is passed to pytest (prevents unintended token usage)chdir option from @tasks (tasks now always chdir during execution)..env files in task directories (all required vars should be specified in the global .env).strict mode for OpenAI tool calls when all function parameters are required.Store for manipulating arbitrary sample state from within solvers and tools.
Store for manipulating arbitrary sample state from within solvers and tools.Transcripts for detailed sample level tracking of model and tool calls, state changes, logging, etc.Subtasks for delegating work to helper models, sub-agents, etc.init value in default Docker compose file so that exit signals are handled correctly (substantially improves container shutdown performance).function field to ChatMessageTool to indicate the name of the function called.Support for tool calling for Llama 3.1 models on Bedrock.
strict mode in OpenAI tool calls (update to v1.40.0 of openai package required).Support for tool calling for Llama 3.1 models on Azure AI and CloudFlare.
max_tokens from 1024 to 2048.--log-images to opt back in).mistralai is now required).Fix issue affecting results of pass_at_{k} score reducer.
pass_at_{k} score reducer.Add pass_at_{k} score reducer to compute the probability of at least 1 correct sample given k epochs.
pass_at_{k} score reducer to compute the probability of at least 1 correct sample given k epochs.value_to_float string conversion (handle numbers, "true", "false", etc.)max_tokens to 4096name parameter with task created from @task decorated function.metadata available in prompt, grading, and self-critique templates.Epochs data type for specifying epochs and reducers together (deprecated epochs_reducer argument).
Epochs data type for specifying epochs and reducers together (deprecated epochs_reducer argument).INSPECT_CACHE_DIR environment variable.prompt attribute of @tool for descriptions.tool_with() function for adapting tools to have varying names and parameter descriptions.@task arguments so that dynamically created tasks can be retried.eval-retry message to terminal for filesystem based tasks.Existing symbols will continue to work but will print deprecation errors.
bootstrap_std metric with stderr for built in scorers (see rationale for details).ToolEnvironment to SandboxEnvironment and tool_environment() to sandbox() (moving the renamed types from inspect_ai.tool to inspect_ai.util). Existing symbols will continue to work but will print deprecation errors.bash(), python(), and web_search() functions from inspect_ai.solver to inspect_ai.tool. Existing symbols will continue to work but will print deprecation errors.chdir option to @task to opt-out of changing the working directory during task execution.ToolInfo parameters to be directly expressed in JSON Schema (making it much easier to pass them to model provider libraries).tool_choice="any" for OpenAI models (requires >= 1.24.0 of openai package).--parallel-tool-calls false.cwd argument to SandboxEnvironment.exec()tee rather than docker cp for Docker sandbox environment implementation of write_file().azure model arg for OpenAI provider to force binding (or not binding) to the Azure OpenAI back-end.setup field to Sample for providing a per-sample setup script.metadata field from samples.TimeoutError if a call to subprocess() or sandbox().exec() times out (formerly a textual error was returned along with a non-zero exit code).example_dataset() (and print available example dataset names).tool_error text for Anthropic tool call responses.…still work fow now, but result in a runtime deprecation warning).
eval().api_key to get_model() for explicitly specifying an API key for a model.network_mode: none for disabling networking by default in Docker tool environments.max_samples (set to 25 for the Docker provider).eval_async() (unsafe because of need to change directories for tasks). Parallel task evaluation will instead be implemented as a top-level feature of eval() and eval_async().inspect_ai.tool module (previous imports still work fow now, but result in a runtime deprecation warning).call_tools() function now operates directly on messages rather than task state.google-generativeai v0.5.3).Optional increased control over the tool use loop via the call_tools() function and new tool_calls parameter for generate().
call_tools() function and new tool_calls parameter for generate().per_epoch option for CachePolicy to allow caching to ignore epochs.choices and files when converting Sample images to base64.Various fixes for the use of Docker tool environments on Windows.
--no-toolenv-cleanup.inspect toolenv cleanup command for manually cleaning up tool environments.ToolError exception type for explicitly raising tool errors to the model. Formerly, any exception would be surfaced as a tool error to the model. Now, the ToolError exception is required for reporting to the model (otherwise other exception types go through the call stack and result in an eval error).INSPECT_LOG_DIR in .env file relative to .env file parent directory.- for delimiting --limit ranges rather than ,.Sandbox Environments for executing tool code in a sandbox.
multiple_choice() solver now has support for questions with multiple correct answers.BadRequestError (400) errors (which were formerly all treated as content moderation errors).TaskState.tools getter/setter (where the setter automatically syncs the system messages to the specified set of tools).use_tools() function now uses the TaskState.tools setter, so replaces the current set of tools entirely rather than appending to it.state.completed = False when max_messages is reached.bytes field in Logprobs and TopLogprobs.truthfulqa benchmark.intercode-ctf example.Stream samples to the evaluation log as they are completed (subject to the new --log-buffer option). Always write completed samples in the case of an
--log-buffer option). Always write completed samples in the case of an error or cancelled task."cancelled" status in eval log for tasks interrupted with SIGINT (e.g. Ctrl-C). Logs are now written for cancellations (previously they were not).--max-samples (maximum concurrent samples) to --max-connections, which will result in samples being more frequently completed and written to the log file.eval_retry(), copy previously completed samples in the log file being retried so that work is not unnecessarily repeated.inspect eval-retry command to retry a log file from a task that ended in error or cancellation.retryable_eval_logs() function and --retryable option for inspect list logs to query for tasks not yet completed within a log directory.shuffled property to datasets to determine if they were shuffled.extensions argument from list_eval_logs().Bugfix: Inspect view was not reliably updating when new evaluation logs were written.
Bugfix: results was not defined when no scorer was provided resulting in an error being thrown. Fixed by setting results = EvalResults() when no score
results was not defined when no scorer was provided resulting in an error being thrown. Fixed by setting results = EvalResults() when no scorer is provided.Update to non-beta version of Anthropic tool use (remove legacy xml tools implementation).
BREAKING: The pattern scorer has been modified to match against any (or all) regex match groups. This replaces the previous behaviour when there was m
pattern scorer has been modified to match against any (or all) regex match groups. This replaces the previous behaviour when there was more than one group, which would only match the second group.any option to indicate the model should use at least one tool (supported by Anthropic and Mistral, mapped to auto for OpenAI).list[ContentText | ContentImage].max_samples option to control how many samples are run in parallel (still defaults to running all samples in parallel).boolq.py benchmark.piqa.py benchmark.tuple return out of ToolResult type.TaskState.input_text.Remove deprecated support for matching partial model names (e.g. "gpt" or "claude").
ollama local model provider.multi_scorer() and majority_vote() functions for combining multiple scorers into a single score.model_graded_qa().TypeError for solvers and scorers not declared as async.NaN or Inf is encountered while reading log file header.Exclude null config values from listings in log viewer.
Show first log file immediately (don't wait for fetching metadata for other logs)
--version CLI arg and inspect info version command for interrogating version and runtime source path.null config values in output from inspect info log-fileFix issue with logs from S3 buckets in inspect view.
sort() method to Dataset (defaults to sorting by sample input length).write_eval_log() now ignores unserializable objects in metadata fields.
write_eval_log() now ignores unserializable objects in metadata fields.read_eval_log() now takes a str or FileInfo (for compatibility w/ list returned from list_eval_logs())..ipynb file due to lack of dependencies (e.g. nbformat).Your coding agent can read these notes before it upgrades. Set up the MCP server →