NewYour coding agent can read the release notes before it upgrades.Set up the MCP server →
PyPI · #2415 most downloaded on PyPI
Framework for large language model evaluations
Last release 2 days ago
02 Oct 2026
Ships on a steady schedule
a new release about every 1 weeks
Nearly every release is documented
notes for 60 of the last 60 stable releases
2 versions withdrawn
withdrawn after publishing
3 years old
253 releases · first in 2024
update changelog for release
update changelog for release
content, as the google provider sends them; LiteLLM otherwise passed the object as the function response itself and Vertex rejected documents with $ref keys (an OpenAPI spec read with curl) with a 400.update changelog for release
update changelog for release
reasoning_effort, mapped to the levels or thinking budgets the native Google provider uses.model_name does not contain "gemini" (LiteLLM otherwise replaced them with a placeholder); requests with tools carry the function-calling hint; and a turn returned as a malformed function call (no text or tool call, or a call written as code) is retried with a corrective exchange, as with the google provider.One column per month.
update changelog for release
update changelog for release
gpt-6.1-sol), including its context window and output limits.resolve_attachments is much faster for long conversations; in full mode, deeply nested model API call content may keep two more nesting levels.self_critique(), and model_graded_qa()/model_graded_fact() with model_role=None, now critique or grade with the correct model when one task is evaluated against several models..eval log from a non-S3 filesystem (e.g. gs://, az://) no longer stalls other running work for the whole download.subprocess() and Docker sandbox exec() no longer intermittently fail with BrokenPipeError when the command exits before reading its input.update changelog for release
update changelog for release
claude-sonnet-5-5): reasoning_effort="none" turns off up-front thinking, forced tool choice degrades to auto, and computer use uses the computer toolset on the Claude API and Vertex.web_browser() is deprecated, logs a warning when called, and will be removed in a future release.<name>.shards/ are hidden once a successful merged log covers them (inspect view --show-shards to list them); bundles leave them out.coordinate now click at the current cursor position instead of failing, and the tool description states which actions require one.back_click and forward_click now work (with a rebuilt aisiuk/inspect-computer-tool image); previously they failed inside the container regardless of arguments.update changelog for release
update changelog for release
--extra-headers and --extra-body options for inspect eval and inspect eval-set, taking an inline JSON or YAML mapping.update changelog for release
update changelog for release
eval_set: a task whose dataset GREW since its last run (a strict superset, with stable sample ids) is now topped up in place — the prior samples are reused and only the newly added samples run — instead of re-running the whole dataset. A shrunk dataset, or growth without stable ids, still triggers a full re-run.-M responses_api=false to opt out).-M thinking_display=omitted to opt out).model_info flags that enable adaptive thinking.web_search() without a search provider usable through the proxy fails when the model is first called, with an error naming the fix.attachment:// references; raw model API calls stay condensed until the sample completes.f1() now counts repeated tokens (multiset overlap, matching SQuAD F1); scores can rise or fall for answers or targets containing repeated words.update changelog for release
update changelog for release
litellm-proxy provider for models served by a LiteLLM proxy, which reads each model's upstream model for context window, cost, reasoning and Claude prompt caching.--env arguments preserve commas as literal text (such as --env 'NO_PROXY="localhost,127.0.0.1"') instead of being coerced into lists. (#5368)deepseek-flash), including image input; model info notes that the retired deepseek-v4-flash and deepseek-v4-flash-vision-exp names are now served by V4.1 Flash.bridged_tools are again denied unless the model proposed the call in a bridged generation, once per proposal, with or without an approval policy (0.3.265 ran them regardless as a stopgap); BridgedToolsSpec(require_proposal=False) opts a server out.SampleSource adds (including an empty-seed task with sandbox="docker") no longer fail with a LookupError, and their containers and generated compose files are cleaned up at the end of the run.match(), includes(), exact(), f1(), pattern() and answer() now record reason="no_response" when the raw model completion is empty or whitespace only, so a model that returned nothing is distinguishable from one that answered wrong. Score values are unchanged. (#5376)attempt_timeout are now retried, and ones cut off by a sample time_limit are recorded as that limit, instead of failing with a bare cancellation.inspect ctl task list and the inspect ctl sample reads take --log-dir <dir> to read task and sample status, events, messages and store from the .eval logs in a local or S3 directory, without a live eval process.update changelog for release
update changelog for release
update changelog for release
update changelog for release
computer_toolset_20260801 toolset; pass -M computer_toolset=true to use the toolset on Claude Opus 4.8, Sonnet 5 and Opus 5.key action now honors repeat./root/.cache/inspect), or a parent /root/.cache that other users could modify, now fails checkpoint setup and restore with a clear error instead of being reused.call_mcp_tool) by the tool's own name, and approvers see and modify the tool's own arguments.sample_abandoned() hook reports a sample cancelled before anything was logged (cancelled while queued, or before its retry_on_error re-run), so a source waiting on it no longer stalls.inspect CLI startup and import inspect_ai; InputRequest and request_input now annotate schema by name only, so typing.get_type_hints and pydantic schema generation for InputRequest are unsupported.AttributeError.read_eval_log_sample_summaries_async and three other async log readers to the public inspect_ai.log exports.meta provider for Muse Spark models on the Meta Model API, which streams by default, preserves model reasoning across turns, and reports policy-blocked prompts as content filter stops.gpt-6-sol) and GPT-6 Luna (gpt-6-luna), including reasoning_effort="none", which these models accept (GPT-6 Astra does not).model from its response.sandbox_agent_bridge(bridged_tools=...) now fails the sample as a native tool exception does, and malformed arguments are reported to the model as a parsing error.bash_session no longer sends the literal string "None" to the shell when type_submit is called without input.find only from /usr/sbin:/usr/bin:/sbin:/bin, not the image PATH.default_headers TypeError when ANTHROPIC_AUTH_TOKEN is set and the caller supplies its own default headers.sample_cleanup(), which could leak sandboxes on certain providers.code_execution() tool (native execution disabled) are now executed instead of being silently dropped when xAI reports them as its built-in tool.claude-opus-5-5): thinking can't be disabled, forced tool choice degrades to auto, and computer use is not yet supported on the Claude API and Vertex.update changelog for release
update changelog for release
inspect score --scorer pkg/name now resolves @scanner functions from installed packages (e.g. inspect_petri/audit_judge) instead of failing with LookupError; unknown names now report the "scorer couldn't be loaded" guidance rather than a raw traceback.inspect ctl task score interim-metrics payload now reports each entry's originating scorer under scorer and the score's name under name (previously scorer held the score name). This lets consumers disambiguate scores from dict-valued scorers — where several scorers can share a score name — and reconstruct the EvalScore needed to resolve a task's headline metric. Breaking for clients that read the old scorer field as the score name.update changelog for release
update changelog for release
update changelog for release
update changelog for release
function_call and custom_tool_call output items now carry a non-null item id, and streamed custom tool calls now report completed status so client SDKs dispatch them.AssertionError: Session was never entered in unrelated code.cache_ttl now defaults to "auto", which switches a sample's prompt-cache TTL from 5 minutes to 1 hour after a >5 minute gap between its requests; pass "5m" or "1h" to pin.literal: task targets now keep the rest of the value when it contains additional colons.human_reviewer() lets an operator review a tool call together with its result and continue or terminate the sample, on the same surfaces as the human approver.react() accepts review policies, which apply to the agent's tool calls in place of any eval-level or task-level reviewers, as approval does for approvers.aiobotocore instead of aioboto3, which is no longer installed, and Inspect no longer holds botocore back to an old release.eval_retry now reuses the model roles recorded in the original log, including roles the task set itself in Task(...).Reviewer protocol and review policies (Task(review=), eval(review=), --review) run after a tool call executes and before the model sees its result, and can continue, terminate, or escalate; each decision is recorded as a ReviewEvent.<think> tag, so reasoning no longer leaks into assistant output text. As in OpenRouter's own SDK, only signed text and encrypted entries are replayed: Gemini no longer sees its own readable prior thinking on later turns, only the thought signature. Reasoning replayed from another provider (no OpenRouter details) now goes into the <think> tag as readable text only, and is omitted entirely when it has none (e.g. a redacted block with no summary), so no signature or opaque payload enters the assistant text channel for any model family.RuntimeError: Event loop is closed when a memoized model is used across multiple eval() calls or event loops.ValueError("Unexpected output type: ResponseToolSearchOutputItem") when an agent uses native OpenAI deferred tool search; the response-item handler now recognises the tool_search_output item without overwriting the cached tool_search_call. (#4968)--sample-id now accepts ids containing colons (e.g. user:cybergym/arvo_6008); a task: prefix is stripped only when it names a task in the run.0 now raises an error instead of being interpreted as option Z on tasks with 26 or more choices.multiple_choice() now recognizes answer letters wrapped in LaTeX or markdown ($B$, **B**, (B)), which previously scored INCORRECT.perplexity() and target_perplexity() now record infinite perplexity for a sample whose NLL is too large to exponentiate instead of losing the sample to an OverflowError.csv_dataset() now honors the dialect's delimiter when no explicit delimiter is supplied, including tab-separated and registered custom dialects.csv_dataset() now loads UTF-8 files with a byte-order mark, including Excel CSV exports, without requiring an explicit encoding.file_dataset() now reads .tsv and .tab files as tab-delimited instead of rejecting them.ask_user prompts are no longer hard-wrapped by the console, so long commands copy out of the terminal intact.inspect ctl sample cancel now works on a sample that is still initializing (e.g. waiting on sandbox provisioning) — the cancel applies the moment the sample starts, and inspect ctl sample list marks the pending cancel.eval(), which dominated the wall time of very small evals during tests.INSPECT_EVAL_CTL_SERVER is now honored by eval() and eval_set() called from Python, not only by the CLI.sample_complete() now fires for a running sample cancelled individually, so a source waiting on that sample no longer stalls; a blocking callback can no longer hang a task cancel.math() now records reason="invalid_response_format" when no answer can be extracted, so format failures are distinguishable from wrong answers.math() now gives a symbolic answer the same verdict whether it and the target are written in LaTeX or plain notation (e.g. \frac{x}{2} vs x/2), plain-notation symbols match case-insensitively like LaTeX ones, and an answer is parsed under its target's symbol assumptions so cancelling imaginary terms (e.g. x+i-i vs x) still match.choice() now tags empty completions as NOANSWER with reason="no_response" and records reason="invalid_response_format" when no choice can be parsed, so format failures are distinguishable from wrong answers.grouped()) now report it on an all-unscored run instead of collapsing to a synthesized flat NaN, on both the list and dict metric paths; metrics that raise on empty input still report NaN, with a one-time warning. (#5150)file:// URIs (e.g. paths with spaces, as Path.as_uri() produces) are now accepted by eval(), Task(), and --approval././ sandbox_paths entry now fails when the sample starts instead of after its checkpoints have been taken and cannot be restored./, \, ~ or NUL, or longer than 200 bytes, get a hashed checkpoint directory name and no longer resume checkpoints from earlier versions.model_length output.tool_choice/toolConfig now return a 400 naming the bad field instead of a status-less error, and a non-string tool name no longer poisons the sample transcript.retry_cleanup, lose) samples an earlier attempt completed..json log by id now matches the id's string form exactly, as .eval logs always have (1 finds "1"), instead of also matching zero-padded numeric forms such as "001".log_images=False keeps the images already recorded in the prior attempt's reused samples.pids_limit, read_only, cgroup, stop_grace_period, build.no_cache, or build.pull no longer fail validation when starting an eval.bash_session(), text_editor(), exec_remote() and sandboxed MCP servers now run as the sandbox's default user instead of always as root; the three tools accept user="root" to restore the old behavior.INSPECT_SANDBOX_TOOLS_STRICT_DIGESTS has been removed.PATH; images must provide /bin/sh and the coreutils Inspect uses (including zstd for existing .tar.zst checkpoints) in the system bin/sbin directories, not /usr/local.eval_set() calls instead of being re-run every time.input that exit without reading stdin now return their exit status and stderr instead of raising, and commands that fill stdout before reading stdin no longer hang.task CLI now refuses a pre-planted or non-root-owned /opt/human_agent, errors instead of silently skipping when it cannot be created, and no longer runs a staged install script.task shell hook is now appended to the login user's .bashrc as that user, so a .bashrc that user cannot write (e.g. root-owned) or that is a symlink fails installation instead of being written by the sandbox default user (root in most images).custom_outputs now populate default token usage when the returned ModelOutput omits usage, matching iterable/generator behavior.eval_retry now reuses the model roles recorded in the original log, including roles the task set itself in Task(...)..eval files no longer fails with AttributeError: ... '_needs_input' on Python builds that include CPython's gh-156002 zipfile change.--log-level info no longer prints a line for every OpenAI and Anthropic HTTP request.math() now raises an error when no reference answer can be parsed instead of silently excluding the sample from metrics.choice() now raises an error for samples without answer options instead of silently scoring them incorrect.fallback_models) with a 400 error.fail_on_refusal generate config option (--fail-on-refusal) fails a sample with a ModelRefusalError when a model refuses a request, settable eval-wide, per task, per model, per model role, or per call./var/tmp/sandbox-services parent already exists with the wrong owner, mode, or type.nonroot account (UID/GID 65532; default user unchanged) and installs the web browser's Playwright browsers to a shared path so the browser tool works under a non-root user.update changelog for release
update changelog for release
gpt-6-astra).temperature, top_p, and logprobs causing a 400 error on GPT-5.5+ models when no reasoning_effort is set (they are now dropped with a warning).stop_reason="model_length" instead of failing with a generation error under xAI's current context-overflow wording.Score.reason, so an explicit reason is no longer dropped (and a stale metadata["unscored_reason"] cannot replace it).finish_reason="length" instead of "stop".system_instruction as REST Content instead of a bare JSON array, fixing batch submission failures; batch results are now parsed through the SDK converter so unknown REST fields (e.g. usageMetadata.serviceTier) no longer fail the whole batch. (#5100)chain_of_thought() now uses format_template() for prompt interpolation, matching the other prompt solvers: custom templates with extra {name} placeholders no longer raise KeyError (unknown placeholders pass through unchanged; JSON-style braces still need {{ }} escaping). (#5166)claude-mythos-preview now reports its release date, context window and output limit instead of resolving to no model info.inspect ctl task drain command stops dispatching a task's queued samples while in-flight samples finish naturally, completing the task with an ordinary log whose abandoned remainder a later inspect eval-set re-invocation or inspect eval-retry still runs.inspect eval-set re-invocation now re-runs samples left queued (never dispatched) when a task was ended with inspect ctl task cancel --action score|error; previously the cancelled task's log read as complete and was reused, so those samples never ran.EvalResults gains an optional logged_samples field, set on logs finished by a graceful inspect ctl task cancel or inspect ctl task drain, recording how many samples the log actually resolved.inspect ctl task cancel of a task between retry attempts (including one still writing its errored attempt's final log) now abandons the queued retry (the task ends with its last attempt's error log) instead of asking to re-issue once the retry starts.inspect ctl task list rows now report a pending graceful resolution (resolving: drain/score/error), and cancelled samples of a task whose retry was suppressed or abandoned no longer render as pending.attempt_timeout or cache_prompt no longer misses the model cache, since neither changes what the provider returns.hidden_states (from -M hidden_states) as JSON-serializable nested lists instead of silently dropping them to None in the log; the batched Hugging Face path now records each sample's own activations rather than the whole batch's. Note: code reading ModelOutput.metadata["hidden_states"] live (in a solver or scorer) now receives nested lists rather than tensors — wrap with torch.tensor(...) if tensor operations are needed. (#2860)update changelog for release
update changelog for release
reducer="mode" to model_graded_qa()/model_graded_fact() to restore the previous behavior. (#4721)pattern() now falls back to the full regex match when the pattern contains no explicit capture groups (previously such patterns always scored INCORRECT). (#4828)fsspec upper bound from <=2025.9.0 to <=2026.6.0 to align with the current huggingface/datasets cap. (#4761)max_sandboxes and max_subprocesses now default off the processors the eval may actually use, so an eval in a CPU-limited container no longer oversubscribes its quota.grep tool now supports extended regex via a new extended_regexp option (patterns remain basic regex by default).system block boundaries, so instruction blocks are no longer silently discarded by the API.reasoning_effort="minimal" now maps to low (with a warning) on models that reject minimal thinking rather than failing the request.model_graded_qa/model_graded_fact no longer score a malformed multi-character verdict such as GRADE: CI as correct; such verdicts now leave the sample unscored with grade_parse_failure recorded.local_path() now percent-decodes file:// URLs, restoring the to_uri() round trip for paths with spaces or other encoded characters (affects every consumer of file URLs: log recorders, control-channel state, checkpoint restore, approval policy files). (#5025)files values that are empty strings now create empty files in the sandbox rather than recursively copying the runner's working directory (or the dataset's directory) into it.task_info() and log_viewer() raising ValueError when applied to an empty DataFrame. (#4826)eval_set() now defaults log_dir to INSPECT_LOG_DIR or ./logs, as eval() does, rather than requiring it.eval_set() can now override any argument that does not change task identity, rather than only five.INSPECT_EVAL_* environment variables with the same meanings inspect eval-set gives them.inspect ctl config --max-samples now works for tasks using adaptive connections — an integer pins sample concurrency, and clear resumes adaptive tracking..eval log uploads to S3 now appear in inspect trace output.ToolSource no longer serves one sample's tools to another; tools are now resolved fresh through the server on every call rather than cached. Sessions are scoped per async task rather than per sample, so a handoff or other subagent's turn (which runs in its own task) now gets its own connection to a local MCP server rather than reusing its parent's live session.ValueError: seek of closed file when writing a .eval log smaller than 8MB to S3.inspect ctl sample score interim-scores a single running sample's work-so-far on demand (briefly held while scored; the sample keeps running).inspect score now reports samples that errored or were stopped early, so a re-scored partial run is no longer displayed as if it were complete.cascade() scorer runs scorers in order and reports the first that settles a sample, so an expensive grader only runs on samples cheaper scorers left unsettled.usage.output_tokens_details, so bridged clients can distinguish a reasoning response from a plain one.output_config.format, OpenAI text.verbosity, or any of Gemini's thinkingConfig.thinkingBudget, presencePenalty, frequencyPenalty, responseLogprobs and logprobs now gets it, instead of silently receiving the model default.thoughtsTokenCount, so a client can see how many thinking tokens a request used.stop_sequences: 5) is now rejected with a 400 instead of being recorded into an event that cannot be read back, which made the whole sample transcript unreadable.candidatesTokenCount no longer double-counts thinking tokens alongside thoughtsTokenCount.$ref) now warns, instead of silently constraining the model more weakly than asked.TypedDicts rather than pydantic models, cutting manifest parse time and GC pressure on the sync thread for runs with many segments.log_shared sync.bridged_tools by calling the bridge service directly; unapproved host tool calls are now denied.Task are now recorded in the eval log.id key on synthesized message and reasoning items rather than sending an explicit null, which backends like vLLM reject.inspect ctl sample cancel --action cancel now works on samples that haven't started — cancelling a never-started sample before it runs and withdrawing (un-requeuing) a queued re-run so its prior outcome stands.inspect log recover, eval-retry, and eval_set gain --incomplete-action error (with an --incomplete-max guard) to resolve samples that were in progress at a crash as errors and finalize the eval as success instead of endlessly re-running hung samples.inspect log recover --list now reports total_samples for evals run with --limit or --sample-id as the selected samples times epochs, rather than the full dataset size.ci_wilson() metric reporting the Wilson score confidence interval for the mean of binary scores (as {"lower", "upper"}), with bounds always within [0, 1]; cluster= computes an effective-sample-size interval accounting for within-cluster correlation.inspect_ai.scorer now exports Reference, the model for the message/event references that scanner scores store in metadata (previously importable only from Inspect Scout).temperature, top_p, penalties) set on Kimi models with thinking disabled are now warned about and ignored instead of causing a 400 error.stream_idle_timeout option (also retunable live via inspect ctl config) abandons and retries a model call whose streaming response stalls, instead of waiting out the whole-attempt timeout.top_logprobs is now honored instead of being silently capped at 1.on_stream now delivers stream events from the Bedrock, Groq, Mistral (completions API), and Azure AI providers.tool_choice="any") on Kimi models other than K3 no longer fails with a 400 — the request falls back to "auto" with a warning.value_to_float() now maps numeric custom correct/incorrect/partial/noanswer values to 1/0/0.5/0 as documented, instead of silently passing them through (e.g. a custom incorrect=-1 no longer produces negative accuracy). Default string sentinels and non-finite values are unaffected. (#4928)value_to_float() matches sentinels with ==, so custom numeric sentinels also map equal bools (True == 1.0) and equal elements of list/dict values in score reducers — the same reach the default string sentinels already had; this is now documented. Mapped values are also always returned as floats (the incorrect/noanswer branch previously returned an int). (#4928)update changelog for release
update changelog for release
match(numeric=True) now parses numbers with attached sentence or enclosing punctuation (e.g. "42!", "(42)"); operator prefixes (<42, ~42) still do not match, and location="exact" remains strict. (#4742)tools: web_browser*, bash syntax; policy file paths given as file:// URLs are normalized to local paths. (#5025)AzureAuthError with remediation guidance instead of being downgraded to a warning and an empty listing, matching how S3 surfaces auth failures. (#4914)task submit/task quit confirmation prompt now declines cleanly instead of raising an EOFError traceback.as_tool() returns its full response rather than one clipped at the maximum tool output size.agent_list() now reports one line per background agent instead of embedding every completed agent's full report.@tool(max_output=...), overriding max_tool_output (0 disables truncation for that tool).on_stream is passed, so it can no longer fail model calls that stream without a callback.write_eval_log() and write_eval_log_async().inspect_ai.util._sandbox.self_check is now a collection of plain pytest tests. See docstring for migration instructions.ask_user() in a detached worker parks for someone to attach (inspect acp --server <socket>) rather than erroring on a console that isn't there.activity of approval (naming the tool being decided) or question, where it previously showed nothing at all — an approval is awaited before its tool call is recorded, so the wait left no trace and the sample read as idle.sun_path limit a long home directory could otherwise push the old default past.limit and max_sandboxes.call_id (optional as of openai 3.5.0) no longer error in the agent bridge or token-count padding.content_filter stop reasons instead of failing the sample.INSPECT_EVAL_SET_SELECTION), so an external runner can execute one task per process into a shared log directory while owning the eval-set metadata itself. Selected tasks run through the ordinary eval() path with no eval-set orchestration; because the runner owns completion decisions, workers neither fail a task on sample errors nor retry a task in-process.Score.reason field records why a score has an abnormal value (e.g. invalid_response_format, grader_failed), is preserved across score edits, and appears as score_<name>_reason dataframe columns. (#4567)pattern() and answer() now score unmatched output as INCORRECT with reason="invalid_response_format" instead of NOANSWER. Default metrics are unchanged (the default value_to_float already maps NOANSWER to 0); analyses that filter on value == "N", and custom value_to_float mappings that treat noanswer differently, should key on reason instead. (#4567)perplexity() and target_perplexity() now return Score.unscored() with a reason instead of a raw NaN value when logprobs are unavailable. All unscorable states (including an empty completion, which earlier revisions labeled no_response) carry reason="scoring_failed": the sample is excluded from metrics, so the reason reports the instrument declining to run — the empty-completion detail remains in explanation. (#4567)inspect ctl task score scores a running eval's in-flight samples (each briefly held while scored) and reports interim metrics that fold in completed samples' final scores, without ending any sample..json log now reports the requested uuid when the sample is missing, and raises a clear error when neither id nor uuid is provided.max_model_len is now registered as the model's context window, so compaction and context-length handling reflect the served configuration (including LoRA adapters via their parent model). (#4215)INSPECT_EVAL_SET_SELECTION), so an external runner can execute one task per process into a shared log directory while owning the eval-set metadata itself.inspect ctl model throughput (and GET /models/throughput) reports each model's recent output tokens/sec, retries/min, and backoff across the run; trace retry lines and the display footer now carry the current rate.EvalSampleLimit.reason, limit_reason on sample summaries and samples_df()), so operator-terminated samples can be told apart without reading transcript events.message_limit / time_limit fields alongside the existing token_limit), including mid-run inspect ctl config retunes.INSPECT_SANDBOX_TOOLS_STRICT_DIGESTS is set.bash_session no longer stops returning output for the rest of the session when multibyte output happens to be split mid-character across reads.--sandbox-prebuilt option (sandbox_prebuilt on eval()) skips image builds and fails fast at task startup when a prebuilt image is missing.x-local: false on a compose service is now treated the same as omitting x-local (the image is pulled) rather than marking the image as local.inspect ctl sample store command reads a running or just-finished sample's current store directly (with server-side --key exact/prefix filtering).inspect ctl sample cancel-tool-call cancels one hung tool call (the model sees an ordinary tool timeout and the sample continues), with pending tool calls now visible in inspect ctl sample list --json.headline_metric naming which scorer and metric summarise the eval, honored by the log listing, the progress display, and evals_df().Model.generate() and Model.generate_loop() accept an optional on_stream callback that by itself enables provider streaming and receives incremental events (text/reasoning/tool-call deltas and retry boundaries), with streamed progress now also visible on inspect ctl sample list.on_stream now delivers stream events from the OpenAI, OpenAI-compatible (Together, Fireworks, etc.), Grok, and SageMaker providers, in addition to Anthropic and Google.stop_reason="content_filter" (as non-streamed ones do) instead of failing the call.stream=true with non-strict tools no longer fails before the request is sent.stream/streaming model args now accept auto uniformly across providers, and unrecognized values raise an error instead of silently enabling or disabling streaming.ValidationError failing the sample) when native compaction ran with streaming enabled.inspect ctl config --json refusals against an older eval process now emit the structured {"error": ...} envelope instead of only stderr prose.inspect ctl config --max-tasks retunes a running eval's task concurrency mid-flight — raising it starts pending tasks immediately (pass clear to restore launch config).inspect ctl retries them shortly.inspect ctl commands now take a --model disambiguator, so one task run against several models can be selected by name (e.g. inspect ctl task cancel my_task --model gpt-5).model_roles={"grader": [...]}), with model-graded scorers grading by majority vote across the list.INSPECT_HTTP_* environment variables, and connection setup gets 60s rather than the SDKs' 5s.message.get(...)) no longer fail with a Jinja UndefinedError.update changelog for release
update changelog for release
ci() metric reporting a confidence interval for the mean (as {"lower", "upper"}). Defaults to mean ± t · stderr with a Student-t critical value (n - 1 degrees of freedom; clusters - 1 when cluster= is set) so small samples get honest widths; method="bootstrap" gives a percentile (cluster) bootstrap interval. (#4160)temperature/top_p/top_k continue to work on models that support them; browser state tool results and file-based image/document sources no longer fail type checking).cache_write_tokens as ModelUsage.input_tokens_cache_write and excludes it from full-rate input_tokens (generate and compaction responses); compaction usage also now excludes cache reads and records reasoning tokens. (#4855)-dev sandbox-tools binaries when local main refs are missing, stale, or unavailable.Task objects, preventing later evals from inheriting prior overrides.sandbox_unavailable rather than as command output; non-tool callers (e.g. scorers) get a SandboxUnavailableError or PermissionError raise. (#4709)ANSWER: A,) is now scored instead of rejected as no answer.inspect ctl sample requeue can now sweep every currently-errored sample in one command (--errored) or requeue several SAMPLE_ID EPOCH pairs, reporting each sample's result individually.inspect ctl spellings (e.g. ctl tasks, ctl limits); use the noun-group commands (ctl task list, ctl config, ...) instead.inspect ctl config can now retune a running task's per-sample time/token/message limits mid-flight (--time-limit / --token-limit / --message-limit), reaching in-flight samples as well as ones not yet started..eval log now yields between samples, so a large flush no longer stalls in-flight samples and control-channel requests for the whole batch.inspect/turn_state extension notification (started / ended / cancelled) from the agent turn boundary so ACP clients have an exact "agent working" signal.inspect ctl human-readable output, preventing spoofing of the operator's terminal.response_schema) no longer fails with a 400 — it now uses DeepSeek's JSON mode, warning that the schema must be described in the prompt.--now flag on inspect ctl task|model|process pause additionally holds in-flight samples at their next model call until resume, with held-sample counts reported in inspect ctl task list.inspect ctl config no longer version-gates individual knobs — current eval processes reject unsupported knobs atomically server-side, and only processes predating strict config validation are refused as a whole.text_editor, bash_session) failing to install in non-root sandboxes (e.g. Kubernetes pods with runAsNonRoot).retry_refusals gets a fresh model attempt instead of a replayed cached refusal.web_search("exa") no longer fails with a validation error, and Exa citations now include page text by default.materialize_media() before model use; fixed selected-dataset media remains automatic, while sandbox bridges require inline data URIs.data: URIs (e.g. host paths or URLs); such requests are now rejected.exec() output limit is enforced by front-truncating the output streams rather than by raising OutputLimitExceededError (which remains the behaviour for read_file()). (#4778)update changelog for release
update changelog for release
precomputed_scores() scorer applies scores computed outside of Inspect (e.g. human ratings) from a JSON or JSON Lines file, matched to samples by id and epoch.file_dataset() now recognizes JSON and CSV URLs with query parameters while preserving the complete URL passed to the selected dataset reader.model_graded_qa/model_graded_fact now leave a sample unscored when the default grader's final GRADE: verdict is a letter the instructions never offered, instead of scoring a grade mentioned earlier in its reasoning. Note: re-scoring existing logs may shift metrics and sample counts — samples whose grader verdict was off-menu (including P grades when partial_credit is disabled) were previously scored from an earlier on-menu mention and are now unscored with grade_parse_failure recorded.usage.reasoning_tokens and log a token-counting warning on every generate.xhigh reasoning effort and a service_tier model argument for Priority Processing.inspect ctl sample events / messages / show / errors) now return metadata only by default, so monitors that never pass --content cannot be prompt-injected by the evaluated agent's output.inspect ctl mutations with piped or captured output now print one outcome line each instead of repeating the full task header (--terse/--no-terse to override).inspect ctl command's --help now sketches its --json payload's top-level keys.ValidationError with openai >= 3.1.0, which reports MCP call errors as structured objects rather than strings.ValueError naming the file and line, instead of AttributeError or a silent load. (#4546)ANSWER: A, B, and C) are now scored correctly instead of as no answer.AttributeError.update changelog for release
update changelog for release
inspect ctl sample and other unscoped commands no longer appear to hang on eval sets with many running tasks. (#4789)bundle_log_dir() allows output_dir starting with hf/ when log_dir is current directory or parent directory.data_uri_to_base64() strips data: headers from URIs with empty media types.hf_dataset(..., auto_id=True, shuffle=True) now attaches each auto id to its record (matching csv/json) instead of the shuffled position, so a record keeps the same id across seeds and limited slices. (#4459) Note: ids for affected datasets will change once on upgrade (row order for a given seed is unchanged) — avoid retrying an in-flight eval_set across this boundary. Previously, unseeded shuffles assigned irreproducible ids, which silently corrupted eval_set retries on affected datasets; those workflows are now correct.aggregate(key, agg=...) metric factory applying any standard metric (mean, stderr, accuracy, …) to a single key of a dict-valued Score.value.krippendorff_alpha() metric for inter-rater agreement across multiple judges, with nominal / ordinal / interval measurement scales.collect score reducer that preserves each scorer's value as a list instead of aggregating.NaN (as produced by Pandas, Hugging Face, CSV, and PyArrow sources for missing values) are now treated as missing for input, choices, setup, sandbox, files, and metadata, matching the existing target behavior. (#4626)APIConnectionError: Connection error. when openai 3.x is installed.reasoning_effort values are now mapped to valid provider/model tiers for Together, SambaNova, Perplexity, and Fireworks.--model-spec option runs several models in one inspect eval or inspect eval-set command, each with its own generation config, model args, and base url.update changelog for release
update changelog for release
stderr(cluster=...) avoids quadratic time and memory in cluster size; results and warnings can differ for extreme scores.read_file now raises FileNotFoundError when the container reports "no such file or directory". (#4686)pyobjc-framework-AppKit is missing or no display is attached.list_eval_logs_async() now lists remote log directories (S3/GCS/Azure) without blocking the event loop, treats a missing S3 bucket as an empty listing, and downgrades Azure auth errors to a warning.update changelog for release
update changelog for release
math() scorer answers with a non-evaluating grammar under a bounded worker thread, preventing model output from executing Python on the evaluator host.on_task_start now receives the resolved solver plan as data.plan, including any Task.setup solvers.Updated changelog to reflect version 0.3.255 release.
Updated changelog to reflect version 0.3.255 release.
inspect ctl task list now reports per-eval refusals and http_retries.update changelog for release
update changelog for release
subagent_type parameter from the agent() tool when there is only one subagent type.update changelog for release
update changelog for release
ComposeService now accepts the standard compose keys restart, stdin_open, and tty (previously rejected as unknown fields).eval_set retry, or failing log validation.task_retry_attempts now use the same task dispatcher as runs with retries (the separate no-retry dispatcher was removed).inspect ctl sample events --full now pretty-prints the raw events instead of rendering a mostly-empty summary table.inspect ctl sample messages TASK SID [EPOCH] reads a running (or buffered-but-unlogged) sample's current conversation as a snapshot, with --tail/--all/--full.inspect ctl task pause|resume and inspect ctl process pause|resume commands pause a running eval or eval-set (in-flight samples finish; nothing new starts) and resume it in place, with paused/quiesced reported by inspect ctl task list.f1() and exact() now match decimal-number answers wrapped in punctuation (e.g. "(3.14)" or a trailing period) against bare-number targets. (#4620)- DeepSeek: New native deepseek provider with support for DeepSeek V4 models (deepseek/deepseek-v4-pro and deepseek/deepseek-v4-flash), including thinking control via --reasoning-effort.--reasoning-effort now controls thinking on current Mistral reasoning models (Mistral Medium 3.5 and Mistral Small 4), which replace the retired always-thinking Magistral models.exact() now requires word order and word counts to match (previously it compared unordered word sets, so e.g. "world hello" scored as an exact match for "hello world"). Note: exact() scores may decrease on existing evals where answers matched only via word reordering or duplicate collapse; use f1() for order-insensitive scoring.match(numeric=True) now recognizes LaTeX-escaped currency and formatting symbols (e.g. \$20, \€20, 1\,000), which previously failed numeric extraction and could silently match a different number in the output. Scores may shift on affected samples (mostly upward; the previously ignored number is now the one matched).reasoning_effort and effort via adaptive thinking, alongside response_schema; reasoning_tokens is promoted to adaptive on Claude 4.7+. (#3765)stderr(cluster=...) and grouped(all="groups") now return 0.0 for empty score lists without emitting NumPy runtime warnings. (#4718)total_cost now bills Anthropic 1-hour prompt-cache writes at 2x the base input price rather than the 5-minute rate, so runs using -M cache_ttl=1h are no longer understated. (#4703)include_history=True no longer present an empty history for samples without an assistant turn; such samples may now receive parseable grades and enter the metric denominator. (#4722)HF_TOKEN or cached huggingface-cli login credentials instead of authenticating with a placeholder and being rate limited as anonymous. (#4600)data:/file: sources only, gated by the sanitizer and capped to the transcript column width); raw HTML in message content remains escaped as before. Thanks @rasmusfaber. (#505)exec calls render their JavaScript source as a highlighted, expandable input body instead of a truncated single-line header summary; header-only tool blocks no longer sit flush against their bottom edge. (#504)update chanelog for release
update chanelog for release
bool-typed columns now coerce via YAML, so "false"/"0"/"no"/"off" parse to False instead of every non-empty string becoming True.message_limit/token_limit/cost_limit errors not being properly raised when running in a sandbox..eval logs with exclude_fields no longer fails with "integer overflow" when a sample contains a JSON integer larger than 2⁶³−1.inspect ctl sample list/show now report a running sample's in-flight activity (generating 7:12, bash 0:41, retrying in 0:45), and sample events renders pending events, so long model calls and retry backoffs no longer read as silent idle.inspect ctl and viewers) as soon as the reuse sweep completes.inspect ctl sample events --tail (including the default first page) now counts events matching the --type filter, so default reads show a useful recent window instead of a near-empty page. (meridianlabs-ai/inspect_ai#162)inspect ctl sample list, or a busy-skipped process — it now points at inspect ctl process anomalies as the escalation.inspect eval now prints a one-line pointer to inspect ctl monitoring at launch, eval/ctl help text cross-links the two surfaces, and inspect ctl task list flags errored samples.@task, @solver, and @agent decorators now retain their return annotation on Python 3.14 even when re-wrapped by user decorators, so registry_create() and signature introspection keep working. (#4554)@scorer and @tool objects whose decorated function omits its return annotation (previously this silently failed). (#4554)eval_async, recover_eval_log_async, recoverable-log discovery, and accessing EvalLog.samples from their results) no longer raise "read_eval_log cannot be called from a trio async context" under a trio event loop.sample_init event with a null state can now be re-read (the null was previously dropped on write and failed validation on read).list_eval_logs_async(), an async equivalent of list_eval_logs() that is safe to call from any async context (including trio); calling list_eval_logs() with a filter under trio now raises an error directing to it.google web search provider now returns results in search rank order rather than page-fetch completion order.has_client and wait_for_client(), letting agents check for or wait until a fully bound ACP client is attached.EvalResults.completed_samples now counts samples that completed without error rather than scored samples, so score-on-error samples no longer inflate it. (#4602)**kwargs now survive replay from logs (e.g. eval_retry, inspect score), including keyword arguments named name. (#4375)@solver, @agent, @tool, …) that declare **kwargs no longer crash at registration when a keyword argument is named type, o, or info, and approval-policy params named type/name no longer collide on creation. (#4504)frontier() no longer errors when a model release date group has only missing (NA) headline scores; those rows are now skipped instead of crashing idxmax().count_tokens() concurrency (adaptive, like generate()) so large token-count fan-outs no longer overwhelm the provider connection pool.generate() now retries the anyio transport-close race (AttributeError: 'NoneType' object has no attribute 'call_soon' during response cleanup) instead of failing the sample.mcp_server_*() tools fail. (meridianlabs-ai/inspect_ai#170)mcp_server_http() no longer ignores its name argument (the server was always named after the URL).find/find_in_page actions missing a url or pattern now degrade to a search action instead of crashing the sample with a validation error. (#4119)inspect ctl model pause|resume commands pause one model's dispatch (including not-yet-started eval-set tasks) while the rest of the run continues.inspect ctl task list now suggests only the resume commands for latches actually holding a paused task, instead of always listing all three.inspect ctl sample requeue to re-run one errored/cancelled sample inside the still-running eval.--action error now log the error message Sample errored: interrupted by operator and report status error (previously the raw cancellation message, misreported as cancelled).text_editor() paths containing a null byte now return a tool error to the model instead of crashing the tool. (#4659)count_tokens() now sends extra_headers from the generate config (including any anthropic-beta values), matching generate(). (#4606)CompactionEdit clearing server-side tool uses at stale message indices when keep_tool_inputs=False removes client-side tool messages, which raised IndexError or wrote a tool use into an unrelated message. (#4528)start_period, so services with a long startup grace period now start reliably. (#4698)1.5s) in Docker compose files now produce the correct startup timeout instead of a silently wrong one. (#4698)agent-client-protocol >= 0.12 (adapts to its renamed multi-select schema types and new catch-all property type).update changelog for release
update changelog for release
on_model_retry now reports the retry cause via exception_type and status_code on ModelRetry, so hooks can distinguish e.g. timeouts from 429s and 5xxs.additional_tools declarations (used by newer OpenAI models to declare tools via input items rather than the top-level tools array) through the bridge as real tools, preserving each tool's original schema verbatim. Previously these declarations were dropped, so the model received no tools; reconstructing them was also lossy and broke reserved-schema tools such as codex's collaboration.*. (#4556)Sample Sources: Generate a task's samples dynamically while it runs by passing a SampleSource as the task's dataset (with enqueue_sample() for imperat
SampleSource as the task's dataset (with enqueue_sample() for imperative additions) — the sample-level mirror of TaskSource, for RL loops and adaptive evals.task_retry_attempts now use the same task dispatcher as runs with retries (the separate no-retry dispatcher was removed).inspect ctl task pause|resume and inspect ctl process pause|resume commands pause a running eval or eval-set (in-flight samples finish; nothing new starts) and resume it in place, with paused/quiesced reported by inspect ctl task list.inspect ctl sample messages TASK SID [EPOCH] reads a running (or buffered-but-unlogged) sample's current conversation as a snapshot, with --tail/--all/--full.inspect ctl process anomalies [PID] shows a running process's in-flight and anomalous trace actions by reading its trace file directly, so it works against busy, hung, or even exited processes.inspect ctl config changes are now recorded in each affected eval log (EvalLog.config_updates, with author, timestamp, optional --reason, and old → new values); effective_eval_config()/effective_generate_config() fold them over the launch config.inspect ctl sample events --full now pretty-prints the raw events instead of rendering a mostly-empty summary table.reasoning_effort="none".ContentDocument inputs.-M stream=true hitting max_tokens) now degrades gracefully to a max_tokens stop reason instead of raising LengthFinishReasonError and aborting the eval. (#4552)<think> text, rather than HTML-escaped signature JSON and encrypted payloads. (#4320)with_raw_response, which previously failed with 'Message' object has no attribute 'parse'.functionResponse parts in model-role turns are now converted to tool messages instead of dropped (fixes 400 "Requests ending with a model turn are not supported").Sandbox.exec failure errors now include stdout as a fallback when stderr is empty, so commands that report diagnostics on stdout still produce a useful error message./memories paths so equivalent spellings map to one file rather than several, and add an instance parameter to memory() for independent per-instance memory stores.file:// URIs as well as plain local paths.validation_predicate type for Scout extensions.inspect trace commands and inspect ctl process anomalies) now skip truncated, corrupt, or unrecognized records — e.g. from a hard-killed process or a newer inspect version — instead of failing entirely.inspect ctl discovery and pid targeting) no longer risk sending a console Ctrl+C to the probed process or misreporting live processes as dead.chain_of_thought solver. (#4631)is_data_uri() now recognizes base64 data URIs that carry media-type parameters (e.g. ;charset=utf-8) or an empty media type (data:;base64,...), which were previously misclassified as non-data URIs.file_dataset() now recognizes JSON and CSV URLs with query parameters while preserving the complete URL passed to the selected dataset reader.generate() calls with differing system prompts as utility agents; utility classification now requires a tool-calling agent loop.limit=0 instead of returning all samples.at_least(2), pass_at(2), pass_k(2)) now round-trip through the registry instead of raising LookupError when restored from a log.read_eval_log_samples()) was retained in memory forever after the caller released it.container_id is required when a turn mixes code-execution-backed server tools (e.g. web search with dynamic filtering) with client tool calls.Model API: New moonshot provider for Moonshot AI Kimi models (e.g. moonshot/kimi-k3), with built-in model info for Kimi K3 and detection of kimi-* mod
moonshot provider for Moonshot AI Kimi models (e.g. moonshot/kimi-k3), with built-in model info for Kimi K3 and detection of kimi-* model names on hosting providers.inspect ctl sample list no longer recomputes sample summaries on every request, so polling an eval buffering many large samples (e.g. a retry's carried transcripts) can no longer stall the eval process.with_raw_response (e.g. langchain-openai), which previously failed with 'ChatCompletion' object has no attribute 'parse'. (#4341)Logging: Local eval (.eval) and JSON (.json) log files are now written atomically (temp file + fsync + rename), preventing corruption from interrupted
.eval) and JSON (.json) log files are now written atomically (temp file + fsync + rename), preventing corruption from interrupted writes such as disk-full or process crash. (#2949)/opt/inspect-restic to root-only /root/.cache/inspect, and egress scratch files no longer land in /tmp, so agents no longer see evidence of checkpointing.EvalSample and EvalSampleSummary now record turn_count and the sample's token limit (token_limit, token_limit_type, and metered token_limit_usage).samples_df gains default turn_count and token_limit_usage columns, and evals_df configuration columns gain token_limit_type.cloudflare (was cf) and model names are passed verbatim from CloudFlare's catalog (e.g. cloudflare/@cf/meta/llama-3.1-8b-instruct or cloudflare/moonshotai/kimi-k3).model_length stop reason rather than raising an error.inspect ctl sample list now shows per-sample turn count and, when a token limit is configured, its computed usage and configured ceiling.inspect ctl task cancel and inspect ctl sample cancel gain an --action option (score/error/cancel) selecting how samples are resolved, so a cancelled eval can still complete (replaces inspect ctl sample cancel --error).inspect ctl sample errors now filters errored/retried samples server-side, so triage polls of large evals no longer build and ship the full sample listing.inspect ctl sample commands given an exact full task id now stop discovery at the process hosting it, skipping busy-retry delays from older sibling processes.inspect ctl config --key NAME LIMIT retunes any named concurrency() limit mid-flight, without the registering code opting in; the config view lists the registered keys.inspect ctl config can now retune --timeout, --attempt-timeout and --max-retries mid-flight, reaching even generate calls already retrying (pass clear to restore launch config).inspect ctl sample list now caps its listing at the 100 most relevant rows by default (--limit/--all to adjust, --status to filter) and reports a complete status histogram plus a truncated flag in its envelope.inspect ctl sample events now shows the data payload of transcript().info() events in its default compact output (previously only the source was shown).inspect ctl sample events now accepts --type all (shell-safe spelling of --type '*') and --from-start to read the full backlog from the first event.inspect ctl sample events gains --limit N to cap the events returned per page (combinable with --from-start, --tail, or a cursor).--json flag for inspect eval, inspect eval-set, and inspect eval-retry emits machine-readable launch-handoff and done JSON lines on stdout, so agents can use inspect ctl immediately after launch instead of sleep-and-retrying.--detach flag for inspect eval, inspect eval-set, and inspect eval-retry runs the eval in the background — the command prints the launch record and returns once the control endpoint is bound, leaving the eval running detached from the terminal (monitor with inspect ctl task list).inspect trace anomalies and inspect trace http gain a --json flag emitting machine-readable output for agents and shell pipelines; the JSON always includes error/timeout buckets (no --all needed).inspect trace anomalies no longer errors when --filter matches an action's completion record but not its start record.--debug attach diagnostics ("Waiting for debugger attach") now print to stderr, keeping stdout clean for machine-readable output such as --json.inspect ctl sample show now reads the sample's summary and error detail in one atomic request, so its output can no longer mix data from different attempts.inspect ctl sample show now reports a cancelled sample as cancelled/pending (matching sample list) instead of error with the cancellation as its error.inspect ctl task cancel and inspect ctl sample cancel to cancel a running task or one running sample (idempotent, with --dry-run).max_retries to permit N retries (N+1 total attempts); previously it counted total attempts, so e.g. max_retries=1 never actually retried.read_eval_log, eval-retry, events_df/messages_df), a single keyed lookup instead of trying all 23 Event union members in turn; ~6.4x faster per-event validation on large logs, with no change to the log format. (#4445)sample.messages from the first model role's conversation when a solver runs multiple agents concurrently, instead of returning whichever agent's model call happened to fire last. (#4414)score_on_error) now contribute their scores to metrics and scored_samples, matching what the sample log shows. (#4412)score_reducer and task_source values in serialized task args now restore as instances rather than their registered factories. (#4374)log_dir_uri) so the viewer can reliably scope its local cache to the directory./log-headers ~3s -> ~0.3s, /logs ~1.5s -> ~0.06s).resolve_model_roles now copies a Model passed by object (e.g. via eval(model_roles=...)) before stamping its role, so roles supplied through the Python API — not just --model-role — get a distinct instance and are not misattributed. (#4464)Google: The Gemini Developer API endpoint now supports OAuth/ADC authentication (-M use_adc=true or GOOGLE_USE_ADC=true) with automatic token refresh,
-M use_adc=true or GOOGLE_USE_ADC=true) with automatic token refresh, for deployments reachable without an API key.…the old flat spellings remain as hidden deprecated aliases, except ctl sample which is now the group (use ctl sample show).
reasoning_effort with a documented default of high).gpt-5.6.reasoning_effort="max" is now passed through natively for GPT-5.6+ models rather than being clamped to xhigh.cache_write_tokens field).reasoning_mode generation option (--reasoning-mode) for GPT-5.6 pro mode; requesting "pro" defaults to background processing like the gpt-5-pro model line.Task(checkpoint=False) now vetoes checkpointing for that task, overriding an eval-set/CLI enable (previously a no-op).inspect ctl CLI into resource-noun groups (ctl task, ctl sample, ctl config, ctl process); the old flat spellings remain as hidden deprecated aliases, except ctl sample which is now the group (use ctl sample show).inspect ctl sample commands now warn and skip an eval that stays busy through the retries instead of failing outright, with stderr caveats (and an honest non-zero exit when no tasks remain visible) wherever the skip could mislead.inspect ctl sample show/sample events payload reads and process keep/release now retry a busy eval with narrated attempts instead of failing after a single short attempt.inspect ctl config now errors when the target process runs an inspect version too old to support a requested knob, instead of silently ignoring it.inspect ctl --json invocations now emit a structured {"error": {kind, exception, message, status}} envelope on stdout (exit code still non-zero) instead of only stderr prose or a raw traceback; human output is unchanged.max_subprocesses is now retunable mid-flight via inspect ctl config --max-subprocesses and the PATCH /config endpoints, alongside the other concurrency knobs.type.--model-role roles that share the same model and config now report their token usage separately in role_usage instead of collapsing onto one role. (#4450)Model API: New openai-api-completions provider for the legacy /v1/completions endpoint of any OpenAI-compatible server (raw prompts, no chat template)
openai-api-completions provider for the legacy /v1/completions endpoint of any OpenAI-compatible server (raw prompts, no chat template); shares its implementation with vllm-completions.openai-api-completions and vllm-completions accept pre-tokenized prompts via ChatMessage.metadata["prompt_token_ids"].count_tokens so compaction no longer hits 400 errors (thinking blocks "cannot be modified" / "each thinking block must contain thinking") when counting message subsets whose assistant turns contain reasoning blocks.model, missing messages/input, or a non-object body) instead of crashing the sample. (#4187)compaction and fallback content blocks (this handling was present only in the in-repo copy and missing from the injected proxy build).on_checkpoint/on_resume callbacks; on_resume may return a ResumeReport surfaced to agents via checkpointer().restored.inspect ctl tasks now includes model and solver columns.inspect ctl limits to view or retune a running eval's max_samples / max_sandboxes / max_connections concurrency limits mid-flight (with --dry-run).devices field on compose services (device mappings such as /dev/kvm), previously rejected as an unknown field.metadata as ModelOutput.metadata.output_tokens (matching the OpenAI/Anthropic convention where reasoning is a subset of output), so output-metered token limits and cost computations count Gemini thinking tokens.auto_model_class model arg to select the transformers auto-class used to load the model, so architectures not registered with AutoModelForCausalLM (e.g. the Mistral 3 series and other image-text-to-text models) can be loaded with e.g. AutoModelForImageTextToText. (#4438)token_limit(n, type="output"), Task/eval() token_limit=TokenLimit(...) values, or string forms like --token-limit output:1m (with k/m/b magnitude suffixes).stable_message_ids() linear per turn..eval logs now looks up zip members via a cached O(1) name index instead of an O(members) scan per member, removing quadratic (O(members²)) overhead when loading logs with many samples (e.g. read_eval_log, eval-retry, samples_df(full=True)).pattern match_all=True could incorrectly return the target value when no matches were present.model_graded_qa / model_graded_fact now mark a sample unscored (instead of INCORRECT) when the judge's output does not match the grade regex, tagging unscored_reason="grade_parse_failure" so judge-parse failures leave the rate and stay visible rather than inflating the INCORRECT count.read_file() staging to a generated regular file so container paths cannot copy outside the private host temporary directory.log_start() header flush when log storage is unreachable) now yields an errored EvalLog — re-queued by task retries and eval_set() — instead of tearing down the entire run and cancelling sibling tasks.--score-on-error and --continue-on-fail (when absent on the command line) silently overwriting a value set in a @task, a --run-config file, or a prior eval log being retried.mean() and bootstrap_stderr() now return 0.0 for an empty score list instead of nan (with numpy empty-slice warnings), matching the empty-input handling of accuracy()/std()/stderr()/var().max_connections / adaptive_connections settings instead of classifying the adaptive-vs-static path from task-level config alone.RequestTimeTooSkewed) retries of S3 log writes now always get their full 5 attempts instead of being cut short by a 60-second wall-clock stop.eval-retry --max-retries 0 now disables retries as documented instead of inheriting the original eval's retry policy.{type, name, params} are no longer misidentified as encoded registry objects during registry_kwargs() round-trip. (#4374)http:// instead of https:// in the running-samples port-mappings panel.Anthropic: Various changes related to Sonnet 5.
:60 seconds when the elapsed time has a fractional second.react sub-agent nested inside another react (as a tool, handoff, or deepagent task) no longer crashes under checkpointing.Checkpointing: Run restic backup with --quiet so its progress output can't overflow the sandbox output cap on long backups.
restic backup with --quiet so its progress output can't overflow the sandbox output cap on long backups.inspect ctl flush (write a running eval's buffered samples to the log now, e.g. to make S3-backed results analyzable without waiting) and inspect ctl buffer (view or change the --log-buffer / --log-shared sample-buffer params of a running eval).inspect ctl tasks no longer drops eval processes owned by another user, which were previously misreported as dead.EvalSpec.revision, resolving the actual distribution (namespace-package aware).:60 seconds when the elapsed time has a fractional second.react sub-agent nested inside another react (as a tool, handoff, or deepagent task) no longer crashes under checkpointing.Log: Shared sample buffer files synced to S3 (via --log-shared) are now tagged inspect-ephemeral=true so they can be targeted by an S3 lifecycle rule.
--log-shared) are now tagged inspect-ephemeral=true so they can be targeted by an S3 lifecycle rule..eval on a remote filesystem (e.g. S3) now fetches the per-sample journal summary files concurrently, reducing load time for logs with many samples.task_args are passed but cannot be applied to any task. (#4194)KeyError at finalisation when a provider rewrites its own model name mid-run (e.g. vLLM resolving a base:adapter LoRA spec to base).read_eval_log_samples_by_id() to concurrently read a specific subset of samples by (id, epoch) (#2873).api_key is a short-lived credential.CLOUDFLARE_API_TOKEN, HF_TOKEN) to API key override hooks rather than the derived *_API_KEY name.AZUREAI_API_KEY values to API key override hooks.api_key instead of replacing it with an environment credential.ClientPayloadError wrapping a PayloadEncodingError, e.g. a connection reset mid-body) instead of crashing the sample.together, hf-inference-providers, custom routed providers).generationConfig structured-output schema (responseSchema/responseJsonSchema) instead of dropping it.output_config.effort for adaptive thinking.inspect ctl tasks now pins each eval's reported start to its first sample's start instead of letting it drift forward as early samples finish.inspect ctl reads now use a 15s timeout and retry a busy eval up to 8 times (printing a status on each timeout) before failing with a non-zero exit, instead of silently dropping a momentarily-unresponsive eval from the listing.model and model_roles overrides for re-scoring (inspect score --model / --model-role).self_check now verifies that a large (~1 MiB) command argument round-trips correctly through exec.override_sandbox_output_limit() context manager to temporarily raise the exec-output and/or read-file size caps for the current context.git_context() now redacts credentials embedded in the git remote URL (e.g. https://user:token@host) before recording origin, preventing tokens from leaking into eval logs and downstream consumers.data tar filter when extracting sandbox checkpoint egress tarballs on the host, preventing a sandboxed agent from writing files outside the destination repo via crafted ../absolute-path/symlink entries.FileSystem.is_writeable() actually take effect, avoiding a double-separator write-test path for direct callers..json eval logs no longer parse the entire samples array, making header reads of large logs dramatically faster (e.g. a 29MB S3 log: ~47s -> ~0.4s).Task Sources: Drive a running eval from code with TaskSource.
TaskSource.notifications/tools/list_changed from a server that advertises listChanged).sandbox_client teardown so a slow or failing mcp_kill_server no longer escapes the task group as a masking "Attempted to exit a cancel scope" error.call_tool with read_timeout_seconds so a lost response surfaces to the model as a tool timeout error instead of deadlocking the tool call.mean() now maps Value to float via value_to_float() like the other built-in metrics.mean/median/pass_at/pass_k epoch reducers applying a custom value_to_float twice to dict-valued scores.task_identifier now excludes runtime-only GenerateConfig fields from model_roles configsread_eval_log, read_eval_log_async, and samples_df now accept exclude_fields for more memory-efficient loading of large samples.ComposeConfig) when an eval-level sandbox override (--sandbox <provider>) is passed without its own config._reasoning_summaries_lock once the cached value is set, so highly-concurrent generate() calls no longer queue through the lock on every call.GenerateConfig.extra_headers on the chat completions API path (previously only the conversation-api path honored it).trigger when validating fallback blocks from message history, fixing a ValidationError with anthropic>=0.110.0 (which made trigger a required field on BetaFallbackBlock).turn_limit() which tracks total generations.inspect ctl tasks now reports keep-alive status (on / off / mixed) and a new inspect ctl keep command (backed by POST /keep) latches keep-alive on a running process so it parks after its eval.on_model_retry hook, fired before each model retry backoff with the model name, attempt number, and upcoming wait_time (useful for surfacing time spent in rate limiting and other retries)./log-bytes range requests as plain responses instead of line-iterating a BytesIO.read_eval_log, read_eval_log_async, and samples_df now accept exclude_fields for more memory-efficient loading of large samples.subprocess() no longer deadlocks on timeout/cancel when asyncio's child watcher misses the process exit (observed under heavy docker compose exec load); the shielded post-kill process.wait() is now bounded.Anthropic: Support for server-side refusal fallback via the fallback_models generate config (Claude 5+ on the first-party Anthropic API).
fallback_models generate config (Claude 5+ on the first-party Anthropic API).cache_ttl model arg for specifying the prompt cache TTL ("5m" or "1h").reasoning_tokens is set on Claude 4.7+ or Claude 5 (which removed the budget_tokens control); use reasoning_effort instead.GenerateConfig.extra_headers on the chat completions API path (previously only the responses and OpenAI-compatible paths honored it).NamespaceToolParam grouping through agent bridge.ModelOutput.fallback and roll them up per-sample on EvalSample.model_fallbacks (and sample summaries); expose a fallbacks count column in samples_df().NamespaceToolParam grouping through to the OpenAI Responses API.inspect eval / inspect eval-set now bind a per-process control server (AF_UNIX, default on) exposing a read surface for the live run. New inspect ctl commands let another process (CLI, scripts, agents) observe a running eval / eval-set.internal_tool_type() helper and INTERNAL_TOOL_TYPE options key marking server-side tools.source="operator" provenance on operator-injected messages.--onedir bundle instead of a single StaticX executable.self_check now verifies that non-ASCII (UTF-8) command output round-trips correctly on exec stdout/stderr.has_fallbacks/fallbacks filter variables, sample header, transcript fallback marker and model-event badge).ChatMessage objects.OperationalError: unable to open database file when concurrent evals share a log directory.Anthropic: Explicitly classify Claude 5 (Fable/Mythos) models instead of relying on "latest"
Transcript: Continue using the realtime sample buffer database when WAL journal mode cannot be enabled.
claude_code, codex_cli).{{param}} placeholders in tool-call views.user · queued chip visible when agents are idle mid-turn.working limit event recorded alongside the real one when a message/token/cost limit is hit inside a sandboxed agent bridge (e.g. claude_code).Model API: ChatCompletionChoice.stop_details (StopDetails/StopCategory) surfaces a model's refusal/safety category and explanation when available.
ChatCompletionChoice.stop_details (StopDetails/StopCategory) surfaces a model's refusal/safety category and explanation when available.transcript().events now resolves content attachments (large text, images) instead of returning bare attachment:// references when reads are served from the bounded-history provider.OperationalError: database is locked.Model API: Add ModelInfo.family and ModelAPI.model_family(). Provider capability and request-shape checks now consult a registered ModelInfo.family be
ModelInfo.family and ModelAPI.model_family(). Provider capability and request-shape checks now consult a registered ModelInfo.family before falling back to model-name matching, while preserving the configured model name for provider requests.stream model arg (e.g. -M stream=true) to stream completions. A length-truncated streaming response (with structured output or tools) now degrades gracefully to stop_reason="max_tokens" instead of raising.response_schema (structured output) for Claude models via output_config.format.submit() calls are retained in the subagent transcript (previously stripped) and rendered as markdown, so a submit is distinguishable from a normal assistant message. The parent's result is unchanged.sandbox_service() instances running as different users in the same sandbox to share /var/tmp/sandbox-services.sandbox_service() request payloads with a higher output limit (150 MiB) than the default exec output cap.max_tokens, temperature, reasoning effort/tokens) to the resolved Inspect model, leaving these parameters entirely determined by the evaluation config. Pass forward_generation_config=True to agent_bridge()/sandbox_agent_bridge() to restore previous behavior.platform, extra_hosts, cap_add, cap_drop, security_opt, and tmpfs in ComposeService.SandboxTimeoutError now carries truncated_output with the partial command output captured before a timeout (surfaced to tool callers), instead of discarding it.INSPECT_TRANSCRIPT_BOUNDED environment variable). transcript().events remains a full, compatible view; use transcript().history for memory-aware access.ViewerConfig (passed via Task(viewer=...)) lets eval authors customize how a task's sample list, score panel, and scanner results render in the log viewer — including sample-list columns, default sort, score labels, and color scales. See Custom Views.content to avoid server validation errors.OpenAI: Don't filter tool_search tool in agent bridge for providers derived from OpenAIAPI.
tool_search tool in agent bridge for providers derived from OpenAIAPI.set_model_info() when computing input tokens for unknown models.OpenAI: Map Inspect todo_write tool to native OpenAI/Codex update_plan tool.
todo_write tool to native OpenAI/Codex update_plan tool.openai/bedrock/<model> qualifier.bash_session with realtime logging enabled (the default), caused by a non-serializable sandbox object being written to the sample store. The sandbox is now re-resolved per call rather than cached in the store.Anthropic: Render mid-conversation system messages as user turns for pre-4.8 models.
<system-reminder> user turns for pre-4.8 models.role: "system" messages from bridged Anthropic-API clients.GenerateConfig.cache to all models.memswap_limit in ComposeService.OpenAI: Support tool_search tool type from native scaffolds (e.g. Codex CLI).
tool_search tool type from native scaffolds (e.g. Codex CLI).OpenAI: Drop tool_search tool types from native scaffolds (e.g. Codex CLI).
tool_search tool types from native scaffolds (e.g. Codex CLI).Deep Agent: Support for running subagents in the background.
agent dispatch tool now ships with a tool-call viewer.task to agent.background option to prompt the model to use nohup for long-running commands.reasoning_content in addition to reasoning_details for Deepseek v4.Score.answer on model_graded parse failure.--stream N batch of scored samples.resolve_plan().write_eval_log(..., header_only=True) now preserves on-disk samples for JSON logs and remote .eval logs.agent_name for samples (always consult plan).agent_with(name=...)) no longer overrides its registry identity.Anthropic: Update model database / feature enablement for Opus 4.8.
cache-diagnosis-2026-04-07).is_live property to agent channel for detecting whether channel is in use.Anthropic: Improved merging of beta headers.
`ask_user()` tool: model can solicit a structured answer from the operator.
ask_user() tool: model can solicit a structured answer from the operator.notify_user() tool: model can send status notifications to the operator.request_input() public API: programmatic structured prompts from solvers, agents, or tools, using the same dispatch surfaces as ask_user().<think> blocks without exposing inner reasoning text as visible model output.OpenAI: Backfill required 'query' field when service only provides 'queries'.
model_event_sink option for bridge to control order of model event emission (used to nest model events properly in agent spans).Agent Intervention which provides the ability observe a running agent, interrupt it, and redirect it with follow-up messages.
@tool(parallel=True). Builtin tools that are atomic per call (e.g. bash, python, memory, read_file, web_search) are now registered with parallel=True.--reasoning-effort now works uniformly across Claude 3.7–4.5 and Gemini 2.5. Inspect automatically bridges effort to a budget_tokens / thinking_budget value using a fixed table (minimal=2048, low=4096, medium=10000, high=16000, xhigh/max=32000). Previously these models silently ignored reasoning_effort.minimal / xhigh / max) to the low / medium / high tier accepted by upstream APIs for Groq, Ollama, and SageMaker.max effort to xhigh before forwarding (OpenRouter does not accept max).max → xhigh mapping for reasoning_effort.should_retry() so the model-layer tenacity loop can retry network failures that escape the SDK's one-shot inline retry.run_coroutine() (and the sync log/analysis helpers built on it) now honour INSPECT_ASYNC_BACKEND=trio when called with no running event loop, rather than always using asyncio.AsyncFilesystem: Add iter_files() and iter_dirs() methods.
iter_files() and iter_dirs() methods.pass_k reducer for computing the probability that all k epoch attempts succeed (τ-bench reliability metric).get_model() with role.redacted=False when reasoning content, summary, and encrypted exists.responses_phase parameter.id properly).fastapi and uvicorn are now required dependencies.Config: Add inspect log export-config command to export a run config from an existing log file.
inspect log export-config command to export a run config from an existing log file.get_file() and exists() methods.Scanners: Declare Scanner import in a way that's compatible with pyright type checking.
OpenAI: Add GPT 5.5 as computer use model and exclude 'chat' and 'instant' models from computer use.
reasoning_details in OpenAI-compatible responses.extra_body fields from Message response.openrouter/anthropic/* models.top_k correctly for Nova models.prompt_logprobs support in chat mode via GenerateConfig, parse prompt logprobs from completion mode responses, enabling perplexity() and target_perplexity() scorers end-to-end.--adaptive-connections is now enabled by default (defaults to 100 per model connection).count_tokens() (they are already retried and gated by max_samples).suspend_token_limit() context manager for suspending token tracking and limit enforcement within a scope.hf_dataset retries transient Hugging Face errors (rate limits, timeouts, Hub-unreachable cache misses) up to 3 times (5 in CI) with exponential backoff. Pass retry=False to disable.str() coercion.None is treated (converted to "").hf_dataset(..., shuffle=True) to EvalDataset.shuffled.ToolError if there is a null byte in command input.match(numeric=True) no longer matches digit-substrings (e.g. target 5 against 25); now correctly handles negative, decimal, and scientific-notation targets, and recognises unicode-formatted numbers (unicode minus, vulgar fractions like ½, Chinese numerals, fullwidth digits) in both targets and model output.match(numeric=True, location="exact") is now strict — values like "5 some text" no longer match target "5".evals_df() column name when there are multiple reducers.registry_add()).--run-config option to inspect eval for single-file run configuration.eval_set (CLI --scanner / ScannerConfig). Scans incrementally as logs land, reuses prior results across resumes, and renders progress alongside the existing eval view.IndexError inside resolve_tasks after passing an empty task list to eval.score_display argument to eval_set() function.log_file_info() robust to non-standard filenames; added log_file_info_async() / log_files_from_ls_async() so view-server header reads don't block the event loop.inspect_ai module.INSPECT_PY_LOGGER_FORMAT env var (rich/plain/json) for non-TTY-friendly single-line console logs.COLUMNS and LINES for dumb terminals.GenerateConfig fields with an error.None as default (e.g.x: dict = None).OSError does not include .strerror and .filename[Content, str]."" as an answer for `basic_agent().{submit} in react() agent prompt templates (rather than using .format).pd.NA when converting scores to float in analysis df functions.summaries.json when re-scoring or converting logs.fail_on_error fractional threshold now uses the sliced sample count multiplied by epochs (matching the end-of-run check)eval_retry preserves SampleScore.scorer attribution on restored samples (was None instead of the scorer name).time_limit() no longer masks exceptions raised after the deadline (e.g. from finally blocks) with LimitExceededError; the original exception now propagates.multiple_choice()/answer()/choice() now accept lowercase letters, prefer the last ANSWER: occurrence (matching the CoT "last line" instruction), parse "A, B and C" lists, and handle multi-answer targets with separators (Target("A,B")).max_score reducer's dict/list paths now NaN-filter per key/index, making results order-independent (NaN previously kept whichever element came first).pass_at(k) now returns the unscored NaN sentinel when fewer than k epochs were scored (previously inflated to 1.0).create_reducers no longer rewrites custom reducer names ending in _<digits> (e.g. "top_5") into the built-in _k shorthand.multi_scorer returns Score.unscored() when every sub-scorer declines to score (previously crashed with IndexError).at_least / pass_at reducer names now correctly include the _k suffix on the returned reducer (previously leaked onto the module-global factory).model_graded_qa() / model_graded_fact() — default grade pattern now extracts the last GRADE: $LETTER in grader output.f1() / exact() — targets with leading whitespace are no longer silently skipped.math() — brace-delimited set answers like {1, 2} are now compared as multisets.math() — model outputs containing \boxed{inf} / Infinity / -inf no longer crash the scorer with OverflowError.grouped() — raise instead of silently overwriting a per-group metric when a group name collides with all_label.stderr(cluster=...) — return 0 with a single cluster (was NaN/inf from divide-by-zero).value_to_float() — reject "nan"/"inf" string values so a single non-finite Score doesn't poison accuracy() and friends.inspect log recover — preserve sample.uuid for crashed in-progress samples (initial buffer summary now carries state.uuid; recovery synthesizes a fallback uuid for legacy buffer rows).Anthropic: Skip the top-level cache_control auto-caching field on Bedrock and Vertex where they are not supported.
cache_control auto-caching field on Bedrock and Vertex where they are not supported.tool_call_id in tool responses (parallel tool calling).reasoning_effort for Grok 4 models.ToolError when timeout error occurs in MCP tool call.IncompleteBody errors during multipart uploads under concurrent flushes.list[ContentText].Inspect View: Fix extraneous console errors
Google: Support Gemini 3+ native web search and code execution alongside function tools.
trust_remote_code model argument (defaults to False).vllm-completions model provider that uses completions rather than chat endpoint.retry_immediate now defaults to True. Pass retry_immediate=False (or --no-retry-immediate) to restore the previous batch-retry behavior.max_tasks now defaults to the greater of 10 and the number of models being evaluated (was previously 4).model_usage and role_usage from previous log when performing retries (matches existing eval-retry behavior).usage.input_tokens (currently OpenAI Responses with store=false + include=["reasoning.encrypted_content"]).ToolEvent.message_id to reference the correct ChatMessageTool when multiple tool calls occur in one assistant turn.Add adaptive connections option to automatically tune model API concurrency between configurable bounds based on rate-limit feedback.
fail_on_error threshold).media_resolver() context manager for scoped URI resolution for media reading (images, audio, etc.).download() and gdrive_download() helpers for fetching external files with SHA256 verification, caching, and transient-error retry. gdrive_download() requires the optional gdown dependency, installed via pip install inspect_ai[gdown].stop_reason='model_length' in react() agent, force-compact before falling through to the overflow filter.working_time path in SampleSummary columns.scorer and scorer_args to ScoreEventOpenRouter: Escape signature attribute in tag round-trip.
docker pull for service images already present in the local Docker daemon.input_tokens for gpt-5.4, gpt-5.4-pro, gpt-5.5, gpt-5.5-pro (was 778000, now 922000 = 1,050,000 context − 128,000 output).Your coding agent can read these notes before it upgrades. Set up the MCP server →