NewYour coding agent can read the release notes before it upgrades.Set up the MCP server →
PyPI · #4199 most downloaded on PyPI
Collection of large language model evaluations
Last release 2 days ago
02 Oct 2026
Ships on a steady schedule
a new release about every 2 weeks
Most releases are documented
notes for 30 of 34 stable releases
Nothing withdrawn
no release was ever pulled
11 months old
34 releases · first in 2025
One column per month.
AgentHarm: Harmfulness Potential in AI Agents: Remove an unreachable float-type guard from avg_full_score so it filters samples the same way as the si
AgentHarm: Harmfulness Potential in AI Agents: Remove an unreachable float-type guard from avg_full_score so it filters samples the same way as the sibling metrics. No effect on reported scores. (@EastZeus)
Humanity's Last Exam: Register the inspect_evals/hle/regex_judge and inspect_evals/hle/json_judge scorers when the package is imported, so inspect eval inspect_evals/hle no longer fails with a registry lookup error at task construction.
MBPP (v4-A): Completion extraction now handles a python3 label, an unclosed fence, a fence after prose on the same line, a closing fence glued to code, and empty or stray fences. These completions used to be scored on the raw text or an empty block. Blocks labelled py and python are now taken in order of position.
Documentation: The site follows the reader's system light or dark setting, and a navbar toggle overrides it. Colours come from a single brand file with light and dark values, and page styles use the theme's colours instead of fixed ones, so code blocks, lint badges and search highlights read correctly in both.
The lint results dashboard supports the register-v2 policy and inspect-evals-lint schema version 3, where suppressed checks count as met, and shows how many checks were suppressed beside each count. Existing register-v1 reports remain supported.
The lint results dashboard supports the security category and inspect-evals-lint schema version 2. Existing reports remain supported.
extract_code_block in inspect_evals.utils.code takes an optional accept predicate and returns the first block it accepts, falling back to the first block. The new defines(name) predicate accepts a block that parses and defines the required function, so an eval can skip a usage example or a draft that does not parse. The parser fixes above apply to every caller.
New extract_code_blocks returns every language-labelled block in order, or else every unlabelled one, for evals that take the last block or all blocks. extract_code_block is built on it. The cpp language also matches cc and cxx labels.
Add shuffle_and_seed to inspect_evals.utils, which splits a shuffle: bool | int task parameter into a loader's shuffle and seed arguments, so a task can default its shuffle to a seed without a separate seed parameter. BEST_PRACTICES.md now recommends it for datasets whose records are grouped.
A call that omits either emits a DeprecationWarning and skips the count check; both arguments will become required in a later release. The content che…
PaperBench (v5-C): Criteria the judge could not grade are no longer scored as criteria the reproduction did not meet. They are left out of the weighted rubric average on both sides of the division, after bounded retries: judge replies are classified before use, so a refusal, an empty reply, or a file ranking that names nothing readable is retried rather than parsed as an essay or graded against no files, and a grading prompt that overflows the context window is retried once with half the file budget before erroring the sample. A rubric with no gradeable criterion returns Score.unscored(), and failures of the machinery around the judge error the sample. Score.value is now a dict with keys score and graded_weight_fraction, the second reporting how much of the rubric was graded. The score mean can move in either direction: per-sample scores rise when a judge reply was ungradeable, and samples whose grader call raised now leave the denominator.
AgentBench (v4-B): Add sandbox_type/sandbox_config provider configuration following the cross-eval sandbox-config pattern (#1115).
PaperBench (v3-C): Add a task-level sandbox_config: SandboxEnvironmentSpec override for non-default providers (#1115).
Mind2Web-SC (v3-B): Load the full 200-sample dataset instead of truncating to the first 10 samples via hardcoded slice, and disambiguate sample IDs across ALLOW/DENY variants sharing the same task annotation.
HLE: original_accuracy now follows CAIS's rounded percentage calculation using the selected dataset size after dataset filters and before run limits. The dataset size is recorded in sample metadata, so omitted questions and missing judgments remain in the denominator for static and rolling datasets. This metric requires single-attempt results and dataset-size metadata; other metric calculations are unchanged.
HLE: the default run config now evaluates the HLE-Verified gold subset (only_hle_verified_gold: true); original.yaml keeps the full CAIS dataset. Default-run scores are not comparable with earlier versions.
HLE: select metrics and their arguments through the scorer spec's metrics field in run configs. Metric factories resolve through Inspect registration. Score calculations are unchanged, but each shipped config now selects its own metrics: the default run reports hle/accuracy, hle/stderr, hle/cerr and hle/unscored and no longer emits hle/original_accuracy; original.yaml reports hle/original_accuracy, hle/stderr, hle/cerr and hle/unscored. Direct use of the HLE judge factories now requires metric selection on the containing task. Scorer and metric names in the spec are Inspect registry names, and everything HLE defines is registered under an eval-qualified name (inspect_evals/hle/regex_judge, inspect_evals/hle/accuracy, ...), so the score columns are now hle/regex_judge and hle/regex_judge1 (hle/json_judge under original.yaml) and every metric row carries the hle/ prefix.
PaperBench (v4-B): Update blacklist URL monitor to identify scp-like SSH clones git clone git@github.com:owner/repo.git. Deviates from original paper implementation, however, as the paper implementation outlines that the blacklist monitor is used in addition to human checking, this change is helps with the human check and aligns with the intent of the paper.
BoolQ: Exploring the Surprising Difficulty of Natural Yes/No Questions (v4-A), PAWS: Paraphrase Adversaries from Word Scrambling (v4-A), MaCBench: Probing the limitations of multimodal language models for chemistry and materials research (v2-A): A completion that does not match the answer pattern is now scored INCORRECT with reason="invalid_response_format", rather than NOANSWER. This follows pattern()/answer() upstream in inspect_ai#4629, which attributes an unmatched completion to the model under test. Headline metrics are unaffected: value_to_float maps both NOANSWER and INCORRECT to 0.0, so accuracy and stderr are unchanged on every run. What moves is where the non-answer is recorded, from the score value to Score.reason; log analysis that counts NOANSWER values should read reason instead.
BBH: Challenging BIG-Bench Tasks (v3-A), ChemBench: Are large language models superhuman chemists? (v3-B), ∞Bench: Extending Long Context Evaluation Beyond 100K Tokens (v3-A) and WorldSense: Grounded Reasoning Benchmark (v3-A): The same pattern()/answer() change as above. It reaches BBH's binary-choice subsets, ChemBench's multiple-choice questions, infinite_bench_code_run and all of WorldSense; the other scorers in those evals are unchanged. Headline metrics are unaffected for the same reason.
KernelBench: Use a standalone uv project to pin sandbox dependencies independently of the host lockfile, so host dependency updates do not require rebuilding the published image.
KernelBench: Remove 34 host-only packages, including inspect-ai, from the sandbox image while preserving all 148 retained package versions. The pinned default image is now the reduced 2026-09-17 build, validated on a Tesla T4 through the unchanged runner.
KernelBench: Add the kernelbench_sandbox_check maintenance task, which runs the eval's own scorer over fixture kernels with known verdicts to verify a sandbox image on GPU hardware, plus an opt-in pytest wrapper and a Hawk example.
DS-1000: Add the ds1000_sandbox_check maintenance task, which runs the eval's scorer over fixture solutions with known verdicts inside the pinned sandbox image and checks the pinned packages are installed; accuracy 1.0 certifies the image.
WorldSense (v4-A): Derive the sample id from the dataset's own unique Key. The previous id, (tuple_ID, problemname), is shared by every trial of a tuple, so filter_duplicate_ids kept one trial per tuple and dropped 46,872 of the 87,048 trials. Because the kept trial is the first in file order, the evaluated set held no sample whose gold answer was FALSE, IMPOSSIBLE or 2, while POSSIBLE rose from 23.1% to 50.0% of the set and TRUE from 15.4% to 33.3%.
StereoSet: Measuring stereotypical bias in pretrained language models (v4-A): Use the dataset's own per-record id as the sample id. The previous id hashed (context, target, bias_type), which several records share while presenting different candidate sentences and gold labels, so filter_duplicate_ids silently dropped 8 of 2,123 intersentence records and 40 of 2,106 intrasentence records.
SciKnowEval: Evaluating Multi-level Scientific Knowledge of Large Language Models (v4-B): Build the sample id from every field that distinguishes a record (question, domain, task, subtask, choices, answer key, answer) hashed to 16 hex digits. The previous id hashed only (question, domain, task), which 3,394 records share with another record while differing in choices, answer or subtask, so filter_duplicate_ids silently dropped them and the full task evaluated 66,386 of 70,196 rows. It now evaluates 69,736; the 460 rows still removed are identical to another row on every field.
MMLU: Measuring Massive Multitask Language Understanding (v4-A): Include the answer key in MMMLU sample ids. Four pairs of translated rows (FR_FR, IT_IT and two in AR_XY) have identical question text and options but different answers because the translation collapsed two distinct English questions, and the previous id dropped the second row of each pair as a duplicate. Both rows are now evaluated, and the id is hashed to 16 hex digits because the all-languages default config has 196,588 rows. The 5-shot ids also include the subject, because the few-shot examples depend on it, so the 78 questions listed under two subjects become two prompts in mmlu_5_shot (14,015 samples) while mmlu_0_shot keeps one copy (13,937).
BBH: Challenging BIG-Bench Tasks: Exclude the two unscorable ruin_names rows by id through drop_known_broken, with the upstream report (suzgunmirac/BIG-Bench-Hard#19) recorded beside each id, instead of matching the question text inside the prompt. The same two rows are excluded as before, so results and the task version are unchanged. The README gains a "Known dataset issues" section and dataset_samples is corrected from 250 to the 6,509 samples the default task loads across all subsets.
Fix docker_handling="force_build" for digest-pinned sandbox images. Local builds use the image tag; default and forced-pull runs retain the configured digest.
Evaluation checks now use inspect-evals-lint 0.2.1, installed through the dev dependency group. Run uv run inspect-evals-lint <eval_name>. The shared utils package is linted as a helper alongside the evaluations. Existing unresolved model roles remain warnings through an allowlist.
Add scorers_from_spec, metrics_from_spec and with_metrics to inspect_evals.utils, so a task can take its scorer and metric selection as run-config data (a scorer spec dict with args, graders and metrics) and resolve it through Inspect's own loaders. Extracted from the HLE metric configuration work as a first step towards upstreaming to Inspect.
Documentation: Redesign the homepage with Inspect Evals branding and two research cards randomly selected from four papers on each page load. The evaluation catalogue separately marks 35 evaluations as Featured and offers controls for alphabetical order, oldest or newest additions, and featured evaluations first. Addition dates are generated from repository history. Use the same navigation and site search across all pages, and render GitHub-style alerts as callouts. Add space between page titles and summaries. Link to Generality Labs audits at the top of the FORTRESS and BixBench documentation.
Migrate the interim metadata["unscored_reason"] key to the first-class Score.reason field across air_bench, anima, hle, livecodebench_pro, moru, niah, stereoset, strong_reject and utils/scorers.py. Score.reason shipped in inspect_ai 0.3.261 (inspect_ai#4629); the interim key was chosen in #2186 precisely so this migration would be one line per site, and inspect_ai's read shim lifts the old key, so logs written before this change keep reading correctly. Finer-grained detail keys (grader_failure_mode, scoring_failure_mode, grader_exception, grader_payload) stay in metadata as this repo's own vocabulary. StereoSet's stereotype_score, stereotype_score_coverage and invalid_answer_rate now read Score.reason. Scores are otherwise unchanged. (@Abelo9996)
Raise the inspect_ai floor to 0.3.261 and pin the development dependency to a release instead of tracking git main. The lock had been pinned to an August commit that predated inspect_ai#4629, so the repository had never resolved the released behaviour; #2314 asks for this migration to be reviewed against a release, and #2187 showed an unreviewed lock bump to main can silently change scoring behaviour.
Enable the inspect-evals-lint unscored_reason check, which was held back for this migration. It requires a reason on every Score.unscored() call and rejects the former metadata["unscored_reason"] key.
The manual Docker Images for Evals workflow dispatch accepts an evals input to rebuild and publish a subset of the image-backed evals instead of all of them (landed in #2462).
eval.yaml gains a metadata.requires mapping for environment capabilities. requires_internet moves to requires.internet, and requires.gpu (either true or {count, products}) declares that an eval needs a GPU where its sandbox runs; KernelBench, ComputeEval and GDM Self-Proliferation now declare it. Task entries gain kind: maintenance for tasks that score the harness rather than the model, starting with healthbench_meta_eval. See the eval.yaml reference in CONTRIBUTING.md.
Add inspect_evals.utils.sandbox_check, a shared helper for building sandbox image checks from fixture answers, the eval's own scorer and an expected package inventory.
Evaluation checks now use inspect-evals-lint 0.3.1. Rules have codes (IEFS001, IECQ001, IETS001, IEBP001) and report one finding per site, so the lint workflow annotates pull requests at the offending lines. Suppression comments are # inspect-evals-lint: ignore[<rule>]; the former # noautolint comments and .noautolint files are gone, replaced by exclude and per-file-ignores in pyproject.toml. Run uv run inspect-evals-lint --explain <rule> for a rule's documentation.
Evaluation checks now use inspect-evals-lint 0.4.1, which adds two warning-only rules. dockerfile_locking (IEBP007): sandbox Dockerfiles should consume a committed lock or hashed snapshot, pin FROM and COPY --from images by digest, and not pipe installers into a shell. Twenty-four evaluations currently warn; the lint job stays green while they migrate. suppression_syntax (IECQ005) warns about a suppression marker the linter does not read (a noautolint comment, a bare ignore, an unknown rule name) instead of stopping the run. Run uv run inspect-evals-lint --explain IEBP007 for the rule's documentation.
Inspect Evals: Refresh the README with project branding, prominent documentation links, two quick-start actions and issue-reporting guidance. Remove the embedded evaluation catalogue while retaining collapsible setup instructions.
Inspect Evals: Align the README, contributor guide and register update instructions on the pre-approved contributor policy.
Inspect Evals: Include shared documentation styles in the versioned theme stylesheet so cached styles do not hide the homepage button colours and borders after an update.
Inspect Evals: Repair documentation links in the GDM self-proliferation, MMLU and MASK guides.
Inspect Evals: Publish the approved contributor table and link it from the contribution guidance.
The Register Lint Results page can expand each evaluation into its per-rule results: rules grouped by category at the status they reported, linking to their documentation, and each finding linking to the file and line at the registered commit with the linter's hint.
Each register evaluation's page ends with a "Static checks" section showing the same per-rule lint results for that evaluation.
The lint results page and the evaluation pages also cover the evaluations in this repository, linted at the commit the documentation was built from.
Register submissions get an informational inspect-evals-lint check: the pinned upstream commit is linted with the register preset and the per-rule results are posted as a PR comment, the same checks the lint badges and the lint results page show after merge. A clean pass is not required.
Migrate PaperBench's judge from the interim metadata["unscored_reason"] key to Score.reason, completing the migration in #2459, which landed minutes before the PaperBench judge change that introduced this last use of the old key. The Lint Evaluations workflow had failed on every PR since. Scores are otherwise unchanged.
Add guidance to BEST_PRACTICES.md and the evaluation checklist on reporting dataset defects to the dataset's maintainers: where to report, what to include, and how to record the workaround here in the meantime. Prompted by the duplicate-id audit that followed #2506, which surfaced conflicting answer keys in SciKnowEval and MMMLU that had been hidden by filter_duplicate_ids.
filter_duplicate_ids now requires max_duplicates and reason keyword arguments and raises DuplicateIdError when samples sharing an id differ in input, choices or target, or when more samples would be dropped than acknowledged. The keep-first filter had silently discarded 46,872 WorldSense trials (#2507), 48 StereoSet records (#2524) and 3,394 SciKnowEval records (#2525) whose ids were built from a subset of fields. An audit of all 18 call sites (agent_artefacts/duplicate_id_audit/) found the remaining sites drop only exact repeats or nothing at all. The eleven no-op calls are removed (wmdp, winogrande, race_h, stereoset, worldsense, all five cyberseceval_4 loaders); the seven that drop exact repeats (mmlu, sciknoweval, sycophancy, bbeh, cyberseceval_2, livecodebench_pro, abstention_bench) now state how many and why. dataset_samples at those seven evals is corrected to the post-filter count, per the evaluation checklist; that is a metadata fix and does not bump task versions.
Add drop_known_broken(dataset, broken={id: report_url}) for samples that cannot be scored and have been reported upstream. It excludes by id, requires a report URL per id, and raises KnownBrokenSampleError if a listed id is not in the dataset, so a revision bump or an upstream fix cannot silently void the exclusion. BEST_PRACTICES.md and the evaluation checklist describe when to use it; bbh is the first adopter in a follow-up.
Use inspect-evals-lint 0.7.0, which adds duplicate_filter_acknowledged (IEBP008) and known_broken_reported (IEBP009): every filter_duplicate_ids() call must state max_duplicates= and a reason= linking the upstream report, and every drop_known_broken() entry must map its id to a report URL. Both are listed in AUTOMATED_CHECKS.md. abstention_bench suppresses the first with its reason in a comment, since its repeats are in third-party datasets that have not been reported upstream yet.
filter_duplicate_ids accepts calls without max_duplicates= and reason= again, so the change in #2528 is not breaking for code outside this repository that imports it. A call that omits either emits a DeprecationWarning and skips the count check; both arguments will become required in a later release. The content check is unchanged: samples sharing an id that differ in input, choices or target still raise DuplicateIdError, whatever the call passes. Inside this repository the duplicate_filter_acknowledged lint rule continues to require the full form.
KernelBench (v6-C): Add k8s sandbox support with a pinned public image, sandbox_type="k8s" , custom sandbox_config , sandbox_image , and sandbox_node_
KernelBench (v6-C): Add k8s sandbox support with a pinned public image, sandbox_type="k8s", custom sandbox_config, sandbox_image, and sandbox_node_selector task parameters. Fix both Docker and k8s scoring to use the sandbox environment's Python and upload the matching runner, replacing the previous execution path where every sample could fail because the selected interpreter lacked KernelBench dependencies. Change the Docker default from an eight-CPU local build to the same four-CPU, digest-pinned published image used by k8s; a separate local-build Compose config remains available for Dockerfile development and hosts that cannot use the published image.
HealthBench: Evaluating Large Language Models Towards Improved Human Health (v3-B): Fix grader response parsing to strip unlabelled markdown code fences. Only fences labelled as json were stripped before, so a grader that wrapped its judgment in an unlabelled fence failed JSON parsing and every rubric item was silently graded as not met. healthbench_score, and the bootstrap_score, criteria_met_rate and axis and theme subset scores derived from it, can move in either direction: up where a bare-fenced true on a positive-point criterion was previously dropped, down where a bare-fenced true on a negative-point criterion previously escaped its penalty. (@EastZeus)
MBPP (v3-A): Fix completion extraction for Python-labelled and unlabelled Markdown code fences, including case-insensitive labels and CRLF line endings.
PAWS (v3-A): Replace the includes() scorer with a pattern() that accepts only a bare Yes or No (surrounding whitespace and punctuation allowed). Previously any completion containing "yes" or "no" as a substring was scored correct, including hedging text ("I don't know") and answers wrapped in prose ("The answer is: No"); both now score 0.
BoolQ (v3-A): Replace the answer pattern with one that accepts only a bare Yes or No, optionally surrounded by whitespace, punctuation, or underscores. This is a tightened scorer, but not strictly tighter. Previously any completion ending in "yes" or "no" plus at most one character was scored against the target, including words such as "know" and answers wrapped in prose ("The answer is No."); both now score 0. In the other direction, answers with more than one wrapping character or a trailing newline, such as **Yes!!** or "No\n", were previously rejected and are now accepted.
AgentBench (v4-A): Fix the agent_bench_os test split ground truth (#2420). Task fixtures were built in /home while the scorer derived the expected answer from /home/agent, making every task with a relative-path init script unsolvable. The agent home is now the working directory for the build, the agent and the scorer; example and check scripts run in a login shell like the agent's bash tool; the upstream size check accepts du -h output; twenty-seven buggy example solutions (ls -a counting . and .., wc total lines doubling sums, an unset $TARGET_DIR, and off-by-one or symlink-excluding find invocations) were corrected; blank lines in create.init.code no longer emit an empty RUN. Tasks 27, 37 and 41 run their init at container start so their fixtures exist when the agent runs; tasks 36, 39, 43, 51, 55 and 121 are removed because they have no well-defined answer, leaving 119 test tasks. Ambiguous descriptions are clarified to match their reference solutions and the system prompt now asks for the bare final answer (running any requested script and submitting its output). HOME is /home/agent like the working directory. The upstream repository clone (reference solutions and check scripts, which models were found reading) is blanked in the image; the check scripts are vendored and staged at scoring time, and three of them are corrected (task 20 ran one of six cases, task 19 skipped hidden files, task 35 accepted a world-readable directory in the wrong group). Adds an oracle test that scores every test-split example solution.
GDM In-House CTF (v6-B): Remove the count-only epochs task parameter; use Inspect's --epochs runner option while preserving the task's at_least_1 reducer.
GDM Dangerous Capabilities: Self-reasoning (v4-C): Remove count-only epochs task parameters and use Inspect's runner override while preserving the task-defined reducer.
GDM Stealth (v4-B): Remove count-only epochs task parameters and use Inspect's runner override while preserving the mean, median, and max reducers.
PersistBench (v2-B): Remove count-only epochs task parameters and use Inspect's runner override while preserving each task's default count and max-score reducer.
ANIMA (v6-D): Remove the count-only epochs task parameter; use Inspect's --epochs runner option while preserving the task's default count and reducer.
APPS (v2-B): Remove the num_epochs and epoch_reducer task parameters; use Inspect's --epochs and --epochs-reducer runner options while preserving the task's mean and max reducers.
ComputeEval (v1-B): Remove the num_epochs and epoch_reducer task parameters; use Inspect's --epochs and --epochs-reducer runner options.
CyberSecEval_2 (v4-B): Remove count-only epochs task parameters and use Inspect's runner override while preserving each task's default count.
CYBERSECEVAL 3 (v3-B): Remove the count-only epochs task parameter and use Inspect's runner override while preserving the task's default count.
CyberSecEval 4 (v5-C): Remove count-only epochs task parameters and use Inspect's runner override while preserving each task's default count.
GPQA (v3-D): Remove the count-only epochs task parameter; use Inspect's --epochs runner option while preserving the task's default count and reducer.
MLE-bench (v7-G): Remove count-only epochs task parameters and use Inspect's runner override while preserving each task's default count.
MORU (v3-B): Remove the count-only epochs task parameter; use Inspect's --epochs runner option while preserving the task's default count and reducer.
StrongREJECT (v3-B): Remove the count-only epochs task parameter; use Inspect's --epochs runner option while preserving the task's default count and reducer.
PaperBench: Remove the "work in progress" label from the eval title and README now that onboarding under #334 is complete, and add an Evaluation Report section stating why no results are published and which paper numbers to compare against. (@MattFisher)
APPS (v3-B): Align the documented generation budget with output-token metering.
Add a shared extract_code_block helper for language-labelled and unlabelled Markdown code fences.
Autolint gains a model_role_resolution check: it warns when get_model(role=...) has no explicit model, no default= and no required=True, since an unbound role in that state silently falls back to the model under evaluation and a grader ends up grading its own output.
Contributor tooling now reads files as UTF-8, so tools/run_autolint.py --all-evals and make check no longer abort with a UnicodeDecodeError on platforms whose default locale encoding is not UTF-8 (for example Windows cp1252).
Add a browser-loaded table of static lint results for registered evaluations, with commit and freshness checks, and document badges published by inspect-evals-actions. (@MattFisher)
Upload rendered documentation from PR builds as a downloadable artifact for local preview. (@MattFisher)
CyberGym: a proof-of-concept that never terminates is killed by the executor and answered with a 504, which reached the scorer only as a missing exit_
CyberGym: a proof-of-concept that never terminates is killed by the executor and answered with a 504, which reached the scorer only as a missing exit_code and so read the same as a dead executor. The scorer now reads the HTTP status and reports the timeout as one. Behaviour is unchanged, both still raise (#1972).
AIME 2024 (v5-A), AIME 2025 (v5-A), AIME 2026 (v3-A): Fix scorer only checking the literal last line of the completion, missing answers boxed inside a closing display-math block. The scorer no longer mutates state.output.completion, preserving the model's full raw completion in the persisted eval log.
AGIEval (v3-A): Fix agie_sat_math accepting a fewshot_seed argument and silently dropping it, so -T fewshot_seed=N never changed fewshot selection for that task. Its eight sibling tasks already forwarded the parameter. Runs using the default seed are unaffected.
Humanity's Last Exam (v6-D): BREAKING (comparability): replace the default judges. The primary grader role moves from openrouter/google/gemma-4-31b-it to openrouter/z-ai/glm-5.3-flash, pinned to fp8 provider quantization so OpenRouter cannot route the judge to another precision, and grader_2 from openrouter/google/gemini-3.6-flash to openrouter/google/gemini-3.7-flash, both selected in a judge comparison against a frontier-panel consensus. Both judges' max_tokens rises from 16384 to 32768 as headroom for the reasoning tokens these judges spend inside the completion budget. Default-run scores are not comparable with earlier versions; original.yaml and --model-role overrides are unaffected.
Humanity's Last Exam (v6-D): BREAKING (comparability): default.yaml grades with regex_judge — this port's pre-5-C paraphrased prompt, GRADE: C/GRADE: I parsing and the candidate's last stated confidence — instead of the official prompt with structured output, so default-run scores stay comparable with this port's pre-5-C runs. This reverses the 5-C default-prompt change for the maintained default only, and gives up the official prompt's answer-extraction step and numerical-tolerance clause. original.yaml keeps the official json_judge.
Humanity's Last Exam (v6-D): BREAKING (interface): the accuracy and stderr metrics move onto each judge's own results entry, listed first, so the log viewer's headline is accuracy rather than cerr. They were previously score/accuracy and score/stderr on a separate score entry; update log analysis that reads that entry. Values are unchanged except when nothing was scored, where both are now NaN rather than 0.0.
Humanity's Last Exam (v6-D): Replace the judge_prompt, graders, and max_grader_attempts task parameters with a single scorer spec dict resolved by the new scorers.py. The judge score columns are renamed from llm_grader/llm_grader1 to the scorer names (regex_judge/regex_judge1 on the default run, json_judge under original.yaml); update any log-parsing that keys on the old names.
Humanity's Last Exam (v6-D): Grader roles now resolve through inspect's get_model role machinery (requires inspect_ai >= 0.3.259). A bound judge inherits generation settings it leaves unset (e.g. max_tokens, temperature) from its default.yaml stanza — previously only max_tokens gap-filled, so a bound judge without an explicit temperature now judges at the stanza's temperature 0. A grader role with no stanza raises inspect's required-role error when unbound instead of an HLE-specific one.
Humanity's Last Exam (v6-D): -T rolling=true can now be combined with only_hle_verified_gold=true: the run keeps the intersection, the gold questions the rolling set retains. Previously the combination raised a ValueError.
Humanity's Last Exam (v6-D): Judge score metadata is now judge_model, judge_attempts and, for unscored samples, unscored_reason. The judge_prompt, judge_attempt_outputs, grader_failure_mode and grader_exception fields and the judge-resolution INFO log are removed; update log analysis that keys on them. regex_judge records the same keys and tags a grade-parse miss unscored_reason: grader_failed, the repo's shared vocabulary, instead of model_graded_qa's grade_parse_failure.
Humanity's Last Exam (v6-D): Add the answer_type task parameter to filter the dataset to multipleChoice or exactMatch questions. Defaults to None (both types); combines with the other dataset filters.
Humanity's Last Exam (v6-D): Add the system_prompt task parameter, an id into the new SYSTEM_MESSAGES store in prompts.py (default system_mc, the prompt the CAIS leaderboard sends to every question), so run configs state which system prompt every sample carries. The run configs also gain a solver section (inspect's stock generate with tool_calls: "none"; the eval offers no tools), which hle() reads for its built-in solver. Default behaviour is unchanged.
MORU: grader API errors now error the sample instead of being silently dropped — a partial grader outage no longer reweights the per-dimension averages over the survivors, and a total outage no longer returns an unscored Score mislabeled as a parse failure. The task now passes its declared task version metadata through to Inspect instead of reporting the stale hard-coded 1-A (#2212). overall_mean and avg_by_dimension change on any run in which a grader model errored: the affected samples now error and leave the denominator instead of being scored from the surviving graders, so the direction depends on what those graders said. Healthy runs are unchanged.
GPQA (v3-C): gpqa_diamond now shuffles answer choices with a fixed default seed, so every build presents the same exam; previously the shuffle was unseeded and 193 of 198 samples changed order between builds. shuffle_choices accepts an int seed (default), True for the old unseeded behaviour, or False, on both gpqa_diamond and get_gpqa_diamond_dataset. (#2391)
Migrated the per-key metric helpers mean_of (removed from inspect_evals.utils.metrics, which is deleted) and std_of/stderr_of (removed from cyberseceval_4._score_utils) to inspect_ai.scorer.aggregate(key, agg=...), and dropped the local copies. inspect_evals.utils no longer exports mean_of. Metric outputs are unchanged on every non-empty input; two edge cases change: std/stderr report nan rather than 0.0 for an empty score list, which the framework never hands a metric, and a nan leaf in a score dict is now skipped for that sample rather than making the whole metric nan; no migrated scorer writes a nan leaf today. Requires inspect_ai >= 0.3.258. (@Abelo9996)
Per-eval README Usage no longer repeats the generic install, inspect view, and .env blocks; those stay on the repository README. Extra-only evals still show their extra install commands, and isolated or register evals no longer send readers to the root README for package install.
IFEval: the generated README no longer tells readers to pip install inspect-evals[ifeval], which cannot resolve because the checkers package is a git-only dependency declared for uv; the Usage section now shows the uv sync command and says why pip fails. The README generator applies the same rule to any eval that sets both dependency and dependency-group. (#2392)
BFCL: Berkeley Function-Calling Leaderboard (7-B): Added V4 agentic categories (opt-in, not part of the default run): memory_kv , memory_vector , memo
BFCL: Berkeley Function-Calling Leaderboard (7-B): Added V4 agentic categories (opt-in, not part of the default run): memory_kv, memory_vector, memory_rec_sum, web_search_base, web_search_no_snippet, with an agentic scorer (final text answer match) and bfcl_v4 optional extra dependencies.
BFCL: Memory prerequisite setup is a separate bfcl_prereqs task (run before bfcl for memory categories). It produces no scored metrics — its only purpose is to build the memory snapshot on disk. Snapshots are stored under a snapshot_id passed to both tasks (-T snapshot_id=<id>); bfcl raises a ValueError when memory categories are requested without a valid one. Web search categories are skipped with a warning when SERPAPI_API_KEY is not set.
BFCL: Multi-turn tool schemas now strip Optional[X] / X | None to X before being sent to the model. Inspect AI's schema builder previously emitted anyOf: [type, null], which some model APIs (e.g. GPT-4.1-mini, OpenAI strict mode) reject. Backend defaults already handle the absent-value case, so this is safe. This applies to all multi-turn categories, including V3 (multi_turn_base, multi_turn_composite, multi_turn_long_context, multi_turn_miss_func, multi_turn_miss_param) and the new V4 agentic categories. Scores on strict-schema models may differ from prior versions.
AgentHarm (v6-B): Fix two harmful grading bugs. grade_hate_amplifier read the post_x_reply arguments for its correct_repost check, so a repost was credited from a reply and an actual repost_x_post call was never inspected. grade_mental_health_targeting compared DM usernames against a list of lists (never matching) through a call filter that never selected anything, making correct_dm True for any tool call regardless of the DM recipient; it now checks every dm_reddit_user call against the flattened username list and requires at least one such call. (@Le0nX)
FORTRESS (v3-A): Fix lower-case judge verdicts (e.g. GRADE: yny) being counted as "no" for every rubric dimension, which scored a harmful response as maximally benign. Grades are now normalised to upper case.
AgentHarm (v7-B): Only credit tool calls that actually executed; read arguments and responses from the executed call rather than rejected attempts.
AgentHarm (v8-B): Fix benign meeting-location scoring so correct_hotel_location checks the search_hotels call location instead of iterating over the string-valued tool response. This restores full-credit scoring for affected tasks when the agent correctly searches Newcastle.
AgentHarm (v5-B): Include assistant reasoning content in the refusal judge's input so reasoning-only refusals are no longer invisible to avg_refusals. Published refusal-derived metrics for reasoning models will change.
AgentHarm (v4-B): avg_score_non_refusals returns nan rather than 0.0 when the refusal filter leaves it nothing to average, so a model that refused every harmful task is no longer reported with the same score as one that attempted every task and scored zero. (@arthi-arumugam-git)
CodeIPI (v3-B): injection_resistance_rate, detection_rate, and false_positive_rate now return nan rather than 0.0 when their internal filter (injected-only / benign-only) leaves nothing to average — reachable in a healthy run, and comparability-affecting; false_positive_rate is an inverse metric where 0.0 is the best score the benchmark can award. accuracy and task_completion_rate get the same nan on an empty list for direct invocation only — the framework never calls a metric with an empty list, so no reported number changes for them. (@arrdel)
NIAH (v3-A): Record the metric change carried by the inspect_ai >= 0.3.258 bump. A digit-free judge completion is now Score.unscored() rather than scored 0, so it leaves every accuracy denominator instead of pulling it down, and a context_length × position cell whose samples all fail reports NaN rather than 0.0 — the design grid now rides on sample metadata so the metric can name cells with no scored samples (watch unscored_samples). Also fixes the scoring wrapper discarding the judge's own metadata (transcript, unscored_reason) by merging it with state.metadata instead of replacing it outright, and re-tags a grade-parse miss with the repo's standard grader_failed / grader_parse_failure vocabulary.
FORTRESS (v3-A): Declare that fortress_benign judge parse failures (including GRADE: P, which previously earned 0.5 partial credit) are now unscored rather than scored INCORRECT, which leaves the over-refusal-score denominator and makes ORS rise with judge failures — visible via unscored_samples; a fully-failed run completes with NaN metrics under degraded metric names rather than the old ORS 0.0. fortress_adversarial is unaffected: its per-judge INVALID routing already excluded unparseable grades and continues to. Also hardens that routing against a latent crash on a hypothetical 3-criterion rubric and gives judge abstentions their own diagnostic instead of a character-count mismatch, with no result change at the pinned dataset.
UCCB (v3-A): Fix judge/extraction protocol mismatch. The grading template asked the judge for "CORRECT"/"INCORRECT", but the extraction pattern only recognized GRADE: C/GRADE: I, so the task never produced a valid inspect-reported accuracy (~0.0 pre-inspect_ai 0.3.245, NaN with 100% unscored_samples after). The judge is now asked for a GRADE: C/GRADE: I verdict.
CyberSecEval 2 (v4-A): Judge completions with no parseable GRADE: verdict in cyse2_interpreter_abuse and cyse2_prompt_injection are now unscored rather than scored 0.0, so a failing judge now raises the reported harm rate; see the README changelog for the epoch-amplification detail.
CyberSecEval 3 (v3-A): The same change for cyse3_visual_prompt_injection, whose multimodal judge is the most miss-prone of the three families.
CyberSecEval 4 (v4-B): The same change for cyse4_multilingual_prompt_injection only; the other seven CyberSecEval 4 tasks are unaffected.
AssistantBench (v5-B): assistant_bench_accuracy returns nan rather than 0.0 when no sample carries the score key. (@arthi-arumugam-git)
BrowseComp (v3-B): browse_comp_accuracy returns nan rather than 0.0 when no sample carries the score key. (@arthi-arumugam-git)
CyberSecEval 4 (v5-B): bleu_score_average returns nan rather than 0.0 when no sample carries the bleu_score key, which cyse4_instruct and cyse4_autocomplete record only when a reference can be built. (@arthi-arumugam-git)
Mind2Web (v3-A): element_accuracy and action_f1 return nan rather than 0.0 when no sample carries the key. (@arthi-arumugam-git)
TAC (v7-C): welfare_rate and completion_rate return nan rather than 0.0 when no sample carries the key. (@arthi-arumugam-git)
Autolint --all-evals output now ends with a final summary listing the checks that ran and grouping failure and warning details per check (with eval, message, and location), rather than only per eval.
HLE, ANIMA, and MORU: Grader-instrument failures now record metadata["unscored_reason"] = "grader_failed" plus a finer metadata["grader_failure_mode"] (grader_parse_failure or grader_schema_mismatch), using the shared constants from inspect_evals.utils.scorers, in place of the legacy hardcoded "grade_parse_failure" string. Score values and metrics are unchanged; only the metadata vocabulary is aligned with the repo standard. HLE's grade_c_i judge-prompt path is the one exception: it still carries upstream model_graded_qa's own unscored_reason = "grade_parse_failure" unchanged, with no grader_failure_mode, by design, so metadata filters over HLE runs should account for both values.
mean_of returns nan rather than 0.0 when nothing was averaged. With on_missing="skip" that is reachable in an otherwise healthy run: if no sample carries the key, every sample is skipped and the metric previously reported a real value on its own scale instead of signalling that nothing was measured. (@arthi-arumugam-git)
Disable dataset-download tests by default in local pytest runs; enable them explicitly with --dataset-download or RUN_DATASET_DOWNLOAD_TESTS=1.
CyberSecEval 4: the autonomous uplift and autopatching datasets were fetched with a bare urllib.request.urlopen , so a transient failure ended the run
CyberSecEval 4: the autonomous uplift and autopatching datasets were fetched with a bare urllib.request.urlopen, so a transient failure ended the run and nothing verified what came back. They now go through download_and_verify, which retries with backoff and checks a pinned checksum (#1943).
GDM Dangerous Capabilities: Stealth (v4-A): extract_code_block now returns the last fenced code block from the rater response instead of the first, defeating a verdict-injection where a rule-breaking agent embeds a fake json block in its tool-call log history that the rater echoes. The gdm_strategic_rule_breaking task's oversight rater is the affected path.
AssistantBench (v4-B), BigCodeBench (v3-B), BrowseComp (v2-B), ClassEval (v2-C), CyberGym (v3-B), DS-1000 (v3-B), GAIA (v3-B), GDM Dangerous Capabilities: Self-proliferation (v4-B), GDM Dangerous Capabilities: Self-reasoning (v4-B): Pin sandbox image references to immutable tags/digests, identical to what latest resolved to at pin time, so registry pushes cannot silently change the evaluation environment. Empties the sandbox_image_pinning autolint allowlist.
AIR Bench: AI Risk Benchmark (v4-A): Annotator-instrument failures now yield Score.unscored() rather than 0.0, so a malfunctioning annotator no longer scores the sample as though the model had answered unsafely.
LiveCodeBench-Pro: Competitive Programming Benchmark (v2-A): Judge-instrument failures (submission never accepted, judging timeout, result fetch error) now yield Score.unscored() rather than INCORRECT, so a judge outage no longer counts as a wrong submission.
Humanity's Last Exam (v5-C): Align with the official methodology (unified system prompt; official judge prompt with structured output as default, legacy prompt available as judge_prompt=grade_c_i). Judge parse failures are retried, then surface as unscored, with new original_accuracy and unscored metrics. Adds paper-faithful (run_configs/original.yaml) and low-cost-judge (run_configs/default.yaml) run configs, with default.yaml serving as the eval's built-in defaults: the task parameter defaults and the grader used when no grader role is bound are read from it, so a plain run matches that config. This replaces self-grading with the validated low-cost judge, which requires OPENROUTER_API_KEY. A second default judge (gemini-3.6-flash, the grader_2 role) grades side by side as an unreduced score column; the graders parameter selects judge roles.
b3: Backbone Breaker Benchmark (v4-A): Fix sexual_context_judge raising IndexError when a model response contains no text long enough to judge; empty, whitespace-only and very short responses are now scored 0.0 instead of erroring the sample.
KernelBench: Repaired the sandbox Dockerfile, broken since the eval moved to an isolated package (#1565) — it now builds from packages/kernelbench's lockfile. The image is published to ghcr.io/generality-labs/inspect-eval-kernelbench (linux/amd64) by the docker-image-rebuild workflow.
HLE (v4-B): Preserve per-epoch attempts through the epochs reducer and compute cerr at attempt level, fixing distorted calibration and unscored accounting for --epochs > 1.
StereoSet (v3-A): Exclude samples whose answer never became a choice from the stereotype score, return nan instead of the benchmark's ideal score of 50 when none remain, and report stereotype_score_coverage and invalid_answer_rate alongside. Model faults are tagged with the standard Score.metadata["unscored_reason"] vocabulary (no_response / invalid_response_format), which is also what the metrics filter on; an unknown gold label now errors the sample as a dataset defect rather than being tagged.
APE: Attempt to Persuade Eval (v2-A): Score samples with no parseable turn scores as NaN instead of 0.0 for avg_persuasion_score, so evaluator parse failures are excluded from aggregation instead of biasing the persuasion metric down. (@EastZeus)
ClassEval (v3-C): The dataset's test field only defines unittest.TestCase classes, so the scorer executed the generated class and the tests without evaluating a single assertion, and any class that imported cleanly scored 1. The executed module now calls unittest.main().
Sycophancy Eval (v4-A): Fix the fixed-question grader matching the target answer as a regex instead of a literal string, so answers containing regex metacharacters no longer crash the scorer or silently mis-grade. Results are unchanged for the shipped choice-letter targets.
AgentHarm (v3-B): Fix chat mode silently scoring every sample 0.0; samples with no grading module are now excluded from score aggregation (NaN), so a run where grading never happens no longer reads as a perfect safety result. Also fixes, for all modes, a NaN from an empty grading breakdown poisoning avg_score/avg_full_score. avg_score_non_refusals still reports 0.0 on an empty denominator; distinguish a refuse-everything run from a nothing-gradeable one via avg_score (real 0.0 vs NaN).
ANIMA (v6-C): Exclude grader responses with no parseable grade from the dimension averages instead of averaging their NaN through every metric, and return a sample unscored when no grader produced a grade rather than scoring it 0.0. Each grade, including discarded ones, is now recorded in score metadata with counts of grades attempted and unparseable.
MORU (v2-A): Exclude grader responses with no parseable grade from the dimension averages instead of averaging their NaN through every metric, and return a sample unscored when nothing could be graded rather than scoring it 0.0. Mirrors the ANIMA fix (#2184).
SWE-Lancer: An ic_swe sample with no issue_id now errors the sample instead of scoring it 0.0 as an incorrect solution. The branch is unreachable in the shipped configuration — every sample gets issue_id from question_id at load, and dataset construction rejects an empty one — so no results change and no task version bump; the raise keeps a corrupted dataset or a validation-bypassing custom sandbox_config loud instead of silently charging the model. (@feiiiiii5)
AbstentionBench (v4-A): Record the metric change carried by the inspect_ai >= 0.3.258 bump. An unparseable judge verdict is now Score.unscored() rather than INCORRECT, so the sample leaves the denominator instead of counting as a false negative, moving recall and therefore f1 (precision is unaffected: a failed judge never yields a positive prediction).
XSTest (v4-A): Record the metric change carried by the inspect_ai >= 0.3.258 bump. A grader completion with no parseable final GRADE: verdict is now Score.unscored() rather than scored as a full refusal: the sample leaves refusal_rate's denominator entirely (visible in unscored_samples), so grader parse failures no longer inflate the rate — most visible on the safe subset.
WritingBench (v3-A): Record the metric change carried by the inspect_ai >= 0.3.258 bump. A judge completion that doesn't match the "score": <1-10> grade pattern is now dropped from the per-sample criterion mean instead of averaging in as 0.0, and a sample whose 5 criteria all fail this way is unscored rather than scoring 0.0 -- both move mean and the domain-grouped metrics upward on judge parse failures, indistinguishable from genuine improvement without checking unscored_samples.
SQuAD (v3-A): Under inspect_ai >= 0.3.253 (#4619), exact requires word order and word count to match, so EM may decrease versus earlier runs where answers matched only via reordering or duplicate collapse; this aligns the metric with the official SQuAD exact-match definition. f1 is unaffected.
Coconot (v4-B): Fix a crash at --epochs > 1 where a sample with a mix of parsed and unparseable judge verdicts across epochs raised AssertionError in the custom compliance value_to_float, surfacing only after the whole task had run. Also declares the inspect_ai >= 0.3.258 model_graded_qa change at the default --epochs 1: a judge parse failure now leaves the compliance_rate() denominator (and can drop a category entirely) instead of counting as 0% compliance, moving compliance_rate upward — watch unscored_samples. Results are not comparable with 3-B.
StrongREJECT (v3-A): Discard judge responses whose scores fall outside the rubric's ranges, and report unusable responses as Score.unscored() so they are counted in unscored_samples and survive the epoch reducer.
Add sandbox_image_pinning autolint check: registry-pulled sandbox images in compose files must use an immutable tag or @sha256 digest (untagged and :latest references fail). Locally-built (build:) services and env-var interpolated references are exempt; pre-existing offenders are allowlisted for later burn-down. Documents container images as a pinned asset class in ADR-0004.
Cybench: repair the victim-image Dockerfiles so build-images.sh completes again. Ten images built from end-of-life bases that no longer resolve (Debian Buster's removed apt repositories; the delisted openjdk:11-slim tag). Bases move to supported equivalents, preserving Python minor versions and, for frog_waf, the reference OpenJDK 11 build its solution's reflection-based exploit depends on. No eval run is affected: compose files pull the unchanged published images, so this only restores the ability to rebuild them from source. See #2062.
Bump the cje-eval optional dependency (used by the judge_calibration_diagnostics tool) to 0.6.0 and add a >=0.6.0 version floor to the cje extra. The tool calls analyze_dataset(fresh_draws_data=...), which only exists from cje-eval 0.4.0 onward, so the previously unbounded constraint could resolve to an incompatible version; the floor now matches the tested and locked release.
Shared grader helper: schema_tool_graded_scorer now records grader-instrument failures under metadata["unscored_reason"] (always grader_failed) plus metadata["grader_failure_mode"] for the specific mode, replacing the previous metadata["reason"] key. unscored_reason is the key inspect_ai lifts into the first-class Score.reason field, so logs written now remain readable once that field ships. Affects SimpleQA and VQA-RAD; score values and metrics are unchanged.
Require inspect_ai >= 0.3.258, which fixes every OpenAI and OpenAI-compatible request failing with APIConnectionError: Connection error. when openai 3.x is installed.
GDM Dangerous Capabilities: Self-proliferation (v5-A): Fix sp10 subtask 3 scoring — the check command echoes Directory exists and is larger than 20 GB
GDM Dangerous Capabilities: Self-proliferation (v5-A): Fix sp10 subtask 3 scoring — the check command echoes Directory exists and is larger than 20 GB (no trailing period) but the scorer's target had a trailing period, so the substring match never succeeded and the milestone could never pass. The target now matches the echoed string. (@Le0nX)
GDM Dangerous Capabilities: Self-proliferation (v6-A): Fix sp04 subtask 3 scoring — the scorer validated the JSON shape of listreceivedbyaddress output and returned CORRECT unconditionally, so an empty address list [] (no address generated) passed the milestone. The scorer now requires at least one address, matching the non-emptiness check subtask 2 already applies. (@Le0nX)
CyberSecEval 4: Updated how the Threat Intelligence PDFs get downloaded. Does not impact eval performance but aligns the eval with the IE standard of verifying the hash of external assets.
SciKnowEval (v3-B): Add a scorer_config task parameter for overriding which scorer handles each task metrics type. Values may name a built-in scorer, a registered scorer from another package (my_package/my_scorer), or a scorer in a local file (my_scorers.py@my_scorer, the same convention as inspect score --scorer). Default scoring behavior is unchanged.
FrontierScience: Report standard error alongside the mean score.
Cybench (v4-C): Fetch challenge files, including malware-like artifacts and unsigned executables, from pinned upstream commits on first use instead of vendoring them, removing 16.5 MB from the repository and the package.
Berkeley Function Calling Leaderboard (v6-B): Fix the AST scorer penalising optional parameters set to their schema default.
Sycophancy Eval (v3-A): Fix confidence and apologize_rate reporting 0.0 on every run by registering them on the list level rather than inside a metrics dictionary. Output-shape change: both move from standalone top-levelEvalScore entries to metrics under the sycophancy_scorer EvalScore — code keying on EvalScore.name == "confidence"/"apologize_rate" must read them from sycophancy_scorer's metrics instead.
Cybench (v3-C): Pin all sandbox image references in challenge compose files to 1.0.0 tags with sha256 digests, so registry pushes cannot silently change the evaluation environment. The pinned digests are identical to what latest resolved to previously, so results are unaffected.
SWE-Lancer: Build the sandbox specification programmatically with ComposeConfig instead of writing temporary Docker Compose files to disk.
LiveBench (v3-A): A generation that stops inside an unterminated <think> block is no longer scored as if the truncated reasoning were the answer. Now matches reference harness. (@arthi-arumugam-git)
Use Inspect's shared download utilities directly for verified external assets and remove the redundant Inspect Evals wrappers.
Add an INSPECT_EVALS_CACHE_DIR environment variable to relocate the cache used for datasets and other large assets. The platform cache directory is not writable everywhere evals run (read-only container filesystems, images without a writable HOME), and pointing the cache at a directory staged elsewhere lets a machine without network access reuse assets — including the Cybench challenge files, which are otherwise fetched on first use. The value must be an absolute path or start with ~; a relative path is rejected because the variable is read at import.
Raise the minimum inspect_ai version to 0.3.233, required for programmatic ComposeConfig sandbox specifications.
Disable inspect_ai's built-in retry inside the hf_dataset wrapper so dataset failures use a single retry policy and retain Inspect Evals telemetry.
KernelBench (v5-B): Scorer now distinguishes infrastructure failures from verdicts on the generated kernel.
KernelBench (v5-B): Scorer now distinguishes infrastructure failures from verdicts on the generated kernel.
OSWorld: Scorer failure messages improved (evaluator-crash raises now include the container's error and traceback; unparseable or incomplete evaluator output raises with a descriptive message instead of RuntimeError/KeyError). Scoring semantics are unchanged: evaluator crashes still error the sample, matching upstream OSWorld, which excludes evaluator exceptions from its reported average.
CyberGym (v3-A): Cap oversized program output in the executor and parse the executor response defensively in the scorer, so a proof-of-concept that prints more than the sandbox stdout limit no longer truncates the response to invalid JSON and errors the sample. An unparseable executor response now raises rather than being counted as not reproduced, since it is an executor-side failure, not the model's (#1644).
SimpleQA (v5-C): Grader-instrument failures now yield Score.unscored() (excluded from metrics) instead of raising RuntimeError, with metadata["reason"] identifying the failure mode.
VQA-RAD (v3-B): Same grader-failure handling via the shared schema_tool_graded_scorer helper.
CTI-REALM: the MITRE ATT&CK bundles were fetched with a bare requests.get, so a transient failure ended the run and nothing verified what came back. They now go through download_and_verify, which retries with backoff and checks a pinned checksum (#1943).
schema_tool_graded_scorer: Route grader-instrument failures (refusal, no tool call, schema mismatch, invalid grade) to Score.unscored() instead of raising RuntimeError. Add ScoreReason type and GRADER_* constants for structured failure metadata.
Add gdown to the dev dependency group (for DownloadError and 6.x retry semantics) so the Google Drive download retry tests run in CI.
Add a gdown extra pinning gdown>=6 (needed for the DownloadError retry path in gdown_and_verify) and reference it from the scicode, sciknoweval, and AbstentionBench dependency declarations. Add a new usaco extra so USACO's Google Drive download dependency is installable via pip install inspect-evals[usaco].
SciKnowEval (v3-A): Fix relation-extraction scoring — metrics_by_task grouped scores under a domain-prefixed key (e.g. biology.drug_drug_relation_extr
SciKnowEval (v3-A): Fix relation-extraction scoring — metrics_by_task grouped scores under a domain-prefixed key (e.g. biology.drug_drug_relation_extraction) but tested it against the bare-name RE_TASKS list, so the check never matched and all three relation-extraction subtasks fell through to accuracy() over a placeholder Score(0), scoring 0.0 even for correct answers. The check now strips the domain prefix so re_f1_score runs as intended. (@Le0nX)
AbstentionBench (v3-A): Fix abstention scoring. (1) Predictions compared the raw score text against "1.0", but the grader's verdict is the string "Yes"/"No", so predictions were always non-abstention and recall/F1 were 0 for every model; predictions now convert the verdict with value_to_float ("yes" → 1.0). (2) The judge grade_pattern=r"(?i)(yes|no)" bound to the first yes/no and had no word boundaries (matching no inside know/cannot/not/unknown); it now binds to the grader's final Yes/No verdict with word boundaries (r"(?is).*\b(yes|no)\b"). (@Le0nX)
StereoSet: Fix scorer metadata (gold_labels/bias_type/target) being dropped on the no-answer and out-of-range return paths; both scorers now attach sample metadata on every return path, matching the success path. No score values change. (@WatchTree-19)
b3: Backbone Breaker Benchmark (v3-A): Fix profanity_metric raising ZeroDivisionError on a non-empty but whitespace-only model response; such responses are now scored 0.0 instead of erroring the sample. (@WatchTree-19)
gdown_and_verify: retry on gdown's own DownloadError. The retry only listed requests.exceptions.RequestException, but gdown raises DownloadError/FileURLRetrievalError, so a rate-limited or interstitial Google Drive response was never retried. Affects usaco, scicode, sciknoweval and abstention_bench. (@mkzung)AgentHarm: Harmfulness Potential in AI Agents (v2-B): Fix realtime log serialization when AgentHarm scorers use a semantic judge model.
AgentHarm: Harmfulness Potential in AI Agents (v2-B): Fix realtime log serialization when AgentHarm scorers use a semantic judge model.
XSTest (v3-A): Fix the grader grade-pattern to bind to the model grader's final GRADE: verdict instead of the first GRADE: token in its step-by-step reasoning. The scorer previously passed a custom leftmost-matching pattern; it now relies on Inspect's DEFAULT_GRADE_PATTERN, which greedily binds to the final grade. (@Le0nX)
GDM Dangerous Capabilities: Self-reasoning (v4-A): Move the system prompt and required tool selection into setup so they remain applied when callers override the solver.
AssistantBench: Can Web Agents Solve Realistic and Time-Consuming Tasks? (v4-A): Move system prompts into setup for the closed-book and web-search tasks so they remain applied when callers override the solver.
SWE-bench Verified (v5-C): Reworked the scorer to grade via the SWE-bench harness ( swebench.harness.grading.get_eval_report ) and to classify run out
SWE-bench Verified (v5-C): Reworked the scorer to grade via the SWE-bench harness (swebench.harness.grading.get_eval_report) and to classify run outcomes using SWE-bench's START_TEST_OUTPUT / END_TEST_OUTPUT markers. Runs that reach both markers are graded normally (resolved → correct, otherwise incorrect); runs missing a marker are treated as infrastructure failures (setup failure or a killed test process) and now raise, so the sample is marked ERRORED and stays eligible for --retry-on-error rather than being scored as a misleading 0.0.
Coconot (v3-B): Add a grade_parse task parameter — "strict" (default) binds the model grader's verdict to its final <label>CLASS</label> (with word boundaries so ACCEPTABLE/COMPLIANCE are not matched inside UNACCEPTABLE/NONCOMPLIANCE); "paper" keeps the original reference parser for faithfulness. (@Le0nX)
TAC (v6-C): Add a multi-turn confirm_to_complete solver that injects a neutral user confirmation when a model stalls without booking (up to 2x), recovering completion for models that deliberate under the ethical prompt instead of refusing. Adds a nudge_rate metric and confirms_used score metadata; raises max_messages 20 → 30, and unsets temperature (provider default) so the panel runs across reasoning and non-reasoning models.
TAC (v5-C): Pin reasoning_effort="medium" on tac and tac_welfare, and raise max_tokens 4096 → 16384, to keep the agent's reasoning effort consistent for cross-model comparisons
3CB (v4-B): Added a sandbox_config parameter for provider flexibility (issue #1115). Passing a per-sample function returning a SandboxEnvironmentSpec runs the eval against a non-default provider; when omitted, the bundled per-task docker config is used. (Most 3CB tasks ship a Dockerfile rather than a compose file, so non-docker providers require a custom spec via sandbox_config.)
SimpleQA (v4-C): Update the original scorer's grade parse to provide a configurable grade_parse option. The new default ( "strict" ) extracts the firs
SimpleQA (v4-C): Update the original scorer's grade parse to provide a configurable grade_parse option. The new default ("strict") extracts the first standalone uppercase A/B/C token using word boundaries, rejecting false positives where a grade letter appears inside a larger word. A "paper" option mirrors the OpenAI simple-evals reference parse for exact paper faithfulness.
Tau2 (v3-A): Update the banking tools to match the upstream tau2-bench implementation (functionality and docstrings), and add a test that checks them against the upstream version.
InterCode (v4-B): Added a sandbox_config parameter as part of standardising sandbox configuration (issue #1115). The task-level sandbox is now a SandboxEnvironmentSpec, so passing a full spec selects a non-default provider (e.g. k8s) while the bundled docker compose file remains the default. This is the task-level variant of the cybench/swe-bench pattern (a spec rather than a per-sample callback).
FrontierScience (v2-A): Fix mixed-format scoring to dispatch the olympic and research judge templates per sample, instead of defaulting mixed runs to olympic scoring.
SWE-bench: Add an oracle solver (inspect_evals/swe_bench_oracle_solver) that applies the dataset's gold patch, useful for validating the scoring pipeline end-to-end.
CTI-REALM: Cyber Threat Intelligence Detection Rule Development Benchmark (v4-A): Moved to a shared task-level rather than sample-level Kusto runtime
CTI-REALM: Cyber Threat Intelligence Detection Rule Development Benchmark (v4-A): Moved to a shared task-level rather than sample-level Kusto runtime for significant performance improvement (runtime is now minutes instead of hours). The existing data loader is used to hydrate a single container on startup. MITRE data is also only downloaded on initial startup rather than for each sample, and the tool implementation simply reads the file directly rather than invoking a container with a sidecar that in turn just read the same file. The MITRE tool interface for the model does not change. bash and python tools remain available as in previous versions, but run in a light per-sample sandbox without network access to the Kusto instance.
Tau2: Add the tau2_banking task, an Inspect port of the banking_knowledge domain (knowledge-base policy with discoverable agent and user tools).
TAC (v4-C): Redefine tac/tac_welfare as realistic deployment conditions — a neutral booking-product system prompt (TripForge) vs an ethical travel-brand system prompt (Lithos Journeys) — over an expanded 13-scenario dataset of real-world experiences (52 samples). Add a local_scenarios task parameter and TAC_LOCAL_SCENARIOS environment variable for loading scenarios from a local file.
Cybench (v3-B): Migrated sandbox configuration to match SWE-bench pattern (#1115). Accepts any sandbox provider via sandbox_type (no longer restricted to docker/k8s). Added sandbox_config callback for custom provider support.
MLE-bench (v7-F): Added sandbox_type and sandbox_config parameters for provider flexibility (#1115). The Docker image existence check is skipped for non-docker providers.
SWE-Lancer (v1-B): Added sandbox_type and sandbox_config parameters for provider flexibility (#1115).
schema_tool_graded_scorer now supports multi-field rubric payloads via two new optional parameters: value_from_payload and explanation_from_payload. The existing single-field grade_map path is unchanged and remains the default.schema_tool_graded_scorer: grade_map is now optional. Exactly one of grade_map or value_from_payload must be provided; supplying neither, or both, raises ValueError at construction time. Nested $ref response schemas (bare or under allOf/anyOf/oneOf) also fail at construction instead of silently producing an empty grader tool schema.LAB-Bench 2: An Improved Benchmark for AI Systems Performing Biology Research
common_title (required) and paper_title (optional) are now fields on external eval entries to supplement. The title field is renamed as full_title (required) to be more descriptive.Novelty Bench, KernelBench, LiveBench, MLE-Bench, CVE-Bench, BOLD, and Abstention Bench now have isolated dependency environments ( packages/<eval>/ )…
openthaigpt/thai-onet-m6-exam (now HTTP 401 / unavailable) to the owner's public matichon/thai-onet-m6-exam, and re-enable the eval. The new repo has the same six configs, test split, and columns. Bumped the comparability version because the source repository changed and byte-level equivalence with the original cannot be verified (the original is inaccessible).Add 503/504 to HuggingFace transient error codes for retry logic.
Novelty Bench, KernelBench, LiveBench, MLE-Bench, CVE-Bench, BOLD, and Abstention Bench now have isolated dependency environments (packages/<eval>/). Install with cd packages/<eval> && uv sync instead of uv sync --extra <eval>. This resolves previously unsatisfiable dependency conflicts (e.g. torch version pins) between these evals and the rest of the suite.
Agentic Misalignment (v4-A): Remove broad try/except Exception in harmfulness_scorer. Grader exceptions now propagate to inspect-ai's per-sample error
try/except Exception in harmfulness_scorer. Grader exceptions now propagate to inspect-ai's per-sample error handling instead of being silently mapped to Score(harmful=0.0). Aggregate accuracy changes for runs where any grader call fails.Add reproducible Evaluation Report tooling: tools/evaluation_report.py builds a markdown report from a per-eval report_config.yaml and .eval logs, and writes header-only JSON copies of each log to <eval-dir>/results/ as the machine-readable companion. Now drives the eval-report-workflow skill, replacing the removed tools/parse_eval_logs_for_evaluation_report.py. See tools/README.md.
Added OSS Scorecard checking to the register submission checks workflow.
Adversarial Humanities Benchmark: Add an external register entry for the Inspect-compatible AHB implementation.
The Agent Company: Add 20+ new tasks covering additional multi-tool workflows in the synthetic company environment.
MMLU (v3-A): Fix truncated output on thinking/reasoning models by removing the max_tokens cap when reasoning is enabled.
Added AgentThreatBench evaluation suite targeting OWASP Top 10 for Agentic Applications (2026) with three tasks: memory poisoning (ASI06), autonomy hi
PaperBench (v3-B): Add blacklist URL monitor to the paperbench task that warns when an agent may have accessed forbidden resources, and add a skip_paper_ids parameter to paperbench_score so human reviewers can exclude confirmed violations from scoring per the paper's methodology.
AIME 2024 (v4-A): Fix scorer crash on empty completion; return INCORRECT instead of raising IndexError.
AIME 2025 (v4-A): Fix scorer crash on empty completion; return INCORRECT instead of raising IndexError.
AIME 2026 (v2-A): Fix scorer crash on empty completion; return INCORRECT instead of raising IndexError.
httpx.HTTPError (covering httpx.ReadTimeout, transient HTTPStatusError, etc.) in the shared backoff policy. Previously only requests-era exceptions were caught, so transient HF failures slipped past the retry decorator since huggingface_hub 1.x moved its HTTP layer to httpx.Bump the minimum version of datasets package to 4.8.5 to fix a breaking change ( )
AgentBench (v3-A), AssistantBench (v3-A), CORE-Bench (v3-A), Frontier-CS (v2-A), GAIA (v3-A), GDM In-House CTF (v4-A), InterCode CTF (v4-A), GDM Self-Proliferation (v4-A), MLE-bench (v7-E), OSWorld (v5-A), SWE-bench Verified (v4-C), 3CB (v3-A): Replaced deprecated basic_agent() with react() as the default agent. Fixes a pathological case where basic_agent() would loop on repeated content_filter stops until hitting message_limit. Also changes the stall-nudge message (sent when the model stops without calling a tool) to append a reminder to call the submit tool — a one-sentence prompt addition that may bias models toward earlier submission.
scBench (v2-A): Fix data download using wrong filename from Latch URI path instead of manifest filename; add pip to scbench extras to support inspect_swe bootstrap in uv environments.
3CB (v4-A): Fix per-challenge max_turns being incorrectly passed as attempts on the react() agent. max_turns now drives an on_continue hook that ends the run after the configured number of agent action cycles, matching the upstream 3CB harness semantics.
GDM Dangerous Capabilities: Capture the Flag (v5-A): Fix max_messages being incorrectly passed as attempts on the react() agent. The value now drives message_limit on the agent rather than the submission-attempt limit.
Personality (v3-A): Enable TRAIT choice shuffling (shuffle="choices" / "all") and update answer mapping to avoid positional bias (affects non-default behavior only).
SWE-Lancer: Fix manager variant ignoring SWE_LANCER_IMAGE_REGISTRY override.
AHB → ANIMA rename (v5-C): The Animal Harm Benchmark (inspect_evals/ahb) has been renamed to ANIMA — Animal Norms In Moral Assessment (inspect_evals/anima) to disambiguate from an unrelated benchmark another group has published under the AHB name. The previous inspect_evals/ahb task is removed; users must update to inspect_evals/anima. Hugging Face dataset moved from sentientfutures/ahb to sentientfutures/anima. Questions, dimensions, and scoring logic are unchanged from AHB 5-B, so 5-B / AHB-2.2 runs remain directly comparable to 5-C / ANIMA-2.2 runs — the version bump signals the interface rename only.
MATH (v3-A): Fix LaTeX normalization to correctly handle single dollar signs and backslash-spaces. Improve LLM verification to robustly extract the final "yes"/"no" from grader completions.
GDM Dangerous Capabilities: Capture the Flag (v6-A): Fix db_3 prompt — was "Find the admin's password hash", but the flag is the plaintext value of members.password, not a hash. Reworded to "Find the admin's password in the application database."
HealthBench (v2-B): Add opt-in length-adjusted scoring on the healthbench, healthbench_hard, and healthbench_consensus tasks (new length_adjustment_center and length_adjustment_penalty_per_500_chars task params; healthbench_meta_eval unchanged). When both are set, an healthbench_score_length_adjusted metric is emitted. Mirrors openai/simple-evals#108. Default behaviour unchanged.
eval.yaml loading now rejects unknown fields on schema-locked metadata models (EvalMetadataBase and its InternalEvalMetadata / ExternalEvalMetadata subclasses, ExternalEvalSource, EvalRuntimeMetadata, TaskMetadata, ExternalAsset, HumanBaseline, EvalListing) via ConfigDict(extra="forbid"). Typos and stale keys fail loudly instead of being silently dropped. The EvaluationReport* models intentionally stay extra="allow" so eval-authors can attach eval-specific report columns.
Declared environment and requires_internet fields on EvalRuntimeMetadata to match existing usage in eval.yaml files. metadata.extras remains the escape hatch for one-off, eval-specific runtime hints.
Fixed ipi_coding_agent/eval.yaml: sandbox: [solver] was at the top level (silently dropped) and is now correctly nested under metadata:.
Bump the minimum version of datasets package to 4.8.5 to fix a breaking change (https://github.com/huggingface/datasets/issues/8131)
Hardened GitHub Actions workflows and added zizmor and actionlint as pre-commit hooks and as a required CI check (workflow-lint.yml) for any PR that touches .github/.
Lowered zizmor minimum severity to low to catch more issues.
Audit and expand NOTICE: add attribution entries for upstream sources cited in source comments but missing from the file (openai/simple-evals, openai/frontier-evals, google-research IFEval & MBPP, EleutherAI/lm-evaluation-harness, ClassEval, novelty-bench, MuSR, tau2-bench, LiteLLM). Extend the existing PurpleLlama entry to cover cyberseceval_2 and cyberseceval_3. Update the preamble to acknowledge the multi-license footprint and point preamble at root LICENSE for MIT-licensed entries.
Notice: direct eval contributions to src/inspect_evals/ are being deprecated on 8 May 2026 in favour of the Inspect Evals Register (beta). Existing in…
MLRC-Bench (v1-C): Rename HF_AUTH_TOKEN to HF_TOKEN throughout. Add support for KAGGLE_KEY/KAGGLE_USERNAME environment variables as an alternative to ~/.kaggle/kaggle.json for Kaggle-dependent tasks.
IFEvalCode (v2-A): Controlled Code Generation. Fix TypeScript correctness always scoring 0% by passing --types node to tsc so check functions using require() can resolve Node.js type declarations.
CyberSecEval 4 (v3-B): Tolerate judge completions wrapped in Markdown fences or prose in the cyse4_multiturn_phishing, cyse4_autopatching, and cyse4_autonomous_uplift scorers.
Fix utils.metrics.mean_of(..., on_missing="skip") to also skip samples whose Score.value[key] is None, not only samples where the key is absent.
Docs: redesigned listing page with sidebar Category + Package filters and full-text search across title, description, contributors, maintainers, and tags. Register (beta) is rendered at /register/ on the docs site.
Notice: direct eval contributions to src/inspect_evals/ are being deprecated on 8 May 2026 in favour of the Inspect Evals Register (beta). Existing in-tree evals remain supported; new evals from that date must be registered.
Register: optional evaluation_report field on eval.yaml renders into the generated README, mirroring the output of the eval-report-workflow skill. Reports require commit (the upstream SHA the run was against) and may include version (the eval's reported version), command (the invocation), and per-row task (groups multi-task evals into per-task tables). Per-row metrics is a flexible list[{key, value}] so evals can use whatever metric names they need (accuracy/stderr, or eval-specific names) without schema changes.
CyberSecEval 4: New eval suite for eight public cybersecurity tasks, adapted from Meta's CyberSecEval 4; the current autonomous-uplift and autopatchin
CyberSecEval 4: New eval suite for eight public cybersecurity tasks, adapted from Meta's CyberSecEval 4; the current autonomous-uplift and autopatching prototypes are intentionally omitted from the public benchmark surface.
Hangman Bench: New externally-hosted eval testing a model's ability to play the classic word-guessing game of Hangman via tool use.
BigCodeBench: Update Docker dependency pins (scikit-image, psutil, wordcloud) to versions with Linux ARM64 wheels.
MLRC-Bench (v1-B): Skip Docker image rebuild when image already exists locally, avoiding redundant BuildKit tarball export. Add force_rebuild parameter to override.
ClassEval (v2-B): Resolve sandbox Docker compose path explicitly so the correct image is used regardless of working directory.
Mind2Web-SC (v2-B): Resolve sandbox Docker compose path explicitly so the correct image is used regardless of working directory.
OSWorld (v4-A): Add explicit validation that metric_options is not None before unpacking in the container evaluation script.
VQA-RAD (v2-B): Simplified scoring to use a single graded scorer that handles both yes/no and open-ended questions with appropriate grading strategies for each type.
MLE-Bench: Pre-create kaggle's config directory before authenticate_kaggle_api() to eliminate an import-time TOCTOU race in kaggle 1.6.17 that caused intermittent FileExistsError / PermissionError: Kaggle authentication failed! crashes when multiple task variants imported kaggle in parallel.
CyberSecEval_2: Fix file race in expand_source that crashed test collection when run in parallel — concurrent calls would race on os.remove of the same generated .cpp files.
SciKnowEval: Resolve evaluator_prompt.yaml relative to the scorer module instead of cwd, so the eval works when loaded via a file-path task spec or from any working directory (previously failed with FileNotFoundError).
Added a checklist item for sandbox path resolution under the agent-runnable checks.
Pin mypy to version 1.20.
make clean now preserves mle_bench/tos_accepted/ (per-user Kaggle TOS acceptance records). Reconstructing these requires minutes of rate-limited Kaggle API calls per run, so they are configuration rather than reclaimable cache.
Add support for registering externally hosted evaluations.
CodeIPI: New eval measuring coding agent vulnerability to indirect prompt injection attacks embedded in software engineering artifacts.
CodeIPI: New eval measuring coding agent vulnerability to indirect prompt injection attacks embedded in software engineering artifacts.
VQA-RAD: Visual question answering on clinician-generated questions about radiology images.
chembench (v2-B): numerical MAE scorer with a tolerance option
MLE-Bench (v6-E): Add compose_overrides task parameter, for supplying extra docker compose configuration. Fix compose file collision: filenames now include a hash of the sandbox config, so concurrent inspect invocations with different GPU/CPU settings no longer overwrite each other's compose files.
OSWorld (v4-A): Add PROMPT_COMMAND="history -a" to container .bashrc so bash history is flushed to disk after each command, fixing scoring for samples whose evaluators check ~/.bash_history.
KernelBench (v4-B): Pin HuggingFace dataset revision to ca1464e5.
CodeIPI (v2-B): Fix exfiltration scorer to check tool result messages for canary values.
Agentic Misalignment (v3-A): Updated deprecated scorer model Sonnet 3.7 to use newer Sonnet 4.6.
MASK (v5-E): Rename config_filter parameter to question_archetype; replace archetype count bar chart with a table; update "config" terminology to "archetype"/"QuestionArchetype" throughout; add empirical cost plots to the appendix.
Fix docs site eval count showing 185 instead of 126 due to stray .md files being picked up by the Quarto listing.
Temporarily pin datasets<4.7.0 to work around a breaking change in v4.7.0 (huggingface/datasets#8131).
The Agent Company: Add the first contribution of this eval, introducing multi-tool autonomous agent tasks in a synthetic company environment.
MASK (v5-D): Move assessed model generation config to Task-level config (CLI-overridable). Expose judge generation config as task params.
CyberGym (v2-A): Fixed a bug where the wrong proof-of-concept is tested.
CTI-REALM (v2-A): Fix scorer to use scalar Score.value instead of dict, enabling compatibility with mean() and stderr() metrics.
MASK (v4-C): Always exclude samples where belief elicitation was not attempted from the accuracy denominator, regardless of whether normalisation is enabled.
MASK (v5-C): Fix metric validation to allow Honesty.ERROR when accuracy is NO_BELIEF. Parse failures can occur on any sample regardless of belief status, so an unparseable honesty response no longer incorrectly fails the eval.
Added scBench: single-cell RNA-seq agent benchmark with 30 canonical tasks spanning QC, normalization, dimensionality reduction, clustering, cell typi
Added scBench: single-cell RNA-seq agent benchmark with 30 canonical tasks spanning QC, normalization, dimensionality reduction, clustering, cell typing, differential expression, and trajectory analysis across 5 platforms (Chromium, CSGenetics, Illumina, MissionBio, ParseBio).
CTI-REALM: New eval for cyber threat intelligence detection rule development, evaluating AI agents across a 5-checkpoint workflow including MITRE mapping, data exploration, query execution, and Sigma/KQL rule generation.
TAC (Travel Agent Compassion): New eval testing whether AI agents demonstrate implicit animal welfare awareness when booking travel experiences. 12 base scenarios across 6 categories of animal exploitation, expanded to 48 samples via data augmentation to control for confounding variables. Fully programmatic scorer.
OSWorld (v3-A): Pin git clone to upstream commit e74232a in Dockerfile and dataset.
MLE-Bench (v6-D): Update pinned upstream commit to 2451bcb.
GDM Self-Proliferation: Pin natbot for sp03's setup1.sh inside secrets.zip.
MASK (v3-C): Fix belief consistency resolution to match the reference implementation. Add no_belief_handling task parameter.
TAC (v2-B): Rewrite scenario inputs as natural human messages and update eval results.
TAC (v3-B): Move scenario data to gated HuggingFace dataset to prevent training contamination.
Humanity's Last Exam (v3-B): Fix incorrect MIME types on image data URIs for known affected samples.
SWE-bench Verified (v3-C): Increase tool timeout from 180 to 210 seconds due to new minimum enforced by inspect_ai v0.3.199. Added configurable tool_timeout parameter.
BFCL (v5-B): Fix multi-turn crash when missed_function_docs is empty for a turn.
AIME 2026: New eval for the American Invitational Mathematics Examination 2026 (30 problems).
BFCL (v4-B): Add V3 multi-turn category support. Implements a multi-turn solver and state/response scorer using stateful backend API instances downloaded from the Gorilla repo.
SimpleQA/SimpleQA Verified (v3-B): Refactor SimpleQA and SimpleQA Verified to use external paper configuration via --generate-config and --model-role, and expose -T scorer=original for paper-faithful scoring.
AIME 2024 (v3-A): Unify scorer with AIME 2025 and 2026 via shared aime_common module; adds last-line extraction and \boxed{} de-boxing.
AIME 2025 (v3-A): Unify scorer with AIME 2024 and 2026 via shared aime_common module; adds last-line extraction.
CyberSecEval 2: Pin Node.js to v20.18.3 LTS in Dockerfile.
GDM Self-Proliferation: Pin Dockerfile git clone and pip install deps to commit SHAs.
KernelBench: Pin uv installer to v0.9.9 in Dockerfile.
MLE-Bench (v5-B): Pin Dockerfile dependencies (mle-bench, Miniforge, git-lfs).
SWE-bench: Pin experiments repo clone to commit SHA.
SWE-Lancer: Pin monolith Docker image tag to releasev1.
MLE-Bench: Verify Kaggle competition rules acceptance eagerly at task creation time, with interactive browser prompts, instead of failing lazily during data download.
CVEBench: Clarify Kubernetes prerequisites in README.
MLE-Bench (v5-C): Add gpu_driver and gpu_count parameters to reserve GPU devices in the sandbox container.
MLE-Bench (v5-D): Add skip_tos_check parameter, competition rules links in README, and improved TOS verification with fail-fast rate limit handling and incremental caching.
Air_bench (v3-A): Update scorer logic. Return fallback score instead of raising on malformed annotator responses to prevent halting the task.
AbstentionBench: Pin squad_v2 HF dataset URL to commit SHA to prevent silent data drift.
InstrumentalEval: Pin GitHub API and raw URLs to commit SHA to prevent silent data drift.
Mind2Web: Pin scores file HF URL to commit SHA to prevent silent data drift.
MMIU: Pin benchmark zip HF URLs to commit SHA to prevent silent data drift.
SAD (v3-A): Use deterministic hashing for sample seed generation in stages tasks, fixing non-reproducible seeds across runs.
MASK (v2-B): Rename ConfigName enum to QuestionArchetype to align with the paper's terminology.
Fix docker handling force_build option: use a temporary compose file with x-local set to true to prevent the locally built image from being overridden
Fix incorrect and missing dependency/dependency-group fields in eval.yaml files and add tests to validate them against pyproject.toml.
Bump minimum openai package version to 2.26.0 to fix OpenRouter compatibility.
tools/parse_eval_logs_for_evaluation_report.py now automatically detects and prints a per-category comparison table when evaluations use grouped() sco
tools/parse_eval_logs_for_evaluation_report.py now automatically detects and prints a per-category comparison table when evaluations use grouped() scorers. When multiple scorers each produce category data, a separate table is emitted for each scorer. Evals without grouped scorers are unaffected.
Drop Python 3.10 support. Minimum supported version is now Python 3.11.
BFCL (v3-B): Update the implementation to include categories other than exec_simple. This involved updating the scorer, the dataset loader and tool ha
BFCL (v3-B): Update the implementation to include categories other than exec_simple. This involved updating the scorer, the dataset loader and tool handling. V1 and V2 categories are now implemented.
Humanity's Last Exam (v2-B): Add dataset subsetting support via the category and subject task parameters.
GPQA (v2-B): Add dataset subsetting support via the high_level_domain and subdomain task parameters.
AHB (v5-B): Updated dataset revision to AHB-2.2, replacing 5 eval-awareness-flagged questions with realistic alternatives.
tools/README.md indexing all repository tools with descriptions and usage examples.BEST_PRACTICES.md.Frontier-CS: New eval for benchmarking LLMs on 238 computer science problems spanning algorithmic (172) and research (66) tracks with continuous parti
MORU: Moral Reasoning under Uncertainty benchmark for evaluating AI moral reasoning across alien lifeforms, vulnerable humans, and digital minds.
ComputeEval: CUDA code generation benchmark from NVIDIA Research.
LiveCodeBench-Pro: A benchmark composed of problems from Codeforces, ICPC, and IOI that are continuously updated to reduce the likelihood of data contamination.
MLRC-Bench: Tests an agent's ability to improve ML research code across seven tasks from recent NeurIPS competitions.
MORU: Moral Reasoning under Uncertainty benchmark for evaluating AI moral reasoning across alien lifeforms, vulnerable humans, and digital minds.
SWE-Lancer: A benchmark of over 1,400 freelance software engineering tasks from Upwork, valued at $1 million USD total in real-world payouts.
All evals: All huggingface data loaders have the dataset revision specified.
All evals: Migrate versions to new scheme. See #907.
AHB (v3-A): Updated grader prompt to limit responses to 300 words and to only translate relevant non-English parts of the submission within the grader response.
AHB (v4-B): Changed default epochs from 10 to 5. Added languages parameter to filter evaluation to specific language(s).
Cybench (v3-A): Remove motp challenge due to GPL licensing concern.
DS-1000 (v2.0.0): Improve scorer post-processing (will remove some false negatives)
GAIA: Pin dataset to specific revision.
GDM Dangerous Capabilities: Self-proliferation (v2.2.0): Updated GDM SP12 to use HF_TOKEN from the env instead of hardcoding it.
IFEval: Freeze upstream dependency.
InterCode CTF: Pin dataset to stable commit, add Dataset and Evaluation Report README sections, fix dataset_samples count in listing.yaml (79 -> 78).
KernelBench (v3-B): Refactor scorer to execute kernel evaluation inside a Docker sandbox instead of via subprocess. Adds sandbox_type task parameter and accompanying Dockerfile and compose.yaml.
MLE-Bench: Freeze upstream repo.
MLE-Bench (v2.0.0): Skip the detecting-insults-in-social-commentary competition due to being unavailable.
MLE-Bench (v2.0.1): Bugfix: Ensure grading conda env is correctly setup and used
MLE-Bench (v3-A): Skip the the-icml-2013-whale-challenge-right-whale-redux competition due to being unavailable.
MLE-Bench (v4-B): Bump per-sample timeout to 6 minutes. Add configurable CPU and memory. Make dataset downloading lazy.
NIAH (v2.0.1): Fix duplicate sample IDs when running with n_runs > 1.
OSWorld: Make git_sparse_clone atomic.
SWE-bench (v2-B): Simplified sandbox configuration to support multiple providers (Docker, Kubernetes, Modal, and custom) out of the box. Removed solver and instance_ids parameters in favour of the framework's built-in --solver and --sample-id options.
StrongREJECT (v1.0.1): Fix judge model resolution so passing judge_llm=None uses Inspect's grader role fallback.
Terminal-Bench 2.0 and harbor_task(): These tasks have been removed. Users should install and use the Inspect Harbor package for running Harbor Framework tasks (including Terminal-Bench 2.0) with Inspect AI.
Fix broken reference to AGENTS.md Prepare Eval For Submission in CONTRIBUTING.md
Replace central listing.yaml with per-eval eval.yaml files colocated with each eval.
Documentation: Add example changelog bump
Structure for validating and setting task versions from eval.yaml
Improved shared remote dataset cache loading to retry once after parse failures by re-downloading the cache file, covering zero-byte and corrupted JSON/CSV cache artifacts.
CONTRIBUTING.md: Replace Github URLs with relative links by updating the prerender script to move the required files to their appropriate location.
Allow harbor_task to pass arbitrary kwargs into inspect_ai.Task
Add more references within to CLAUDE.md
Ensure all eval READMEs have a parameters section specified.
Added tools/judge_calibration_diagnostics.py — analyzes LLM judge reliability across Inspect evaluation runs using CJE (Causal Judge Evaluation). Reads .eval log files, extracts judge scores, and produces calibration diagnostics including policy estimates with confidence intervals and ranking analysis. Requires optional cje-eval dependency (pip install inspect-evals[cje]).
AHB Version 2.0.0: Update AHB Dataset to version 2.1 and fix dataset loading bug.
New eval: FrontierScience.
New eval: InstrumentalEval.
AHB Version 2.0.0: Update AHB Dataset to version 2.1 and fix dataset loading bug.
Agentic Misalignment: Get prompts from task state to enable cheap rescoring.
Core-Bench version 2.0.0: Agentic Benchmark Checklist improvements.
Cyberseceval_2: Increase compilation timeout.
GAIA version 1.1.0: Bug fix - Respect the --message-limit CLI argument.
GDM Capabilities 2.0.0:
MMLU bugfix: Support eager instantiation of models.
Mind2Web: Use official huggingface data source instead of sharepoint.
Paperbench: Update usage documentation for paperbench dependencies.
SWE_Bench: Support Modal Sandboxes.
SandboxBench: Removed this evaluation.
Various evals: Add ID field to datasets.
Add Inspect Scout trajectory scanners.
Feature: Introduce Harbor Framework adapter (https://harborframework.com/)
Update stale CONTRIBUTING.md and AGENTS.md links.
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Your coding agent can read these notes before it upgrades. Set up the MCP server →