Release Summary
The NeMo Gym 0.5.0 release expands the sandbox ecosystem to seven providers, adds four new general-purpose agent harnesses (Codex, KiloCode, RemoteAgent, and Any-SWE) bringing the total to 20, adds 21 new benchmarks and environments, and wires rollout observability end-to-end from the model server boundary through agent transcripts.
Highlights:
- Seven sandbox providers: Docker, Daytona, ECS Fargate, Enroot, and OpenShell join OpenSandbox and Apptainer; large-scale OpenSandbox reliability significantly improved
- Four new agent harnesses: Codex CLI, KiloCode, RemoteAgent, and
anyswe_agent
- Recompute rewards from stored rollouts without re-running inference with
gym eval reverify
- Rollout observability joined end-to-end: model-call capture, agent observations, and a standardized
ng_trajectory schema
- 21 new environments across six domains: Agentic, Knowledge and instruction following, Long context, Science and coding, Translation and multilingual, and Reasoning
First-Time Contributors
We welcomed 22 new contributors to NeMo Gym with this release:
- @mpatel31415 added
gym eval reverify command and --judge-failed-only flag for recovering failed judge rows
- @Glorf added per-rollout model-call capture, Docker and ECS Fargate sandbox providers, rollout observation contract, Claude Code rollout observations, and standardized
ng_trajectory schema
- @nblintao added per-request policy endpoint override for the SWE agents, enabling RL training frameworks to route each episode through a per-episode recording proxy
- @JeffPengCoder brought OSWorld — a stateful desktop GUI benchmark — into the benchmark catalog
- @rystewart-nvidia added the Legal Agent Bench integration, exposing Harvey's 1,749-task LAB benchmark through standard Gym eval commands
- @fallintoplace fixed the long-standing disagreement between runtime aggregate metrics and the persisted
output.jsonl
- @jonathanlli added RULER pretrain evaluation, enabling text-completion scoring for base and midtraining checkpoints
- @thompsonb overhauled WMT24++ and FLORES translation evaluation (55 locales, 219 language pairs, chrF/spBLEU scoring); expanded MMLU-ProX to 29 languages
- @hkumar92 added the PinchBench agentic benchmark
- @pachmu added ToolSandbox, IHEval, RoleMRC, and RAGTruth benchmarks
Thank you to all 57 NeMo Gym contributors this cycle, including 22 first-time contributors!
Command Line Interface
gym eval reverify — re-run only the verifier on stored rollouts; --judge-failed-only recovers rows that failed due to a flaky judge without re-verifying successful rollouts
gym list and gym search extended to cover models, resources-servers, and agents; gym list <type> <name> drills into a single artifact
- External plugins discoverable via
--search-dir and environment variables; -v/--verbose now accepted before any subcommand
Sandboxing
Five new built-in providers: Docker, Daytona, ECS Fargate, Enroot, and OpenShell join OpenSandbox and Apptainer. Provider choice is a one-line config swap — any agent built on nemo_gym.sandbox works with any provider unchanged.
OpenSandbox reliability at scale is significantly improved: keepalive-bounded transport eliminates silent rollout zeroing at concurrency 300–1500; image registry auth supports private container images; sandbox resources are automatically labeled with team, user, and workload identifiers.
See Available Sandbox Providers for the full list.
Configure Agent Harnesses
New harnesses join the existing set (Claude Code, Hermes, mini-SWE-Agent, OpenClaw, Pi, and more):
- Codex and KiloCode integrate
codex exec and kilo run respectively, routing model calls through Gym for per-rollout capture
anyswe_agent runs any Gym harness inside a SWE task container
swe_agents adds OpenCode as a supported agent framework alongside OpenHands, with DeepSWE and DeNovoSWE dataset support and message replay for trajectory branching
- RemoteAgent drives any external service that implements
POST /v1/responses, with Gym owning the tool loop and verification
Configure Models
- vLLM can now drive
/v1/completions for base and pretrain checkpoint evaluation via opt-in use_completions_api
- All Gym model servers now accept
stream: true on /v1/chat/completions via synthesized SSE, unblocking streaming-first clients such as OpenClaw and Codex
- Add
expose_tools_over_mcp: true to any resources server config to serve its tools over MCP with no handler code changes
Rollout Observability
- Per-rollout model-call capture records requests, responses, token usage, and latency at the model server boundary
- Claude Code transcripts populate
ng_agent_observations; agent observations and model-call capture are joined through a standardized ng_trajectory schema
- Judge failures are routed to a
_failures.jsonl sidecar, keeping aggregate metrics over successfully-judged rows only
New Benchmarks and Environments
21 new environments across six domains:
- Agentic: PinchBench (147 real-world tasks), OSWorld (desktop GUI with VM-backed evaluation), Legal Agent Bench (1,749 Harvey LAB tasks), ToolSandbox (Apple multi-turn tool-use), BrowseComp (web research), BioMNIBench DA, Tau3 banking (BM25+grep offline eval path)
- Knowledge and instruction following: SECQUE, FinanceBench, Finance SEC Search, IHEval (instruction hierarchy, rule-based), Litmus-Bench v0.1, RoleMRC (role-play MRC), RAGTruth (hallucination detection)
- Long context: NIAH (retrieval with overlap penalty)
- Science and coding: CVDP Agentic (expanded to support the agentic subset, harness-agnostic)
- Translation and multilingual: WMT24++ (expanded from 5 to 55 locales), FLORES (expanded from 30 to 219 language pairs, chrF/spBLEU scoring), MMLU-ProX (expanded to 29 languages); RULER now supports pretrain text-completion evaluation
- Reasoning: ReasoningGym environments — six agentic variants: Claude Code, Hermes, and four LangGraph-based variants (orchestrator, reflection, parallel thinking, and ReWOO)
See the Available Environments table for the full list.
Deprecation Notices
- WMT24++ and FLORES scores from prior versions are not comparable with this version's chrF/spBLEU output
- Python 3.13.14 is now required (previously 3.12); users running Gym in Python 3.12 environments must upgrade
Bug Fixes
- SciCode realigned to the AA 65-problem test set with per-rollout subtask accuracy reporting
- Fixed silent rollout zeros at high concurrency against OpenSandbox (keepalive-bounded transport)
- Fixed frozen rollouts in long-running benchmarks (TCP keepalive on global aiohttp connector)
- Fixed
gym eval run failing with FileNotFoundError when the output directory did not exist
- Fixed
tool_choice sent to vLLM without tools, causing request rejection
- Fixed Claude Code
max_turns hardcoded to 30; max_turns: null now removes the cap
- Fixed Apptainer sandbox env vars injected into subprocess argv instead of environment
- Fixed MCQA answer parsing for wrapped formats (
$D$, (D), \boxed{\text{Answer: G}})
- Fixed aggregate metrics including non-persisted rollouts, causing disagreement with
output.jsonl
Documentation
- Rewrote the key terminology glossary with a Gym overview, component map, and links to how-to pages
- New page documenting the Anthropic Messages dialect (
POST /v1/messages) and wiring Claude Code through a Gym model server
- Updated NeMo RL v0.7.0 compatibility guidance
- Migrated all remaining docs examples from legacy
ng_run/ng_collect_rollouts to the unified gym CLI
- Documented
gym list and gym search extensions, external plugin discovery, and MCP auto-exposure
- Added
gym eval reverify and multi-reward verification contract documentation
Release Assets
GitHub Release v0.5.0
Changelog Details
- fix: address CLI issues by @marta-sd :: PR: #1829
- fix: handle malformed yaml when loading extra configs by @marta-sd :: PR: #1854
- fix: print full table for
gym list and create cli/utils.py with helpers by @marta-sd :: PR: #1858
- feat: [GDPval-AA v2 Updates 4 / n] - Re-Run Failed Tasks and Judgements Only by @vadam5 :: PR: #1846
- feat: [GDPval-AA v2 Updates 5 / n] - Multi-Judge Panel by @vadam5 :: PR: #1852
- fix: ERR-225d2c82 Fix HotpotQA dataset license value by @ritaneves :: PR: #1841
- fix: add 'all' extra, surface auth errors in quickstart (QS fixes) by @sephmard :: PR: #1840
- fix(tau3): update repo pin by @cmunley1 :: PR: #1875
- [codex] Add sandbox API docs guide by @hemildesai :: PR: #1717
- docs: add v0.4.0 release notes by @cwing-nvidia :: PR: #1879
- fix: Container guidance is inconsistent across v0.3.0 docs by @ffrujeri :: PR: #1827
- chore: update uv.lock by @kajalj22 :: PR: #1876
- docs: add v0.4.0 highlights to README News and trim archive by @cwing-nvidia :: PR: #1886
- fix(security): bump aiohttp >=3.14.1 and Pillow >=12.3.0 (CVE mitigations) by @kajalj22 :: PR: #1885
- docs: fix typos in README environment table source configs by @cwing-nvidia :: PR: #1892
- ci: use NVIDIA inference for Claude review by @chtruong814 :: PR: #1878
- [mini-swe-agent 2] Quickstart fix + gradeable example data & rollouts by @ananthsub :: PR: #1896
- feat: use all available domain info when listing benchmarks by @marta-sd :: PR: #1857
- feat(benchmarks): Add arguments to preparation script; configurable RULER by @prokotg :: PR: #1711
- docs(fern): add v0.4.0 version snapshot for GA release by @kajalj22 :: PR: #1913
- fix(mini_swe_agent_2): don't install agent deps into root venv (openai pin conflict) by @ananthsub :: PR: #1916
- release: bump main to 0.5.0rc0 for next dev cycle by @ananthsub :: PR: #1921
- ci: enable changelog builder in release workflow by @kajalj22 :: PR: #1923
- fix(stirrup): make persisted deliverables group/world-accessible by @agronskiy :: PR: #1907
- Add Tau3 banking BM25+grep evaluation configs by @jkyi-nvidia :: PR: #1918
- fix: allow Claude Code unlimited turns by @elisam0 :: PR: #1927
- feat(gdpval): resume multi-stage ELO from cache by @agronskiy :: PR: #1933
- fix: load all benchmarks
gym list / gym search commands by @marta-sd :: PR: #1902
- fixing longmt_eval bug introduced by #03d9d9b by @jeffwillette :: PR: #1941
- fix(scicode): run sub-steps in the resources server process, not a Ray worker by @laszkiewiczp :: PR: #1937
- feat(sandbox): ECS Fargate sandbox provider by @Glorf :: PR: #1645
- fix: loose accuracy in ifbench scorer by @e-dobrowolska :: PR: #1936
- Longmt thread safe by @thompsonb :: PR: #1961
- feat(sandbox): local Docker sandbox provider (#1695) by @Glorf :: PR: #1906
- feat: finance sec search environment by @cmunley1 :: PR: #1473
- Add complete SWE rollout timeline metrics by @youngeunkwon0405 :: PR: #1825
- fix: tau2 knowledge deps + repoint tau2-bench pins to stable branches by @e-dobrowolska :: PR: #1954
- SWE: Opencode Integration by @sdevare-nv :: PR: #1302
- Fix : opencode log on debug only by @sdevare-nv :: PR: #1982
- fix(mcqa): boxed answer extraction by @fsiino-nvidia :: PR: #1844
- [GDPVal] Align rubric judge max_tokens default with comparison mode by @Kh4L :: PR: #1234
- Add Daytona sandbox provider by @hemildesai :: PR: #1513
- Feature/biomnibench da harbor by @azkalot1 :: PR: #1897
- Chat completion to Response conversion function by @bxyu-nvidia :: PR: #1998
- fix(opencode_agent): configurable work dir by @cmunley1 :: PR: #1795
- fix(opensandbox): dont override request timeout with connection timeout by @cmunley1 :: PR: #1955
- feat (CritPt): improve AA scoring recovery and add key rotation by @martinagvilas :: PR: #1944
- feat: litmus_agent resources server (domain-agnostic answer verifier + sandbox code-exec tool) by @OliviaViessmann :: PR: #1911
- fix(litmus_agent): satisfy example data validation (follow-up to #1911) by @OliviaViessmann :: PR: #2008
- Browsecomp, baselined by @ritugala :: PR: #1848
- feat(ragtruth): add RAGTruth resources server for case-level hallucin… by @pachmu :: PR: #1947
- feat(rolemrc): add RoleMRC resources server for role-play MRC scoring by @pachmu :: PR: #1946
- openai_model: handle NVIDIA-hosted gpt-oss response quirks (hosted-MCP items + reasoning strip) by @OliviaViessmann :: PR: #1910
- fix: Harden metrics update function by @d-molinari :: PR: #2019
- Add PinchBench agentic benchmark integration by @hkumar92 :: PR: #1810
- fix: 5 example rows for new envs by @cmunley1 :: PR: #2025
- ci: run data validation pre-merge for changed servers by @kajalj22 :: PR: #2023
- feat(observability): add per-rollout model-call capture by @Glorf :: PR: #1715
- fix(server_utils): enable TCP keepalive on global aiohttp connector by @martinagvilas :: PR: #1959
- fix(responses_converter): omit tool_choice when request has no tools by @bg51717 :: PR: #1925
- docs: fix training tutorial links by @mehdiataei :: PR: #2012
- Add locale to target_lang_name in wmt24pp prepare.py by @thompsonb :: PR: #1980
- Align the finance agent loop and finance_sec_search resource server with the vals-ai by @ushnish-de :: PR: #2055
- Add output_regex parser path to litmus_agent by @OliviaViessmann :: PR: #2049
- fix: swe agents opencode file lock by @cmunley1 :: PR: #2021
- feat(gdpval) - Update Multistage ELO sampling Algo by @vadam5 :: PR: #1958
- docs(housekeeping): Add PR SLA tracker by @ritaneves :: PR: #2038
- fix(ci): allow SLA tracker to label pull requests by @ritaneves :: PR: #2066
- fix: observability feature in tau2 harness by @e-dobrowolska :: PR: #2051
- feat(scicode): per-rollout subtask_accuracy + across-run std of both headline metrics by @laszkiewiczp :: PR: #2070
- feat: Add Legal Agent Bench resource server by @rystewart-nvidia :: PR: #1976
- Updates for supporting CVDP Agentic subset by @arti4nvj :: PR: #1744
- feat: Migrate agent skills to a shared canonical directory. by @ffrujeri :: PR: #2007
- fix(scicode): align evaluation with AA setup by @jubick1337 :: PR: #2073
- Add long-context NIAH environment to teach retrieval without scanning by @hsiehjackson :: PR: #1977
- feat: Add Legal Agent Bench to the benchmark catalog by @rystewart-nvidia :: PR: #2075
- Fix the cache path for sec.gov URLs by @ushnish-de :: PR: #2097
- refactor: use framework BaseVerifyResponse in example_session_state_mgmt by @ananthsub :: PR: #1869
- fix(openclaw_agent): pin openclaw_version default for reproducibility by @j-nolan :: PR: #2004
- feat: accept opaque string envelopes for routed_experts by @zyzhou5 :: PR: #2089
- MCP auto-exposure: single module + opt-in flag by @adil-a :: PR: #2059
- feat: enable external plugins for all CLI commands by @marta-sd :: PR: #2061
- feat: extend
gym list and gym search commands by @marta-sd :: PR: #2062
- feat: add
gym list <type> <name> for inspecting by @marta-sd :: PR: #2063
- docs: document extensions to
gym list and gym search by @marta-sd :: PR: #2064
- fix(omniscience): preserve answers without reasoning tags by @martinagvilas :: PR: #2105
- feat: Add AnyTerminal sandbox support by @elisam0 :: PR: #1988
- feat: Implement Gym agent server for Codex CLI agent harness. by @ffrujeri :: PR: #2095
- align MMLU-Pro prompt formatting with NeMo Skills by @gchlebus :: PR: #1853
- Align aggregate metrics with persisted rollouts by @fallintoplace :: PR: #1799
- Updating IFEval to keep verifier information in verifier_metadata field by @arti4nvj :: PR: #2102
- fix: improve MCQA answer parsing by @ka00ri :: PR: #1945
- feat: add Litmus-Bench v0.1 benchmark by @danecor :: PR: #2107
- fix(vllm): handle null request metadata by @macandro96 :: PR: #2139
- feat: reasoning gym agentic environments by @cmunley1 :: PR: #1463
- IHEval — instruction-hierarchy benchmark integration by @pachmu :: PR: #2068
- trial tau 2 turncount limit by @jkyi-nvidia :: PR: #1236
- Sbatch scripts for Super 3.5 evals by @bxyu-nvidia :: PR: #1991
- Rfneves/clean up front page by @ritaneves :: PR: #2152
- feat: accept stream:true on /v1/chat/completions via synthesized SSE by @ananthsub :: PR: #2130
- feat: observability feature for gdpval benchmark by @e-dobrowolska :: PR: #2134
- Add workplace-claude demo notebook and config by @arti4nvj :: PR:
Note truncated.