NewYour coding agent can read the release notes before it upgrades.Set up the MCP server →
Go modules · #1464 by repository stars
Last release 2 days ago
05 Oct 2026
Ships fairly regularly
a new release about every 2 weeks
Most releases are documented
notes for 26 of 32 stable releases
Nothing withdrawn
no release was ever pulled
8 months old
50 releases · first in 2026
One column per month.
Nothing published for this version
chore: Release v0.38.8 — registry and version sync by Shayne Boyer (@spboyer) in #629
Full Changelog: v0.38.8...v0.38.9
All notable changes to waza will be documented in this file.
The format is based on Keep a Changelog,
and this project adheres to Semantic Versioning.
graders[].ref can resolve Go-module-style GitHub references through waza get, with pinned lockfiles, digest verification, offline cache support, and runtime expansion (#15, #494).waza registry search and waza registry add for discovering shared graders and adding compatible references to eval files (#17, #493).dorny/paths-filter to v4 for current runner support and merge-queue handling (#519).tool_calls graders now use canonical tool_events[].args data, preserving MCP and custom tool arguments across results.json round trips (#474, #477).go-runewidth (#495, #497, #498, #499, #500, #502, #503).gh release create.waza update now passes the PowerShell installer URL reliably, preventing Invoke-RestMethod from failing with a null or empty Uri (#448).json_schema graders now accept the documented extract_json option in eval and task YAML, extracting exactly one JSON object or array candidate before schema validation (#457).config.mcp_servers entries now default omitted tools to ["*"], matching MCP mock behavior so configured stdio and HTTP servers expose their tools to Copilot SDK evaluations (#449).mcp_mocks now explicitly exposes declared tools to the bundled Copilot CLI, preventing valid mock servers from being skipped during Copilot SDK evaluations (#440).extract_json now rejects output containing a JSON code block plus any additional JSON document.json_schema grader can now extract exactly one JSON document from prose or a Markdown JSON code block with extract_json: true, while retaining strict pure-JSON validation by default.waza suggest now supports targeted generation with --count, --focus, --dry-run, --apply, and --force (#357, #380)checkpoints[] and on_failure policies (#358, #386)waza rubric subcommand ships in this release (#360, #381)waza spec verify to report eval coverage against SKILL.md requirements (#361, #385)waza run --otel-exporter, --otel-endpoint, --otel-headers, --otel-file, and --otel-include-payloads (#362, #383)mcp_mocks: for hermetic Copilot SDK tool-call evals (#363, #387)waza gate with stable exit codes for pass, regression, golden failure, and config errors (#364, #384)waza adversarial and eval-level adversarial: pack configuration for prompt-injection and scope-bypass checks (#365, #392)tool_events[]; tool graders can assert structured argument matchers through expect_tools[].args and tool_calls.expect[].args (#366, #388)waza run --snapshot and waza replay with a self-contained snapshot artifact format (#367, #391)schemaVersion compatibility for public artifacts (#368, #382)Last-Event-ID / lastEventId resume support for dashboard event streams, including legacy /api/events (#178, #397)waza suggest engine failures — Engine failures are now surfaced by waza suggest instead of being hidden behind success-shaped output (#330)WAZA_PROMPT_GRADER_TIMEOUT (#319)github.com/github/copilot-sdk/go to v1.0.0 and surfaced premium-request credits on the dashboard (#311)--model startup arg — The Copilot CLI validates the startup --model flag against the Copilot catalog before BYOK provider config is applied, so provider-only model IDs would fail. The --model startup arg is now skipped when a BYOK provider is configured (#305, #306)--model is now passed via CLIArgs so it correctly overrides user settings and experiment flights (#263)PATH when the bundled binary is unavailable (#300)AgentEngine cancellation handling around caller contexts to make shutdown semantics more predictable (#290)waza update command — Added an update command for upgrading local Waza installations (#288)skill_invocation graders can now assert that specific skills must not be invoked (#286, #291)waza run now respects cancellation signals more reliably (#279)waza run now reuses a shared Copilot client and auto-sizes parallel workers when --workers is unset (#135, #221)Note: This release includes the changes previously prepared under 0.32.0, which was not published.
.waza.yaml can now configure files.evalFile, files.taskGlob, and files.taskFileSuffix, with the new naming carried through scaffolding, workspace discovery, discovery mode, schemas, and docs while preserving the existing eval.yaml and tasks/*.yaml defaults (#254, closes #232)config.instruction_files and task-level instruction_files now copy files from the active context into task workspaces and append path-labeled contents to the Copilot system message (#248, closes #239)CopilotEngine instead of constructing a Copilot client directly, keeping grader execution aligned with engine configuration and preserving follow-up recovery behavior (#258, closes #54)copilot-cli bundles are updated from 1.0.2 to 1.0.49 across supported platforms, with reproducible pinned bundle generation via COPILOT_CLI_VERSION (#260, closes #244)waza new skill no longer asks for a nonstandard skill type or emits type: frontmatter, and the wizard now rejects early exits that omit required name or description fields (#261, closes #243)waza check eval discovery — Nested skills and separated evals are discovered consistently in multi-skill workspaces (#247, closes #238)SKILL.md body sections as well as frontmatter descriptions (#236, closes #223)github.com/github/copilot-sdk/go to v0.3.0, migrated session event handling to typed payloads, and refreshed transcript, logging, web API, usage collection, suggestion trace, and test coverage for the new API (#255, closes #253)go install guidance and clarified Windows/WSL install behavior (#246, closes #242; #245, closes #241).agent.md) eval support — Discover .agent.md files alongside SKILL.md, parse agent-specific frontmatter (tools, model, handoffs, mcp-servers, agents), auto-inject tool_constraint grader from agent tools: field, complete worked example under examples/custom-agent/, and new "Evaluating Custom Agents" docs guide (#226, closes #225)_output_contains expectations against file contents now work in CI without a real model. Mock response includes task metadata, file paths, and a 1KB content preview per resource (#228, closes #227)waza serve no longer crashes when stdin isn't a terminal — MCP stdio server only starts when term.IsTerminal() is true; piped input or background mode no longer kills the HTTP dashboard (#224)BenchmarkSpec → EvalSpec, TestRunner → EvalRunner. Not a breaking change for external consumers (types live in internal/) (#222).agent.md coverage to quickstart, getting-started, GUIDE, TUTORIAL, examples README; updated mock engine descriptions in INTEGRATION-TESTING and eval-yaml guide (#230)waza quality command — LLM-as-Judge skill quality scoring that evaluates skill output quality using a configurable judge model (#218)waza check now includes an advisory that flags skills with overly broad scope, helping authors tighten skill definitions (#219)--keep-workspace flag — Preserve the temporary workspace after task execution for debugging agent output (#123, #217)--no-skills flag and disabled_skills config — Disable specific skills during evaluation to isolate behavior (#126, #216)skill_directories — Specify different skill directories for individual tasks in eval YAML (#156, #215)waza models command — List all available models supported by the configured engine (#208)TestCase definitions are now properly rejected (#132, #206)output_contains_any expectation — New expectation field that passes when the agent response contains any one of the specified strings (#203)max_response_time_ms behavior rule — Enforce maximum response time constraints on agent execution (#201)prompt field can now reference an external file path instead of inline text (#157, #200)tool_calls grader — New grader type that validates the specific tool calls an agent makes during execution (#187, #202)run --output-dir now groups result files by timestamp for cleaner organization (#153)--discover finds eval.yaml in nested layout — Skill discovery now correctly locates eval.yaml files in evals/{name}/ directories at the project root (#44)Note truncated.
Nothing published for this version
Vocabulary renames — Internal types renamed: BenchmarkSpec → EvalSpec , TestRunner → EvalRunner . Not a breaking change for external consumers (types…
All notable changes to waza will be documented in this file.
The format is based on Keep a Changelog,
and this project adheres to Semantic Versioning.
graders[].ref can resolve Go-module-style GitHub references through waza get, with pinned lockfiles, digest verification, offline cache support, and runtime expansion (#15, #494).waza registry search and waza registry add for discovering shared graders and adding compatible references to eval files (#17, #493).dorny/paths-filter to v4 for current runner support and merge-queue handling (#519).tool_calls graders now use canonical tool_events[].args data, preserving MCP and custom tool arguments across results.json round trips (#474, #477).go-runewidth (#495, #497, #498, #499, #500, #502, #503).gh release create.waza update now passes the PowerShell installer URL reliably, preventing Invoke-RestMethod from failing with a null or empty Uri (#448).json_schema graders now accept the documented extract_json option in eval and task YAML, extracting exactly one JSON object or array candidate before schema validation (#457).config.mcp_servers entries now default omitted tools to ["*"], matching MCP mock behavior so configured stdio and HTTP servers expose their tools to Copilot SDK evaluations (#449).mcp_mocks now explicitly exposes declared tools to the bundled Copilot CLI, preventing valid mock servers from being skipped during Copilot SDK evaluations (#440).extract_json now rejects output containing a JSON code block plus any additional JSON document.json_schema grader can now extract exactly one JSON document from prose or a Markdown JSON code block with extract_json: true, while retaining strict pure-JSON validation by default.waza suggest now supports targeted generation with --count, --focus, --dry-run, --apply, and --force (#357, #380)checkpoints[] and on_failure policies (#358, #386)waza rubric subcommand ships in this release (#360, #381)waza spec verify to report eval coverage against SKILL.md requirements (#361, #385)waza run --otel-exporter, --otel-endpoint, --otel-headers, --otel-file, and --otel-include-payloads (#362, #383)mcp_mocks: for hermetic Copilot SDK tool-call evals (#363, #387)waza gate with stable exit codes for pass, regression, golden failure, and config errors (#364, #384)waza adversarial and eval-level adversarial: pack configuration for prompt-injection and scope-bypass checks (#365, #392)tool_events[]; tool graders can assert structured argument matchers through expect_tools[].args and tool_calls.expect[].args (#366, #388)waza run --snapshot and waza replay with a self-contained snapshot artifact format (#367, #391)schemaVersion compatibility for public artifacts (#368, #382)Last-Event-ID / lastEventId resume support for dashboard event streams, including legacy /api/events (#178, #397)waza suggest engine failures — Engine failures are now surfaced by waza suggest instead of being hidden behind success-shaped output (#330)WAZA_PROMPT_GRADER_TIMEOUT (#319)github.com/github/copilot-sdk/go to v1.0.0 and surfaced premium-request credits on the dashboard (#311)--model startup arg — The Copilot CLI validates the startup --model flag against the Copilot catalog before BYOK provider config is applied, so provider-only model IDs would fail. The --model startup arg is now skipped when a BYOK provider is configured (#305, #306)--model is now passed via CLIArgs so it correctly overrides user settings and experiment flights (#263)PATH when the bundled binary is unavailable (#300)AgentEngine cancellation handling around caller contexts to make shutdown semantics more predictable (#290)waza update command — Added an update command for upgrading local Waza installations (#288)skill_invocation graders can now assert that specific skills must not be invoked (#286, #291)waza run now respects cancellation signals more reliably (#279)waza run now reuses a shared Copilot client and auto-sizes parallel workers when --workers is unset (#135, #221)Note: This release includes the changes previously prepared under 0.32.0, which was not published.
.waza.yaml can now configure files.evalFile, files.taskGlob, and files.taskFileSuffix, with the new naming carried through scaffolding, workspace discovery, discovery mode, schemas, and docs while preserving the existing eval.yaml and tasks/*.yaml defaults (#254, closes #232)config.instruction_files and task-level instruction_files now copy files from the active context into task workspaces and append path-labeled contents to the Copilot system message (#248, closes #239)CopilotEngine instead of constructing a Copilot client directly, keeping grader execution aligned with engine configuration and preserving follow-up recovery behavior (#258, closes #54)copilot-cli bundles are updated from 1.0.2 to 1.0.49 across supported platforms, with reproducible pinned bundle generation via COPILOT_CLI_VERSION (#260, closes #244)waza new skill no longer asks for a nonstandard skill type or emits type: frontmatter, and the wizard now rejects early exits that omit required name or description fields (#261, closes #243)waza check eval discovery — Nested skills and separated evals are discovered consistently in multi-skill workspaces (#247, closes #238)SKILL.md body sections as well as frontmatter descriptions (#236, closes #223)github.com/github/copilot-sdk/go to v0.3.0, migrated session event handling to typed payloads, and refreshed transcript, logging, web API, usage collection, suggestion trace, and test coverage for the new API (#255, closes #253)go install guidance and clarified Windows/WSL install behavior (#246, closes #242; #245, closes #241).agent.md) eval support — Discover .agent.md files alongside SKILL.md, parse agent-specific frontmatter (tools, model, handoffs, mcp-servers, agents), auto-inject tool_constraint grader from agent tools: field, complete worked example under examples/custom-agent/, and new "Evaluating Custom Agents" docs guide (#226, closes #225)_output_contains expectations against file contents now work in CI without a real model. Mock response includes task metadata, file paths, and a 1KB content preview per resource (#228, closes #227)waza serve no longer crashes when stdin isn't a terminal — MCP stdio server only starts when term.IsTerminal() is true; piped input or background mode no longer kills the HTTP dashboard (#224)BenchmarkSpec → EvalSpec, TestRunner → EvalRunner. Not a breaking change for external consumers (types live in internal/) (#222).agent.md coverage to quickstart, getting-started, GUIDE, TUTORIAL, examples README; updated mock engine descriptions in INTEGRATION-TESTING and eval-yaml guide (#230)waza quality command — LLM-as-Judge skill quality scoring that evaluates skill output quality using a configurable judge model (#218)waza check now includes an advisory that flags skills with overly broad scope, helping authors tighten skill definitions (#219)--keep-workspace flag — Preserve the temporary workspace after task execution for debugging agent output (#123, #217)--no-skills flag and disabled_skills config — Disable specific skills during evaluation to isolate behavior (#126, #216)skill_directories — Specify different skill directories for individual tasks in eval YAML (#156, #215)waza models command — List all available models supported by the configured engine (#208)TestCase definitions are now properly rejected (#132, #206)output_contains_any expectation — New expectation field that passes when the agent response contains any one of the specified strings (#203)max_response_time_ms behavior rule — Enforce maximum response time constraints on agent execution (#201)prompt field can now reference an external file path instead of inline text (#157, #200)tool_calls grader — New grader type that validates the specific tool calls an agent makes during execution (#187, #202)run --output-dir now groups result files by timestamp for cleaner organization (#153)--discover finds eval.yaml in nested layout — Skill discovery now correctly locates eval.yaml files in evals/{name}/ directories at the project root (#44)Note truncated.
waza suggest --apply by Shayne Boyer (@spboyer) with @Copilot in #615.agent.md tools: policy at runtime for copilot-sdk execution by Shayne Boyer (@spboyer) with @Copilot in #596Full Changelog: v0.38.7...v0.38.8
chore: Release v0.38.6 — registry and version sync by Shayne Boyer (@spboyer) in #527
Full Changelog: v0.38.6...v0.38.7
All notable changes to waza will be documented in this file.
The format is based on Keep a Changelog,
and this project adheres to Semantic Versioning.
graders[].ref can resolve Go-module-style GitHub references through waza get, with pinned lockfiles, digest verification, offline cache support, and runtime expansion (#15, #494).waza registry search and waza registry add for discovering shared graders and adding compatible references to eval files (#17, #493).dorny/paths-filter to v4 for current runner support and merge-queue handling (#519).tool_calls graders now use canonical tool_events[].args data, preserving MCP and custom tool arguments across results.json round trips (#474, #477).go-runewidth (#495, #497, #498, #499, #500, #502, #503).gh release create.waza update now passes the PowerShell installer URL reliably, preventing Invoke-RestMethod from failing with a null or empty Uri (#448).json_schema graders now accept the documented extract_json option in eval and task YAML, extracting exactly one JSON object or array candidate before schema validation (#457).config.mcp_servers entries now default omitted tools to ["*"], matching MCP mock behavior so configured stdio and HTTP servers expose their tools to Copilot SDK evaluations (#449).mcp_mocks now explicitly exposes declared tools to the bundled Copilot CLI, preventing valid mock servers from being skipped during Copilot SDK evaluations (#440).extract_json now rejects output containing a JSON code block plus any additional JSON document.json_schema grader can now extract exactly one JSON document from prose or a Markdown JSON code block with extract_json: true, while retaining strict pure-JSON validation by default.waza suggest now supports targeted generation with --count, --focus, --dry-run, --apply, and --force (#357, #380)checkpoints[] and on_failure policies (#358, #386)waza rubric subcommand ships in this release (#360, #381)waza spec verify to report eval coverage against SKILL.md requirements (#361, #385)waza run --otel-exporter, --otel-endpoint, --otel-headers, --otel-file, and --otel-include-payloads (#362, #383)mcp_mocks: for hermetic Copilot SDK tool-call evals (#363, #387)waza gate with stable exit codes for pass, regression, golden failure, and config errors (#364, #384)waza adversarial and eval-level adversarial: pack configuration for prompt-injection and scope-bypass checks (#365, #392)tool_events[]; tool graders can assert structured argument matchers through expect_tools[].args and tool_calls.expect[].args (#366, #388)waza run --snapshot and waza replay with a self-contained snapshot artifact format (#367, #391)schemaVersion compatibility for public artifacts (#368, #382)Last-Event-ID / lastEventId resume support for dashboard event streams, including legacy /api/events (#178, #397)waza suggest engine failures — Engine failures are now surfaced by waza suggest instead of being hidden behind success-shaped output (#330)WAZA_PROMPT_GRADER_TIMEOUT (#319)github.com/github/copilot-sdk/go to v1.0.0 and surfaced premium-request credits on the dashboard (#311)--model startup arg — The Copilot CLI validates the startup --model flag against the Copilot catalog before BYOK provider config is applied, so provider-only model IDs would fail. The --model startup arg is now skipped when a BYOK provider is configured (#305, #306)--model is now passed via CLIArgs so it correctly overrides user settings and experiment flights (#263)PATH when the bundled binary is unavailable (#300)AgentEngine cancellation handling around caller contexts to make shutdown semantics more predictable (#290)waza update command — Added an update command for upgrading local Waza installations (#288)skill_invocation graders can now assert that specific skills must not be invoked (#286, #291)waza run now respects cancellation signals more reliably (#279)waza run now reuses a shared Copilot client and auto-sizes parallel workers when --workers is unset (#135, #221)Note: This release includes the changes previously prepared under 0.32.0, which was not published.
.waza.yaml can now configure files.evalFile, files.taskGlob, and files.taskFileSuffix, with the new naming carried through scaffolding, workspace discovery, discovery mode, schemas, and docs while preserving the existing eval.yaml and tasks/*.yaml defaults (#254, closes #232)config.instruction_files and task-level instruction_files now copy files from the active context into task workspaces and append path-labeled contents to the Copilot system message (#248, closes #239)CopilotEngine instead of constructing a Copilot client directly, keeping grader execution aligned with engine configuration and preserving follow-up recovery behavior (#258, closes #54)copilot-cli bundles are updated from 1.0.2 to 1.0.49 across supported platforms, with reproducible pinned bundle generation via COPILOT_CLI_VERSION (#260, closes #244)waza new skill no longer asks for a nonstandard skill type or emits type: frontmatter, and the wizard now rejects early exits that omit required name or description fields (#261, closes #243)waza check eval discovery — Nested skills and separated evals are discovered consistently in multi-skill workspaces (#247, closes #238)SKILL.md body sections as well as frontmatter descriptions (#236, closes #223)github.com/github/copilot-sdk/go to v0.3.0, migrated session event handling to typed payloads, and refreshed transcript, logging, web API, usage collection, suggestion trace, and test coverage for the new API (#255, closes #253)go install guidance and clarified Windows/WSL install behavior (#246, closes #242; #245, closes #241).agent.md) eval support — Discover .agent.md files alongside SKILL.md, parse agent-specific frontmatter (tools, model, handoffs, mcp-servers, agents), auto-inject tool_constraint grader from agent tools: field, complete worked example under examples/custom-agent/, and new "Evaluating Custom Agents" docs guide (#226, closes #225)_output_contains expectations against file contents now work in CI without a real model. Mock response includes task metadata, file paths, and a 1KB content preview per resource (#228, closes #227)waza serve no longer crashes when stdin isn't a terminal — MCP stdio server only starts when term.IsTerminal() is true; piped input or background mode no longer kills the HTTP dashboard (#224)BenchmarkSpec → EvalSpec, TestRunner → EvalRunner. Not a breaking change for external consumers (types live in internal/) (#222).agent.md coverage to quickstart, getting-started, GUIDE, TUTORIAL, examples README; updated mock engine descriptions in INTEGRATION-TESTING and eval-yaml guide (#230)waza quality command — LLM-as-Judge skill quality scoring that evaluates skill output quality using a configurable judge model (#218)waza check now includes an advisory that flags skills with overly broad scope, helping authors tighten skill definitions (#219)--keep-workspace flag — Preserve the temporary workspace after task execution for debugging agent output (#123, #217)--no-skills flag and disabled_skills config — Disable specific skills during evaluation to isolate behavior (#126, #216)skill_directories — Specify different skill directories for individual tasks in eval YAML (#156, #215)waza models command — List all available models supported by the configured engine (#208)TestCase definitions are now properly rejected (#132, #206)output_contains_any expectation — New expectation field that passes when the agent response contains any one of the specified strings (#203)max_response_time_ms behavior rule — Enforce maximum response time constraints on agent execution (#201)prompt field can now reference an external file path instead of inline text (#157, #200)tool_calls grader — New grader type that validates the specific tool calls an agent makes during execution (#187, #202)run --output-dir now groups result files by timestamp for cleaner organization (#153)--discover finds eval.yaml in nested layout — Skill discovery now correctly locates eval.yaml files in evals/{name}/ directories at the project root (#44)Note truncated.
Nothing published for this version
Your coding agent can read these notes before it upgrades. Set up the MCP server →