NewYour coding agent can read the release notes before it upgrades.Set up the MCP server →
Go modules · #1464 by repository stars
Last release 2 days ago
05 Oct 2026
Ships fairly regularly
a new release about every 2 weeks
Most releases are documented
notes for 26 of 32 stable releases
Nothing withdrawn
no release was ever pulled
8 months old
50 releases · first in 2026
One column per month.
Note: This release includes the changes previously prepared under 0.32.0, which was not published.
Note: This release includes the changes previously prepared under 0.32.0, which was not published.
.waza.yaml can now configure files.evalFile, files.taskGlob, and files.taskFileSuffix, with the new naming carried through scaffolding, workspace discovery, discovery mode, schemas, and docs while preserving the existing eval.yaml and tasks/*.yaml defaults (#254, closes #232)config.instruction_files and task-level instruction_files now copy files from the active context into task workspaces and append path-labeled contents to the Copilot system message (#248, closes #239)CopilotEngine instead of constructing a Copilot client directly, keeping grader execution aligned with engine configuration and preserving follow-up recovery behavior (#258, closes #54)copilot-cli bundles are updated from 1.0.2 to 1.0.49 across supported platforms, with reproducible pinned bundle generation via COPILOT_CLI_VERSION (#260, closes #244)waza new skill no longer asks for a nonstandard skill type or emits type: frontmatter, and the wizard now rejects early exits that omit required name or description fields (#261, closes #243)waza check eval discovery — Nested skills and separated evals are discovered consistently in multi-skill workspaces (#247, closes #238)SKILL.md body sections as well as frontmatter descriptions (#236, closes #223)github.com/github/copilot-sdk/go to v0.3.0, migrated session event handling to typed payloads, and refreshed transcript, logging, web API, usage collection, suggestion trace, and test coverage for the new API (#255, closes #253)go install guidance and clarified Windows/WSL install behavior (#246, closes #242; #245, closes #241)Note: This release includes the changes previously prepared under 0.32.0, which was not published.
Read more
Assets 8 Loading
There was an error while loading. Please reload this page .
All reactions
Vocabulary renames — Internal types renamed: BenchmarkSpec → EvalSpec, TestRunner → EvalRunner. Not a breaking change for external consumers (types li…
.agent.md) eval support — Discover .agent.md files alongside SKILL.md, parse agent-specific frontmatter (tools, model, handoffs, mcp-servers, agents), auto-inject tool_constraint grader from agent tools: field, complete worked example under examples/custom-agent/, and new "Evaluating Custom Agents" docs guide (#226, closes #225)_output_contains expectations against file contents now work in CI without a real model. Mock response includes task metadata, file paths, and a 1KB content preview per resource (#228, closes #227)waza serve no longer crashes when stdin isn't a terminal — MCP stdio server only starts when term.IsTerminal() is true; piped input or background mode no longer kills the HTTP dashboard (#224)BenchmarkSpec → EvalSpec, TestRunner → EvalRunner. Not a breaking change for external consumers (types live in internal/) (#222).agent.md coverage to quickstart, getting-started, GUIDE, TUTORIAL, examples README; updated mock engine descriptions in INTEGRATION-TESTING and eval-yaml guide (#230)Updated README with missing CLI commands — Added documentation for recently-added CLI commands that were missing from the README
`waza quality` command — LLM-as-Judge skill quality scoring that evaluates skill output quality using a configurable judge model
waza quality command — LLM-as-Judge skill quality scoring that evaluates skill output quality using a configurable judge model (#218)waza check now includes an advisory that flags skills with overly broad scope, helping authors tighten skill definitions (#219)`--keep-workspace` flag — Preserve the temporary workspace after task execution for debugging agent output (#123, #217)
--keep-workspace flag — Preserve the temporary workspace after task execution for debugging agent output (#123, #217)--no-skills flag and disabled_skills config — Disable specific skills during evaluation to isolate behavior (#126, #216)skill_directories — Specify different skill directories for individual tasks in eval YAML (#156, #215)Nothing published for this version
Nothing published for this version
Follow-up prompts in eval YAML — Tasks can now include pre-written follow-up prompts for multi-turn evaluation conversations (#189, #209)
waza models command — List all available models supported by the configured engine (#208)TestCase definitions are now properly rejected (#132, #206)`output_contains_any` expectation — New expectation field that passes when the agent response contains any one of the specified strings
output_contains_any expectation — New expectation field that passes when the agent response contains any one of the specified strings (#203)max_response_time_ms behavior rule — Enforce maximum response time constraints on agent execution (#201)prompt field can now reference an external file path instead of inline text (#157, #200)tool_calls grader — New grader type that validates the specific tool calls an agent makes during execution (#187, #202)Timestamped output directories — run --output-dir now groups result files by timestamp for cleaner organization
run --output-dir now groups result files by timestamp for cleaner organization (#153)--discover finds eval.yaml in nested layout — Skill discovery now correctly locates eval.yaml files in evals/{name}/ directories at the project root (#44)Eval coverage grid generator — New coverage output that visualizes which skills have eval coverage across grader types
waza run now correctly injects SKILL.md content into the evaluation context, loads trigger test fixtures, and passes MCP server configuration to the engine (#191)Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
`waza new task from-prompt` command — Record Copilot sessions into task YAML files for eval creation
waza new task from-prompt command — Record Copilot sessions into task YAML files for eval creation (#110)waza eval new generates eval.yaml scaffolding for skills (#94).waza.yaml (#96)waza tokens compare supports skill-specific threshold configuration (#93)waza init inventory with FileWriter abstraction (#63)waza suggest deadlock — Execute() now applies the request timeout before calling Start(), preventing goroutine deadlock (#43)ResourceFile.Content type — Changed from string to []byte for proper binary file handling (#117)tokens compare in subdirectory — No longer shows all files as "added" when run from a subdirectory (#105)--output-dir ignored — Fixed --output-dir having no effect for single-skill runs (#109)config.schema.json defaults with Go source of truth (#65).github/skills/ directory (#69)config node max_workers to workers for consistency across all config types
.waza.yaml first (#64)@wbreza added to CODEOWNERS (#111)Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Grader showcase examples demonstrating all grader types
waza dev command and compliance scoring (#131)--debug flag) (#130)waza generate --repo - Scan GitHub repos for SKILL.md files
Skill Discovery (#3)
waza generate --repo <org/repo> - Scan GitHub repos for SKILL.md fileswaza generate --scan - Scan local directory for skillswaza generate --all - Generate evals for all discovered skills (CI-friendly)--allGitHub Issue Creation (#3)
--no-issues flag to skip prompts (CI-friendly)New Modules
waza/scanner.py - Skill discovery from GitHub repos and local directorieswaza/issues.py - GitHub issue creation and formattingNothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Renamed project from `skill-eval` to `waza` (技 - Japanese for "technique/skill")
skill-eval to waza (技 - Japanese for "technique/skill")
waza (previously skill-eval)waza (previously skill_eval)wazaIf you were using skill-eval, update your scripts:
# Old
skill-eval run ./eval.yaml
pip install skill-eval
# New
waza run ./eval.yaml
pip install waza
--suggestions-file option to save improvement suggestions to markdown file
--suggestions-file option to save improvement suggestions to markdown filefrom copilot import CopilotClient not copilot_sdk)waza run - Run evaluation suites against skills
CLI Commands
waza run - Run evaluation suites against skillswaza generate - Auto-generate evals from SKILL.md fileswaza init - Initialize new eval suites interactivelywaza report - Generate reports from resultsEval Generation
--assist flag for better tasks/fixturesExecutors
Graders
Features
-v)--log)--context-dir)--suggestions)Documentation
str, int, bool, etc.Your coding agent can read these notes before it upgrades. Set up the MCP server →