Test the work you actually ship
Compare agent setups on your tasks, with your quality bar.
Choose a better setup for the work your project needs. Start with a small set of your own tasks, then compare the quality of completed attempts.
A useful case includes the original request, starting revision, necessary context and clear acceptance checks. Run the same case on the setups you want to compare, then review the completed work before comparing effort.
Prepare one fair comparison
- Open a source report or improvement plan in Agent work. Choose Save benchmark candidate.
- Add the original brief, full starting commit, context, environment, and acceptance checks. Review the case before marking it ready.
- Export the task packet, run it through your chosen host, then record the completed output and evidence.
Change one thing
To compare models, keep context and allowed tools consistent. To compare a skill or tool, keep the task and agent configuration consistent.
Repeat attempts. Inspect quality before effort, and include reviewer corrections in delivery cost. Use pass, fail, or unknown for each criterion and owner review where judgment matters.
Task packets, results, and exact versions
The task packet contains title, brief, starting revision, environment, and context. Acceptance checks, evaluator notes, source reports, and solutions stay outside the execution packet.
A completed run belongs to an exact case version. Every criterion needs its matching ID, verdict, and evidence. Agent-reported, owner-reviewed, and imported provenance are separate. Unavailable runtime, tokens, and cost stay unrecorded.
Editing the definition creates a new version. Existing results stay with their original version. The benchmark feature does not launch models or run a bakeoff; use the host and billing arrangement you already choose.
Read and export cases with the CLI
Replace <case-id> with the ID of your real case.
phasedrift learning benchmark list
phasedrift learning benchmark show <case-id>
phasedrift learning benchmark export <case-id>
phasedrift learning benchmark runs <case-id>Use runs <case-id> --version <version> to inspect a particular case version. These commands read and export saved evidence. They do not execute the task.
Carry the useful result forward
Use a supported result to choose the setup for a known task type. Review and activate a scoped lesson separately, then inspect its recorded use during later work.