How improvement compounds
Make a useful change part of the next agent's starting point.
Your codebase is part of the harness. A useful improvement changes the environment agents work in, then earns its place through real work.
From a report to a better default
- A task report names a repeated obstacle and a concrete suggestion.
- The agent prepares one change with a measurement and quality condition.
- You authorize it. The agent implements it and records matching evidence.
- Review the result. Keep the useful implementation and activate scoped guidance where appropriate.
- Later relevant tasks retrieve the guidance and report whether the problem recurs.
Make a decision from the whole result
Read the sources. Compare relevant work and actual checks. Choose what to keep, what to undo, or what needs another task before you can decide.
For a source-linked improvement, Keep saves the reviewed evidence, Adjust continues the experiment, and Stop ends tracking. None of these actions reverts your code. Older direct effort trials close on review, including a more-evidence decision; use capture to keep those trials open.
Walk through your first trial →
Improve more than instructions
| What you change | A useful outcome |
|---|---|
| Code structure | Relevant changes are easier to implement and verify |
| Tools and checks | Fewer setup errors or misleading failures |
| Skills and instructions | The right guidance reaches the right task |
| Agent configuration | Comparable work completes reliably with less effort |
Choose a measure that matches the problem. The quality bar still holds. For a controlled comparison, use repeatable cases in Benchmarks.
Know what is happening automatically
Collection, agent reporting, and your decisions
| Source or action | What it contributes |
|---|---|
| Code scans | Repository measurements and analyzer coverage. |
| Supported native logs | Runtime, tool activity, and tokens where recorded. |
| Agent reports | The agent's assessment, checks, friction, and suggestions. |
| Your review | Decisions about a change and separate lesson activation. |
Project skills ask the agent to retrieve context and report its work. Verify saved reports and recorded uses; instructions alone do not guarantee execution.
Optional machine classification has its own configuration and budget. Reports and manually reviewed proposals work without it. Suggested changes do not silently rewrite your code, install tools, or switch models.
Which work belongs in a comparison?
Compare similar tasks, models, harness settings, and actor scope. Include retries, corrections, and rescue work.
Effort comparisons support time windows, model filters, and primary or all-actor scope. Automatic task-filtered comparisons are not implemented. That limitation applies to older per-turn effort trials. Source-linked task-outcome plans use explicit eligibility and known issue observations. Inspect the contributing sessions and your trial plan to judge relevance.