---
name: phasedrift
description: Use PhaseDrift to understand a project's agent effort and code-health trends, investigate recurring friction, and track a tool, instruction, or workflow improvement with before-and-after evidence. Use when asked how agent work is going, what to improve, or whether a recorded change helped.
---

# PhaseDrift: use evidence to improve agent work

PhaseDrift is a local dashboard connecting **agent effort, code health, delivery results, and agent feedback over time**. Help the user make a concrete decision: what to investigate, what change to try, or whether to keep a change. You operate the product and explain the evidence; do not make the user learn its APIs.

Default to inspecting existing evidence. A request to understand the project does not mean change its instructions, run every analyzer, or start a trial. When asked to improve something, carry out the authorized change and measure it through the workflow below. Use the current project, not a project hardcoded into this skill.

## Connect to the right project

From the target repository:

```sh
phasedrift status
```

Use its returned `repo`, `repoKey`, `url`, and collection status. A browser tab may point at a different project. Verify the tracker with `GET <url>/api/bootstrap` and compare `repoKey`; don't print its CSRF token. Never assume ports 4395/4396 or use a demo as real evidence.

If not connected, use `phasedrift projects` and `phasedrift doctor` to diagnose. For an authorized setup, `phasedrift init --repo <absolute-target>` connects it; `phasedrift open` opens its dashboard. Init writes the project review workflow and queues analysis, so do not run it just to inspect an unconnected repo. If setup needs more detail, read the project's `phasedrift-project-setup` skill when available. If the CLI is absent, report that specific prerequisite rather than inventing an installation command.

A stopped tracker is a collection gap, not proof no work occurred. Don't delete locks/databases or kill unrelated processes. Use the supported project lifecycle commands when restart is authorized. One failed request merits checking status, not endless retries.

## Check capabilities before drawing conclusions

After identity verification, probe the relevant documented endpoints: agent-trends, improvements, and feedback-trend for an agent-work inspection. A 404/not_found is a missing route, not an empty dataset. Matching appVersion values do not guarantee the same installed build. Compare the running process release with the installed launcher when local diagnosis is authorized. Never claim trials exist only in the UI because an API returned 404.

If the tracker is older, explain the specific next step: restart this project on the installed compatible release, preserving its data, then verify the endpoints. `phasedrift doctor` reports readiness (derived turns, checkpoints, trials, check collection) and names the next remedy; `phasedrift upgrade` restarts a managed tracker on the installed release with a backup and a derived-data check. An explicitly read-only inspection must not upgrade or restart it; report the blocker and ready next action. When repair is authorized, use the supported lifecycle, verify project identity again, and confirm data survived. If migrated turn tables are empty but native sessions exist, an authorized forced native reimport can backfill derived turns; a normal unchanged-file import may skip them. Verify counts afterward. Never substitute a full code scan for native activity import.

Zero newly imported sessions with skipped unchanged files is normally an up-to-date incremental import. It does not itself establish missing activity. Inspect last import errors, source freshness and actual series before reporting a collection gap.

## Choose the useful view

| Question | View and evidence |
|---|---|
| Is agent work getting easier or more expensive? | **Agent work** (`#activity`): runtime, tool calls, token components, search-before-edit, time to first edit; filter range/model/grain and inspect contributing turns. |
| What keeps slowing the agent down? | **Agent work**: feedback trend, recurring suggestions and linked checkpoint examples. |
| What code changed or got harder to maintain? | **Overview** (`#overview`): selected code metrics over time; **Code health** (`#findings`): current findings; **Code explorer** (`#codebase`): affected files/functions. Verify route availability in the installed version. |
| Did the work pass checks or reach a PR? | **Agent work → Delivery**: imported PR/check/review evidence and source links. |
| Did our tool/instruction change help? | **Agent work → Improvements → Improvement loop**: change, matching operation/task evidence and Keep/Adjust/Stop. Legacy turn comparisons remain under All tracked changes. |

For an ordinary inspection, load only the relevant series and a few supporting examples. Start with two useful metrics, not every overlay. Use the browser if available, or the read-only API reference below. Give the user direct links to the relevant project/view.

## Turn observations into a decision

1. **Inspect freshness and coverage.** Identify project, range, model and primary/all-agent scope. Read source status: native activity, code scans, GitHub evidence and analyzed checkpoints refresh separately.
2. **Identify a concrete pattern.** Inspect underlying turns/reports/findings. Explain the problem in working terms, such as repeated wrong commands or repeated searches before editing. A high raw count alone is not a problem.
3. **Choose one useful intervention.** Prefer a small change tied to the evidence: improve an instruction, fix a tool setup, narrow a noisy command, or simplify a repeatedly troublesome function. Check for an existing trial first. Don't propose generic cleanup unrelated to observed friction.
4. **When authorized, use the source-linked improvement.** Inspect `phasedrift learning proposals` and `phasedrift learning proposal show <id>`. Prepare one executable plan and set it with `phasedrift learning proposal plan <id> --stdin`; begin with `proposal try <id>`. Use repeatable operations for tooling/startup and eligible task issue observations for workflows. Fix input, recipe, scope, examples and quality before comparing. The user should not fill measurement forms or copy IDs between unrelated objects.
5. **Do the real work.** Verify checkout ownership and run the project's checks. Mark implementation through `phasedrift trial implement <linked-trial-id> --change-ref "actual reference"`; this supplies the timeline dot and boundary. Record independent operation runs with `learning operation run`, or let normal task entry/closeout bind eligible work observations. Keep source native measurements separate from agent assessment. A retrospective original-version replay is labeled retrospective, never an earlier prospective pin.
6. **Review the actual result.** The owner screen names the project and separates changes that helped, changes being tested, review results and suggested changes. Keep requires the selected comparison's minimum examples and passing quality; the decision freezes its evidence. Use `learning proposal review <id> --stdin` with `decision` keep/adjust/stop, a concrete `note`, and `summary: {title, before, after, benefit}`. Write the summary for an owner who has not read the code: a short title, one sentence about the old behavior, one about what happens now, and why it matters. Explain the actual change without promising wider savings than the evidence supports. Use `learning proposal summary <id> --stdin` to improve an existing explanation without changing its measurements. A new compatible failure reopens Review. Stop records a decision and does not revert code. Preserve owner acceptance separately. Existing per-turn exploratory trials use the legacy trial commands and pinned definitions under All tracked changes.

For exact trial operations and a worked example, read [references/improvement-workflow.md](references/improvement-workflow.md). For programmatic inspection, read [references/local-api.md](references/local-api.md). Do not load both unless needed.

## Interpret the measurements correctly

- A session is a conversation; a turn is a run within it. Conversation span includes gaps. Runtime per turn is not CPU time or guaranteed human-wait-free time.
- Search/read calls before first edit and time to first edit are navigation signals. They do not prove the right file was found or a fix was finished.
- Cached input is part of input; reasoning output can be part of output. Do not add all token components together or claim dollar savings without supported pricing.
- Missing/partial data is not zero. Check comparable samples and exclusions. Primary and all-agent cohorts can differ drastically; don't compare them as if identical.
- Lower effort is useful only alongside adequate outcomes. Check verification, review findings, completed requirements or owner feedback. A shorter failed task isn't an improvement.
- Code scans describe a revision/worktree. Historical rescans are measurements made now, not proof checks ran at the original date. Full current scan does not fill all historical diagnostics.
- Agent reports are reported observations; model classifications are interpretations; native records/checks are separate evidence. A merged PR isn't automatically defect-free or owner-accepted.
- Compare frozen definitions and versions. If a metric/scope/window changes, explicitly recompute a new version rather than treating incompatible evidence as a continuous experiment.

## Checkpoints and collection

Use the installed `phasedrift-session-review` skill. Begin a meaningful objective once:

```sh
phasedrift session start --task "Short objective" --area "Affected area"
```

The returned handle owns this objective, its delivered lesson/proposal versions and its eligible changes. A known host identity resumes its own task automatically. Without one, use the returned review command carrying `--attempt`. Do not share handles across agents. The supported host hooks deliver instructions and attach exact session identities; skills alone do not prove hook execution.

Close with the actual concise outcome, verification, friction and learning. Include exact delivered `contextUses` dispositions, and `observations` for eligible issues with present/absent/unknown and quality passed/failed/unknown. Unknown is outside the success denominator. Use `suggestions: []` unless a concrete repeated obstacle warrants a source-linked proposal. Optional suggestion `category` and `issueKey` help grouping; they are not an impact score. Report once at meaningful completion, pause or handoff, not after each call.

Offline closeout returns one retry command for the frozen envelope. It never rebinds the claim to later conversation text. Missing native identity does not prevent bounded task reporting; a report still creates no native runtime or tokens. Do not print secrets or raw private logs. Legacy registration/conversation reviews remain supported for older integrations.

Refresh **activity** to import native work; **scan code** to measure code; **refresh GitHub** for delivery; **analyze feedback** for saved report interpretation. Local import/scans use CPU/disk, not new model calls. Writing reports uses the current agent's tokens; optional Jev analysis may use a paid provider. Never silently enable paid bulk analysis.

For Jev, inspect `phasedrift learning status` and `phasedrift learning analysis policy` first. The server loads `TYPESAFE_API_KEY` from the connected project's root `.env.local` on restart; key presence and enabled policy are separate. Never print credentials. Follow [the Jev setup guide](references/jev-analysis.md) for separately authorized activation and request limits. The queue can process existing reports as well as new ones. Classifications interpret reports; they do not prove a fix worked.

## Finish with an answer the user can use

Lead with the practical payoff: what repeated problem this evidence helps prevent, what one change would address it, and how we would tell whether that change helped. An inspection finds a candidate; it does not itself implement an improvement or prove savings.

If the user replies “??”, “how does this help?”, or similar, explain the previous result in plain language and name the concrete next step. Do not respond with “unclear what you are asking” or a generic menu. For example: “Your reports repeatedly say mock tests passed but the real flow broke. Try one representative real-flow check for relevant integration changes, then compare recurrence and rework on similar tasks. The change has not been applied yet.” Keep verification proportional to the task; do not impose live provider checks on every edit.

State the project and evidence window, the useful finding, the supporting examples, and the next action or trial status. Include a link. Explain a material coverage gap once with its remedy. If evidence is insufficient, say exactly what needs to be recorded; do not manufacture a recommendation just to fill the page.

For a first real-project trial of this skill, read [references/first-use.md](references/first-use.md). Success means the agent independently finds the correct project, explains one real pattern and proposes an executable next step. Reciting these instructions is insufficient.

Check-recovery trials are prospective: receipts and explicit retry sequences can be inspected, but without a pre-change sequence baseline they remain Collecting evidence. Do not claim reduced rework or comparative success from a passing check or deliberate guard rejection. Each registration creates a new identity; identical labels never implicitly resume another agent.
