Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
24 changes: 24 additions & 0 deletions CHANGELOG.md
Original file line number Diff line number Diff line change
Expand Up @@ -676,6 +676,30 @@ All notable changes to this project are documented here. The format is based on
positives. A structural test rejects plant keywords that appear verbatim in
the planted file, and per-fixture sample reports check that a correct report
passes and a finding about a neighbouring defect satisfies no plant.
- Reproducible eval runs (#74): every `evals/run_evals.py` invocation writes
a provider-neutral, schema-versioned `run-record.json`
(`evals/run-record.schema.json`, sorted keys) to its work dir. It holds the
skilldeck version and git commit, each fixture's content digest (files,
exec bits and content), each skill's version, canonical digest and rendered
digest (the install stamp's hash), the harness, exact command template,
version and model, and one entry per planned run
(passed, failed, agent failure, timeout, error, or not run) with timing,
exit code, finding count and the scorer's reasons. Raw reports and stderr
stay in separate files it points to (`--include-reports` embeds them).
`--harness claude|codex|custom` presets pair each agent CLI's
non-interactive command with its adapter, and `--model` passes a model
through `{model}`. `--dry-run` validates the fixtures and prints the plan
without running anything; `--max-runs N` (default 50) refuses an oversized
plan before it starts; runs stay sequential. `--replay RECORD` re-runs a
record's configuration and refuses if any fixture, skill or prompt changed
since, or if its command is not a built-in preset's (unless
`--trust-record-command`). A missing agent command, a repo the adapter
can't install into, a scoring error, an interrupt or a runner bug is now
recorded instead of losing the record; a passing run keeps its record and
raw output (only the review repos are deleted). Agents now run with stdin
closed (`/dev/null`), so `codex exec` no longer waits on or ingests the
runner's stdin, and a review repo is built from exactly the fixture files
its digest covers (no `__pycache__` or `.DS_Store`).
- `skilldeck provenance --verify` re-hashes each installed skill's `meta.yaml`
and `skill.md` and exits 1, naming the skill, when one no longer matches its
recorded canonical digest, is missing, or has unexpected files beside it.
Expand Down
143 changes: 128 additions & 15 deletions evals/README.md
Original file line number Diff line number Diff line change
Expand Up @@ -8,28 +8,40 @@ when changing a skill's wording.

## Running

Requires the [Claude Code CLI](https://claude.com/claude-code) (or another
agent CLI, see `--adapter`) and an API key; **it calls a real agent and costs
real money**, which is why it is manual and not part of CI.
Requires an agent CLI — [Claude Code](https://claude.com/claude-code) by
default, or the OpenAI Codex CLI, or any other through `--agent-cmd` — and its
credentials; **it calls a real agent and costs real money**, which is why it is
manual and not part of CI. Check the plan first with `--dry-run`, which invokes
no agent.

```bash
python evals/run_evals.py # all fixtures
python evals/run_evals.py --dry-run # validate fixtures, print the plan
python evals/run_evals.py # all fixtures, Claude Code
python evals/run_evals.py --skill logging # one skill's fixtures
python evals/run_evals.py --skill authentication-review-saml # one fixture
python evals/run_evals.py --repeat 5 # run each fixture 5 times
python evals/run_evals.py --agent-cmd 'claude -p {prompt}' # default
python evals/run_evals.py --adapter codex --agent-cmd 'codex exec {prompt}'
python evals/run_evals.py --keep # keep temp repos + reports
python evals/run_evals.py --repeat 5 --skill logging # pass rate per fixture
python evals/run_evals.py --harness codex # Codex CLI + the codex adapter
python evals/run_evals.py --model sonnet # request a model from the harness
python evals/run_evals.py --agent-cmd 'my-agent --print {prompt}' # any CLI
python evals/run_evals.py --replay path/to/run-record.json # same config again
python evals/run_evals.py --keep # keep the review repos
```

| Option | Meaning |
| --- | --- |
| `--skill NAME` | Run only the fixtures that exercise skill `NAME`, or the one fixture whose directory is `NAME`. |
| `--agent-cmd CMD` | Agent command line; `{prompt}` is replaced by the review prompt. Default `claude -p {prompt}`. |
| `--adapter NAME` | Which skilldeck adapter installs the skill into the temp repo (`claude`, `codex`, `copilot`, `cursor`, `kiro`; default `claude`). The prompt names the installed file's path, so pair it with that agent's `--agent-cmd`. |
| `--harness NAME` | Agent CLI preset: `claude` (default), `codex`, or `custom` (see [Harnesses](#harnesses)). A preset sets the command and the adapter that matches it. |
| `--agent-cmd CMD` | Agent command line, overriding the preset's; `{prompt}` is replaced by the review prompt and `{model}` by `--model`. Without `--harness`, it makes a `custom` harness. |
| `--adapter NAME` | Which skilldeck adapter installs the skill into the temp repo (`claude`, `codex`, `copilot`, `cursor`, `kiro`). Defaults to the harness's (`claude` for `custom`). The prompt names the installed file's path, so it must be the adapter the agent reads. |
| `--model NAME` | Model to request, substituted for `{model}` (the presets pass it as `--model NAME`) and recorded. Without it the harness's own default is used, and not recorded. |
| `--repeat N` | Run each fixture `N` times (fresh repo each time) and print its pass rate — agents are nondeterministic, so one run says little about a borderline fixture. |
| `--timeout S` | Per-run agent timeout in seconds (default 600). |
| `--keep` | Keep the temp repos even when every run passes. |
| `--max-runs N` | Refuse to start if more than `N` runs (fixtures × repeats) are planned (default 50). |
| `--dry-run` | Validate the fixtures and print the planned runs; no agent or version probe is run (only read-only `git ls-files`, to list fixture files). Exits 2 on an invalid fixture or a plan over `--max-runs`. |
| `--replay RECORD` | Re-run a [run record](#run-records)'s exact configuration; see [Replaying](#replaying). |
| `--trust-record-command` | With `--replay`: run the record's command even though it is not a built-in preset's command (see [Replaying](#replaying)). |
| `--include-reports` | Also copy each run's raw stdout and stderr into the run record. |
| `--keep` | Keep the review repos even when every run passes. |

For each fixture (and each repeat) the runner:

Expand All @@ -42,9 +54,108 @@ For each fixture (and each repeat) the runner:
4. scores the agent's **stdout** (see [Scoring](#scoring)).

A run fails outright — without scoring — if the agent exits non-zero or times
out. Failing runs print the agent's stderr and keep their temp directory; each
repo contains the raw `report.txt` (stdout) and `stderr.txt`. The process exits
non-zero if any run failed.
out; failing runs print the agent's stderr. Runs are **sequential** (concurrency
1): one agent at a time, so `--max-runs` bounds the spend and the wall-clock
time together.

Everything lands in a temp work dir, printed at the start:

```
skilldeck-evals-XXXX/
├── run-record.json # the run record (below)
├── artifacts/run-<N>/<fixture>/report.txt # raw stdout, the scored report
├── artifacts/run-<N>/<fixture>/stderr.txt
└── repos/run-<N>/<fixture>/ # the review repos
```

The review repos are deleted when every run passes (unless `--keep`); the
record and the raw output are always kept. The agent gets no stdin (it reads
`/dev/null`), so a CLI that reads a piped stdin neither blocks nor ingests the
runner's. The process exits 0 when every run passed, 1 when any failed, 2 when
the evals could not run (invalid fixture, budget, changed digests or prompt on
replay, agent command not found) and 130 when interrupted. Once runs start, the
record is always written: a run whose repo can't be built (the adapter refuses
to install, git fails) or whose report can't be stored or scored is recorded as
`error` and the rest go on; an interrupt or a runner bug records the runs so
far, the one in flight as `error` and the rest as `not_run`.

## Harnesses

| Harness | Command | Adapter | Version probe |
| --- | --- | --- | --- |
| `claude` | `claude -p {prompt}`; with a model `claude --model {model} -p {prompt}` | `claude` | `claude --version` |
| `codex` | `codex exec {prompt}`; with a model `codex exec --model {model} {prompt}` | `codex` | `codex --version` |
| `custom` | `--agent-cmd`, verbatim | `--adapter` (default `claude`) | none |

Both presets run the agent's documented non-interactive mode with no other
flags: Claude Code's print mode, and `codex exec`, which prints the final
message on stdout (progress goes to stderr, which is not scored) and runs in a
read-only sandbox by default. The Codex preset is best-effort — it has not yet
been exercised in a recorded run. `--agent-cmd` overrides a preset's command
but keeps its name and adapter; add flags there, such as an approval or sandbox
mode. The version probe runs the command's own executable with `--version`,
and only when that executable is the preset's (`claude` or `codex`, by any
path, with or without `.exe`/`.cmd`): through a wrapper such as
`env FOO=1 claude …` or `npx @openai/codex …` it would report the wrapper's
version, so the record's `version` is `null` instead. A harness that reports
token usage or cost would fill the record's `usage` and `cost_usd`; no preset
parses them yet, so both are `null`.

## Run records

Every run writes `run-record.json`, a provider-neutral, schema-versioned record
([`run-record.schema.json`](run-record.schema.json), JSON Schema 2020-12) with
sorted keys and LF newlines, so equal records are equal bytes on every
platform:

- **what ran**: the skilldeck version, the checkout's git commit and whether
it had uncommitted changes (`null` outside a checkout), and the digest of
`run_evals.py` itself (the scorer and the prompt);
- **against what**: per fixture, a digest of its files — in a git checkout
the files git tracks plus untracked ones it doesn't ignore, elsewhere every
file, never `__pycache__`, `*.pyc` or `.DS_Store` — covering each path,
executable bit (from the git index when tracked) and content (newlines
normalised, so a Windows checkout agrees); the review repo is built from
exactly those files. Then the skill's name, version, canonical digest (the
one in `src/skilldeck/_content_manifest.json`) and `rendered_sha256`: the
rendered skill content the adapter installs, excluding the install stamp
(it equals the stamp's `hash=` and `skilldeck catalog`'s
`rendered_sha256` for that adapter); plus the exact prompt;
- **how**: the harness name, the exact command template, the model (if
requested), the harness version (the probe's first line, `null` if it
failed), the adapter, and the repeat, timeout and budget;
- **every planned run**: its status — `passed`, `failed` (scored and
missed), `agent_failed` (non-zero exit), `timed_out`, `error` (the repo
could not be built, or the agent could not start) or `not_run` (the
invocation stopped early) — with start and end timestamps, duration, exit
code, parsed finding count, the scorer's reasons, and the paths of its raw
output, relative to the record;
- a **summary**: planned, attempted, passed, failed and not-run counts.

The record never contains the raw reports or stderr unless you pass
`--include-reports`, so it can be shared as-is; the raw files stay in the work
dir beside it.

## Replaying

`--replay RECORD` re-runs the recorded fixtures with the recorded harness,
command, model, adapter, repeat count and timeout (so it takes no other
configuration options; `--max-runs`, `--dry-run`, `--keep` and
`--include-reports` still apply). Before anything runs it recomputes every
fixture digest, skill digest and review prompt and **refuses** (exit 2) if any
fixture, skill, rendered skill or prompt changed since the record — the agent
would see a different input, so it is a different experiment. A different
harness version or a changed runner (`run_evals.py`, which holds the scorer) is
printed as a note, not refused. The new record's `config.replay_of` holds the
digest of the record it replayed; comparing the two records' runs is the
variance check. `--replay RECORD --dry-run` verifies all of this without
running an agent.

**A record is data, and replaying it runs its command.** A record can be
edited, or come from someone else, so replay accepts only a built-in preset's
own command (`claude` or `codex`, with or without a model) under that preset's
name. A custom harness's command, or a preset name on any other command, is
refused unless you read the command and pass `--trust-record-command`.

## Scoring

Expand Down Expand Up @@ -122,7 +233,9 @@ calls. Each planted fixture also has sample reports there
(`SAMPLE_REPORTS`): a correct report must pass, and a finding about a
different real defect in the same file must satisfy no plant, which catches
keywords that are too narrow to match or generic enough to match the wrong
finding. The scorer itself is unit-tested in `tests/test_eval_scoring.py`.
finding. The scorer itself is unit-tested in `tests/test_eval_scoring.py`, and
the run records, harness presets, budgets, dry runs and replay, with stand-in
agents, in `tests/test_eval_runs.py`.

A skill may have more than one fixture: name the directory for the skill, or
add a `-<variant>` suffix (e.g. `ci-workflow-review-gitlab`) and set the
Expand Down
Loading
Loading