Skip to content

Record every eval run as a schema-versioned, replayable run record - #133

Merged
richardmhope merged 5 commits into
mainfrom
claude/codebase-review-o7y3i9
Sep 24, 2026
Merged

richardmhope merged 5 commits into
mainfrom
claude/codebase-review-o7y3i9

Conversation

@richardmhope

Copy link
Copy Markdown
Collaborator

What & why

Part of #74.

Every evals/run_evals.py run now writes a run record, run-record.json. It isn't tied to one AI provider and has a versioned schema (evals/run-record.schema.json); keys are sorted and line endings are LF. It records:

  • the skilldeck version, the git commit, whether the checkout had uncommitted changes, and a digest of the runner script itself;

  • for each fixture, a digest of its files:

    • git-tracked files plus untracked files git doesn't ignore;
    • junk such as __pycache__ and .DS_Store excluded;
    • executable bits included.

    The review repo is built from exactly these files;

  • each skill's version, canonical digest and rendered digest. The rendered digest is the install stamp's hash=, which is also skilldeck catalog's rendered_sha256 for that adapter; a test checks all three agree;

  • the harness: its exact command template, version and model;

  • one entry per planned run, with its outcome (passed, failed, agent_failed, timed_out, error or not_run), timing, exit code, number of findings and the scorer's reasons.

Raw reports and stderr stay in the work directory, and the record points to them by relative path. --include-reports embeds them in the record instead.

The record is always written once runs start. Failures to build the repo or score a report, including SkillError, are recorded against that run. An interrupt (exit 130) or a runner bug still writes the record, keeping the runs already done.

New options:

  • --harness claude|codex|custom: claude -p and codex exec presets, each paired with its adapter. --model is passed through {model}; a model value starting with - is refused. The version is probed only from the preset's own executable, so a wrapper command records null.
  • Agents run with stdin closed. Without this, codex exec blocks reading non-TTY stdin.
  • --dry-run validates fixtures and prints the plan without running an agent or the version probe. It runs only read-only git ls-files.
  • --max-runs (default 50) refuses an oversized plan before anything starts. Runs stay sequential.
  • --replay RECORD re-runs a record's configuration:
    • it refuses if any fixture, skill, rendered skill or prompt has changed;
    • it runs only a built-in preset's own command, unless --trust-record-command is given. A record is data, and replaying a custom command executes it.

Tests (tests/test_eval_runs.py, 76 tests) use stand-in agents only. A new autouse guard in tests/conftest.py fails any test that sends a bare claude or codex to subprocess.run.

Review. An independent review raised findings, all fixed here:

  • stdin inheritance;
  • lost records after unexpected errors;
  • replay executing arbitrary commands;
  • the wording of rendered_sha256;
  • version probes crediting wrapper programs;
  • recorded prompts not compared on replay;
  • the fixture digest including junk files and ignoring the exec bit;
  • nits: schema_version: true, SHA-256 commit IDs, absolute paths in messages.

Still open for #74, because these need paid runs: variance reports from repeated real runs, a real check of the codex preset, and token/cost capture (for example by parsing Claude's --output-format json).

Tracking: #94

Type of change

  • Bug fix
  • New feature
  • New or updated skill
  • Docs only
  • Refactor / internal

Checklist

  • Ran uv run --extra dev ruff check . && uv run --extra dev ruff format --check . && uv run --extra dev mypy && uv run --extra dev pytest (895 passed, also with -W error::EncodingWarning and --resolution lowest-direct; plus mypy --strict evals/run_evals.py)
  • Added or updated tests
  • Updated docs where relevant (evals/README.md)
  • Added a CHANGELOG.md entry under ## [Unreleased]
  • For a skill change: bumped that skill's version in meta.yaml (no skill content changed)

🤖 Generated with Claude Code

https://claude.ai/code/session_01HtiCGzpikMrkDYBkfQG5CX


Generated by Claude Code

Every evals/run_evals.py invocation now writes a provider-neutral,
schema-versioned run record (run-record.json, sorted keys, LF) to its work
dir, described by the committed evals/run-record.schema.json. It captures
the skilldeck version, checkout commit and dirty state, the runner's own
digest, each fixture's content digest, each skill's version, canonical
digest and installed-file digest, the harness name, exact command template,
version probe and model, and one entry per planned run: status (passed,
failed, agent_failed, timed_out, error, not_run), timestamps, duration,
exit code, finding count, scorer reasons, and work-dir-relative paths to
the raw report and stderr. Raw output is only embedded with
--include-reports; usage and cost are null until a harness reports them.

- --harness claude|codex|custom presets pair each agent CLI's
  non-interactive command (claude -p, codex exec) with its adapter;
  --agent-cmd still overrides, and --model is passed through {model}.
- --dry-run validates fixtures (expected.yaml, layout, installability) and
  prints the plan without running any subprocess; exits 2 on problems.
- --max-runs (default 50) refuses an oversized plan before anything runs;
  execution stays sequential.
- --replay RECORD re-runs a record's configuration and refuses if any
  fixture, skill, or installed skill file digest changed.
- A missing agent command, an unbuildable repo, or an interrupt is recorded
  instead of aborting without a record; the remaining runs are not_run.

Stand-in-agent tests cover the record fields and ordering (validated
against the schema with a minimal in-test validator), failures, timeouts,
missing commands, dry runs making no subprocess calls, budgets, harness
presets, and replay digest mismatches.

Part of #74.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01HtiCGzpikMrkDYBkfQG5CX
Address review findings on the reproducible-eval-runs change (#74):

- Agents, version probes and git run with stdin=DEVNULL, so `codex exec`
  neither blocks on nor ingests the runner's stdin.
- A run whose repo can't be built (e.g. a SkillError from the adapter) or
  whose report can't be stored or scored is recorded as `error` and the
  rest go on; the run loop writes the record in a finally block, so an
  interrupt (exit 130) or a runner bug keeps the runs already paid for.
- --replay only runs a built-in preset's own command under that preset's
  name; a custom or altered command needs --trust-record-command. Model
  names that look like options are rejected.
- rendered_sha256 is the install stamp's hash (rendered content, stamp
  excluded); docs and schema say so, and a test checks it against the
  installed file.
- The version probe runs only the preset's own executable, so wrappers
  like `env ... claude` or `npx @openai/codex` record version null.
- Replay refuses a changed review prompt, and notes a changed runner in
  --dry-run too.
- Fixture digests cover git-tracked plus untracked-not-ignored files (all
  files outside git), skip __pycache__/*.pyc/.DS_Store, and include the
  exec bit (from the index when tracked); review repos are built from
  exactly those files.
- schema_version rejects `true`, commit ids may be SHA-256, error
  problems carry work-dir-relative paths, and the docstring example uses
  --harness codex.
- An autouse test guard fails any test that would run a bare `claude` or
  `codex` from PATH.

Part of #74.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01HtiCGzpikMrkDYBkfQG5CX
Now that main has skilldeck.stamp.content_hash (#131), the run record's
rendered_sha256 comes from it instead of a local copy of the formula, and
a test checks it against both the installed file's stamp and skilldeck
catalog's rendered_sha256 for the claude and codex adapters.

Part of #74.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01HtiCGzpikMrkDYBkfQG5CX
@richardmhope
richardmhope merged commit 50b3442 into main Sep 24, 2026
17 checks passed
@richardmhope
richardmhope deleted the claude/codebase-review-o7y3i9 branch September 24, 2026 09:53
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants