Record every eval run as a schema-versioned, replayable run record - #133
Merged
Merged
Conversation
Every evals/run_evals.py invocation now writes a provider-neutral,
schema-versioned run record (run-record.json, sorted keys, LF) to its work
dir, described by the committed evals/run-record.schema.json. It captures
the skilldeck version, checkout commit and dirty state, the runner's own
digest, each fixture's content digest, each skill's version, canonical
digest and installed-file digest, the harness name, exact command template,
version probe and model, and one entry per planned run: status (passed,
failed, agent_failed, timed_out, error, not_run), timestamps, duration,
exit code, finding count, scorer reasons, and work-dir-relative paths to
the raw report and stderr. Raw output is only embedded with
--include-reports; usage and cost are null until a harness reports them.
- --harness claude|codex|custom presets pair each agent CLI's
non-interactive command (claude -p, codex exec) with its adapter;
--agent-cmd still overrides, and --model is passed through {model}.
- --dry-run validates fixtures (expected.yaml, layout, installability) and
prints the plan without running any subprocess; exits 2 on problems.
- --max-runs (default 50) refuses an oversized plan before anything runs;
execution stays sequential.
- --replay RECORD re-runs a record's configuration and refuses if any
fixture, skill, or installed skill file digest changed.
- A missing agent command, an unbuildable repo, or an interrupt is recorded
instead of aborting without a record; the remaining runs are not_run.
Stand-in-agent tests cover the record fields and ordering (validated
against the schema with a minimal in-test validator), failures, timeouts,
missing commands, dry runs making no subprocess calls, budgets, harness
presets, and replay digest mismatches.
Part of #74.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01HtiCGzpikMrkDYBkfQG5CX
Address review findings on the reproducible-eval-runs change (#74): - Agents, version probes and git run with stdin=DEVNULL, so `codex exec` neither blocks on nor ingests the runner's stdin. - A run whose repo can't be built (e.g. a SkillError from the adapter) or whose report can't be stored or scored is recorded as `error` and the rest go on; the run loop writes the record in a finally block, so an interrupt (exit 130) or a runner bug keeps the runs already paid for. - --replay only runs a built-in preset's own command under that preset's name; a custom or altered command needs --trust-record-command. Model names that look like options are rejected. - rendered_sha256 is the install stamp's hash (rendered content, stamp excluded); docs and schema say so, and a test checks it against the installed file. - The version probe runs only the preset's own executable, so wrappers like `env ... claude` or `npx @openai/codex` record version null. - Replay refuses a changed review prompt, and notes a changed runner in --dry-run too. - Fixture digests cover git-tracked plus untracked-not-ignored files (all files outside git), skip __pycache__/*.pyc/.DS_Store, and include the exec bit (from the index when tracked); review repos are built from exactly those files. - schema_version rejects `true`, commit ids may be SHA-256, error problems carry work-dir-relative paths, and the docstring example uses --harness codex. - An autouse test guard fails any test that would run a bare `claude` or `codex` from PATH. Part of #74. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01HtiCGzpikMrkDYBkfQG5CX
Now that main has skilldeck.stamp.content_hash (#131), the run record's rendered_sha256 comes from it instead of a local copy of the formula, and a test checks it against both the installed file's stamp and skilldeck catalog's rendered_sha256 for the claude and codex adapters. Part of #74. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01HtiCGzpikMrkDYBkfQG5CX
This was referenced Sep 24, 2026
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
What & why
Part of #74.
Every
evals/run_evals.pyrun now writes a run record,run-record.json. It isn't tied to one AI provider and has a versioned schema (evals/run-record.schema.json); keys are sorted and line endings are LF. It records:the skilldeck version, the git commit, whether the checkout had uncommitted changes, and a digest of the runner script itself;
for each fixture, a digest of its files:
__pycache__and.DS_Storeexcluded;The review repo is built from exactly these files;
each skill's version, canonical digest and rendered digest. The rendered digest is the install stamp's
hash=, which is alsoskilldeck catalog'srendered_sha256for that adapter; a test checks all three agree;the harness: its exact command template, version and model;
one entry per planned run, with its outcome (passed, failed, agent_failed, timed_out, error or not_run), timing, exit code, number of findings and the scorer's reasons.
Raw reports and stderr stay in the work directory, and the record points to them by relative path.
--include-reportsembeds them in the record instead.The record is always written once runs start. Failures to build the repo or score a report, including
SkillError, are recorded against that run. An interrupt (exit 130) or a runner bug still writes the record, keeping the runs already done.New options:
--harness claude|codex|custom:claude -pandcodex execpresets, each paired with its adapter.--modelis passed through{model}; a model value starting with-is refused. The version is probed only from the preset's own executable, so a wrapper command records null.codex execblocks reading non-TTY stdin.--dry-runvalidates fixtures and prints the plan without running an agent or the version probe. It runs only read-onlygit ls-files.--max-runs(default 50) refuses an oversized plan before anything starts. Runs stay sequential.--replay RECORDre-runs a record's configuration:--trust-record-commandis given. A record is data, and replaying a custom command executes it.Tests (
tests/test_eval_runs.py, 76 tests) use stand-in agents only. A new autouse guard intests/conftest.pyfails any test that sends a bareclaudeorcodextosubprocess.run.Review. An independent review raised findings, all fixed here:
rendered_sha256;schema_version: true, SHA-256 commit IDs, absolute paths in messages.Still open for #74, because these need paid runs: variance reports from repeated real runs, a real check of the
codexpreset, and token/cost capture (for example by parsing Claude's--output-format json).Tracking: #94
Type of change
Checklist
uv run --extra dev ruff check . && uv run --extra dev ruff format --check . && uv run --extra dev mypy && uv run --extra dev pytest(895 passed, also with-W error::EncodingWarningand--resolution lowest-direct; plusmypy --strict evals/run_evals.py)evals/README.md)CHANGELOG.mdentry under## [Unreleased]versioninmeta.yaml(no skill content changed)🤖 Generated with Claude Code
https://claude.ai/code/session_01HtiCGzpikMrkDYBkfQG5CX
Generated by Claude Code