Sync with upstream ClawBench - #2
Merged
Merged
Conversation
run-meta.json recorded what was run — task, model, harness name, image ids — but not the revisions those names resolved to. Two runs of "openclaw on v2" a month apart are indistinguishable in the artifact even when the agent, the corpus, and ClawBench itself have all moved, which makes a published leaderboard row labelled rather than reproducible. This collects the ClawBench version, commit, branch and dirty state; the corpus suite and the revision of the commit that last touched it; and the agent and plugin versions pinned by the harness Dockerfile — reading the pins the image was built from rather than asking a running container. Every lookup is best-effort and returns None rather than raising: a PyPI install has no git repository and a container host may have no git at all, and a missing provenance field must never fail a run that otherwise succeeded. `dirty: None` deliberately means "the lookup failed", which is not the same claim as False. Lookups are cached because a batch run builds one block per task and the answers cannot change within a process.
Adds `provenance` alongside `runtime`, reusing the harness image id that _runtime_meta already resolves rather than inspecting the image twice.
Adds a Provenance section to the trace cookbook covering the block's shape, what corpus.revision and harness.pinned_versions actually mean, and why every field can be null — with the filter to apply before comparing runs across code versions. Tests cover pin extraction for each bundled harness, that base-image and COPY lines are not mistaken for agent pins, pip-style pins, corpus resolution for bundled and external case dirs, and that a missing git checkout or git binary yields nulls rather than an exception.
Four defects Perry2004 caught, all real: `dirty` could never be False. `_git()` collapsed empty output into None, so a clean `git status --porcelain` — which succeeds with no output — was indistinguishable from a failed lookup. `_git()` now returns None only when the command could not run, and "" when it ran and said nothing; callers that want a non-empty value ask for it explicitly. Dockerfile pins were reported as fact even under --no-build, where the image can be arbitrarily older than the Dockerfile on disk. Pins are now claimed only when this run built the image, with a `pins_source` field of "dockerfile" or "unverified" saying which. The signal is set by docker_build() rather than read off the --no-build flag, because clawbench-batch builds once and then runs every child with --no-build — keying off the flag would have marked the entire batch path unverified. Pip extras dropped the pin entirely: `litellm[proxy]==1.77.3` matched nothing at all, so the pin vanished silently rather than being recorded. Only `name@version` was recognised, so revision pins — `pkg@github:o/r#sha`, `pkg@git+https://…#ref`, and pip's PEP 508 `pkg @ git+…@ref` — were missed. That is exactly how a preview-stage agent like DeepSeek Harness is pinned, which is the case TIGER-AI-Lab#309 needs. A floating dist-tag (`@next`) is still not collected: it names a moving target, so recording it as a pin would be a false claim.
The clean/dirty test monkeypatched `_git` to return "", bypassing the very `or None` conversion that made `dirty` unable to ever be False — the mock asserted the intended behaviour while the real code did the opposite. It now runs against a real temporary repository, and fails against the old implementation. Adds coverage for pip extras, all three revision-pin spellings, dist-tags being excluded, URL userinfo not being mistaken for a pin, and pins being claimed only for a harness whose image this run built.
Records which pin spellings are collected and why a floating dist-tag is not, and explains when pinned_versions describes the image that actually ran — including that a batch run reports "dockerfile" while a bare clawbench-run --no-build reports "unverified".
…scovery-metadata docs: align public metadata and distinguish historical benchmark scores
…-owned-cache fix(reproduce): preserve user files when cleaning download caches
…enance feat(runner): record code, corpus, and agent revisions in run-meta.json
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
TIGER-AI-Lab/ClawBenchat187cd252bc60af8ac3a2c98a87c9316e49a5ac75Verification