Skip to content

Sync with upstream ClawBench - #2

Merged
rgarcia merged 12 commits into
mainfrom
hypeship/sync-upstream-187cd25
Oct 1, 2026
Merged

rgarcia merged 12 commits into
mainfrom
hypeship/sync-upstream-187cd25

Conversation

@rgarcia

@rgarcia rgarcia commented Oct 1, 2026

Copy link
Copy Markdown

Summary

  • synchronize the fork with upstream TIGER-AI-Lab/ClawBench at 187cd252bc60af8ac3a2c98a87c9316e49a5ac75
  • include run provenance, reproduction-cache cleanup, and documentation updates

Verification

  • branch points exactly to the upstream commit
  • fork is 12 commits behind upstream before this sync

vaibhavdabas16 and others added 12 commits September 9, 2026 20:45
run-meta.json recorded what was run — task, model, harness name, image ids
— but not the revisions those names resolved to. Two runs of "openclaw on
v2" a month apart are indistinguishable in the artifact even when the
agent, the corpus, and ClawBench itself have all moved, which makes a
published leaderboard row labelled rather than reproducible.

This collects the ClawBench version, commit, branch and dirty state; the
corpus suite and the revision of the commit that last touched it; and the
agent and plugin versions pinned by the harness Dockerfile — reading the
pins the image was built from rather than asking a running container.

Every lookup is best-effort and returns None rather than raising: a PyPI
install has no git repository and a container host may have no git at all,
and a missing provenance field must never fail a run that otherwise
succeeded. `dirty: None` deliberately means "the lookup failed", which is
not the same claim as False. Lookups are cached because a batch run builds
one block per task and the answers cannot change within a process.
Adds `provenance` alongside `runtime`, reusing the harness image id that
_runtime_meta already resolves rather than inspecting the image twice.
Adds a Provenance section to the trace cookbook covering the block's
shape, what corpus.revision and harness.pinned_versions actually mean,
and why every field can be null — with the filter to apply before
comparing runs across code versions.

Tests cover pin extraction for each bundled harness, that base-image and
COPY lines are not mistaken for agent pins, pip-style pins, corpus
resolution for bundled and external case dirs, and that a missing git
checkout or git binary yields nulls rather than an exception.
Four defects Perry2004 caught, all real:

`dirty` could never be False. `_git()` collapsed empty output into None, so
a clean `git status --porcelain` — which succeeds with no output — was
indistinguishable from a failed lookup. `_git()` now returns None only when
the command could not run, and "" when it ran and said nothing; callers
that want a non-empty value ask for it explicitly.

Dockerfile pins were reported as fact even under --no-build, where the
image can be arbitrarily older than the Dockerfile on disk. Pins are now
claimed only when this run built the image, with a `pins_source` field of
"dockerfile" or "unverified" saying which. The signal is set by
docker_build() rather than read off the --no-build flag, because
clawbench-batch builds once and then runs every child with --no-build —
keying off the flag would have marked the entire batch path unverified.

Pip extras dropped the pin entirely: `litellm[proxy]==1.77.3` matched
nothing at all, so the pin vanished silently rather than being recorded.

Only `name@version` was recognised, so revision pins — `pkg@github:o/r#sha`,
`pkg@git+https://…#ref`, and pip's PEP 508 `pkg @ git+…@ref` — were missed.
That is exactly how a preview-stage agent like DeepSeek Harness is pinned,
which is the case TIGER-AI-Lab#309 needs. A floating dist-tag (`@next`) is still not
collected: it names a moving target, so recording it as a pin would be a
false claim.
The clean/dirty test monkeypatched `_git` to return "", bypassing the very
`or None` conversion that made `dirty` unable to ever be False — the mock
asserted the intended behaviour while the real code did the opposite. It
now runs against a real temporary repository, and fails against the old
implementation.

Adds coverage for pip extras, all three revision-pin spellings, dist-tags
being excluded, URL userinfo not being mistaken for a pin, and pins being
claimed only for a harness whose image this run built.
Records which pin spellings are collected and why a floating dist-tag is
not, and explains when pinned_versions describes the image that actually
ran — including that a batch run reports "dockerfile" while a bare
clawbench-run --no-build reports "unverified".
…scovery-metadata

docs: align public metadata and distinguish historical benchmark scores
…-owned-cache

fix(reproduce): preserve user files when cleaning download caches
…enance

feat(runner): record code, corpus, and agent revisions in run-meta.json
@rgarcia
rgarcia requested a review from bmsaadat October 1, 2026 12:54
@rgarcia
rgarcia merged commit 8ee03a0 into main Oct 1, 2026
6 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants