diff --git a/CHANGELOG.md b/CHANGELOG.md
index cb130c2c..e321b51d 100644
--- a/CHANGELOG.md
+++ b/CHANGELOG.md
@@ -8,6 +8,7 @@ and this project adheres to [Semantic Versioning](https://semver.org/).
## [Unreleased]
### Added
+- `run-meta.json` now carries a `provenance` block: the ClawBench version, commit, branch, and dirty state; the corpus suite and the revision of the commit that last touched it; and the agent and plugin versions pinned by the harness Dockerfile, with a `pins_source` saying whether those pins describe the image that actually ran. Every field is best-effort and null outside a git checkout, so a run never fails on a missing one. See [`docs/trace-cookbook.md`](docs/trace-cookbook.md#provenance).
- Added `scripts/export_openeval.py`, an additive script exporting a batch's `rescore-summary.json` as an [EvalPort](https://github.com/adhabnr-ux/evalport) `ResultSet` Thanks to [@adhabnr-ux](https://github.com/adhabnr-ux).
- Added a `--browser-runtime kernel` mode to the Harbor adapter that runs each task against one Kernel cloud browser, exposing only a credential-free CDP bridge to the agent, and finalizes the replay and deletes the browser during verification.
@@ -18,6 +19,8 @@ and this project adheres to [Semantic Versioning](https://semver.org/).
- Changed the default Harbor version to `0.22.0`.
### Fixed
+- Isolate `clawbench-reproduce` downloads in a per-invocation cache directory so cleanup preserves existing work-directory files and removes only owned downloads, including on failure.
+- Align public discovery metadata with the canonical repository and shipping corpus, label historical V1 scores in both READMEs, and correct the v0.10.0 citation release date.
- Host-timeout container termination now uses the lazy container-engine resolver.
- Added host-side container and batch-job timeouts so a wedged run cannot stall a batch indefinitely.
- Fixed a judge-provider outage (or an unparseable judge reply) being recorded as an agent failure. `run.py` now exits 3 instead of 1 when the judge never renders a verdict, `batch.py` gives it its own `judge_inconclusive` bucket in `batch-summary.json` instead of folding it into `failed`, and `clawbench-rescore` now retries a cached `match: null` verdict even without `--force`.
diff --git a/CITATION.cff b/CITATION.cff
index a8db5a09..79b45d16 100644
--- a/CITATION.cff
+++ b/CITATION.cff
@@ -6,7 +6,7 @@ version: "0.10.0"
license: Apache-2.0
url: "https://claw-bench.com"
repository-code: "https://github.com/TIGER-AI-Lab/ClawBench"
-date-released: "2026-06-22"
+date-released: "2026-08-30"
identifiers:
- type: swh
value: "swh:1:snp:4baa3fe53cbce2b1abfae44a21170ab079e351a3"
diff --git a/CONTRIBUTING.md b/CONTRIBUTING.md
index 6373412c..f8e51d70 100644
--- a/CONTRIBUTING.md
+++ b/CONTRIBUTING.md
@@ -11,7 +11,7 @@
| Add a new model config (in `models/models.yaml`) | ~1 hour | Get your model on the leaderboard |
| Fix a flaky task (find a broken task, propose a fix) | ~20 min | Keeps the leaderboard fair |
| Translate docs into a new language | ~1 hour | Chinese / Japanese / Korean / Spanish welcomed |
-| Report a bug via [issue template](https://github.com/reacher-z/ClawBench/issues/new/choose) | ~5 min | Helps us prioritize |
+| Report a bug via [issue template](https://github.com/TIGER-AI-Lab/ClawBench/issues/new/choose) | ~5 min | Helps us prioritize |
## Recognition for contributors
@@ -22,7 +22,7 @@
## Good first issues
-The issue tracker has a [`good first issue`](https://github.com/reacher-z/ClawBench/labels/good%20first%20issue) label for contributions sized at "30 minutes, no container experience required." Typical entries:
+The issue tracker has a [`good first issue`](https://github.com/TIGER-AI-Lab/ClawBench/labels/good%20first%20issue) label for contributions sized at "30 minutes, no container experience required." Typical entries:
- Add a new test case for a site we don't yet cover (list in the issue description)
- Verify a flagged-flaky task still works
@@ -142,7 +142,7 @@ Released versions should not be modified by contributors other than maintainers.
## Reporting issues
-Please use the [issue templates](https://github.com/reacher-z/ClawBench/issues/new/choose) to report bugs or propose new test cases.
+Please use the [issue templates](https://github.com/TIGER-AI-Lab/ClawBench/issues/new/choose) to report bugs or propose new test cases.
## Community
diff --git a/README.md b/README.md
index 756c5082..dfaff623 100644
--- a/README.md
+++ b/README.md
@@ -78,13 +78,13 @@
-**ClawBench is an open-source benchmark that evaluates AI browser agents on everyday online tasks — booking travel, ordering food, applying for jobs, managing email — across live websites. V1 lives in `test-cases/v1/`, V2 in `test-cases/v2/`. It measures end-to-end task success with a 5-layer recording pipeline and an agentic evaluator that compares each run against human references. Top score to date: 33.3%.**
+**ClawBench is an open-source benchmark that evaluates AI browser agents on everyday online tasks — booking travel, ordering food, applying for jobs, managing email — across live websites. V1 lives in `test-cases/v1/`, V2 in `test-cases/v2/`. It measures end-to-end task success with a 5-layer recording pipeline and an agentic evaluator that compares each run against human references. See the [live leaderboard](https://huggingface.co/spaces/TIGER-Lab/ClawBench) for scores by corpus, harness, and scoring rubric.**

We asked frontier AI agents to do what people do every day --
order food, book travel, apply for jobs, write reviews, manage projects.
-**Even the best agent only completes about 1 in 3.**
+**In the historical V1 paper evaluation, the strongest evaluated agent completed about 1 in 3 tasks.**
---
@@ -746,7 +746,7 @@ Yes for a V1 benchmark signal: the tasks span 143 live websites and 15 life cate
What's the current top score?
-33.3% — roughly one task in three — from the strongest frontier model we evaluated on V1. The majority of tasks still defeat every model we've tested; the headroom is real, and the benchmark is not saturated.
+The [live leaderboard](https://huggingface.co/spaces/TIGER-Lab/ClawBench) reports scores by corpus, harness, and scoring rubric. The 33.3% result quoted above belongs to the historical V1 paper evaluation; it is not a current overall maximum. Compare results using the same corpus revision, metric, and denominator.
diff --git a/docs/README.zh-CN.md b/docs/README.zh-CN.md
index d5b930e6..60086380 100644
--- a/docs/README.zh-CN.md
+++ b/docs/README.zh-CN.md
@@ -26,13 +26,13 @@
-**ClawBench 是一个开源基准,用于评测 AI browser agent 在日常在线任务上的表现 —— 订酒店、点外卖、投简历、管理邮件 —— 全部在真实网站上进行。V1 位于 `test-cases/v1/`,V2 位于 `test-cases/v2/`。它通过 5 层录制管线和对照人工参考轨迹的 agentic evaluator 衡量端到端任务完成率。目前最高分:33.3%。**
+**ClawBench 是一个开源基准,用于评测 AI browser agent 在日常在线任务上的表现 —— 订酒店、点外卖、投简历、管理邮件 —— 全部在真实网站上进行。V1 位于 `test-cases/v1/`,V2 位于 `test-cases/v2/`。它通过 5 层录制管线和对照人工参考轨迹的 agentic evaluator 衡量端到端任务完成率。按语料、harness 和评分规则划分的结果请见[实时榜单](https://huggingface.co/spaces/TIGER-Lab/ClawBench)。**

我们让前沿 AI 智能体去做人们每天都在做的事 --
点外卖、订酒店、投简历、写评价、管理项目。
-**即使最强的模型,也只能完成其中约三分之一。**
+**在历史 V1 论文评测中,表现最好的受测 agent 完成了约三分之一的任务。**
---
@@ -639,7 +639,7 @@ ClawBench 的定位:**真实消费级网站、日常任务、端到端录制**
目前最高分是多少?
-33.3% —— 大约三分之一的任务完成率 —— 来自我们在 V1 上评测过的最强前沿模型。大多数任务仍能击败我们测试过的每一个模型;提升空间真实存在,基准尚未饱和。
+请查看按语料、harness 和评分规则划分的[实时榜单](https://huggingface.co/spaces/TIGER-Lab/ClawBench)。上文的 33.3% 属于历史 V1 论文评测,并非当前所有结果中的最高分。比较成绩时,请使用相同的语料版本、指标和分母。
diff --git a/docs/trace-cookbook.md b/docs/trace-cookbook.md
index 8fc0f38e..8629f381 100644
--- a/docs/trace-cookbook.md
+++ b/docs/trace-cookbook.md
@@ -27,7 +27,7 @@ Each run directory contains:
| `screenshots/*.png` | Timestamped PNG per action | Vision grounding, GUI datasets |
| `recording.mp4` | Full session video (H.264, 15 fps) | Qualitative analysis, demos |
| `interception.json` | The final blocked request | Outcome labels (Stage-1) |
-| `run-meta.json` | Model, harness, task, timing | Joins and filtering |
+| `run-meta.json` | Model, harness, task, timing, provenance | Joins and filtering |
Pull a single model or task without downloading everything:
@@ -56,6 +56,65 @@ outcome = json.loads((run / "interception.json").read_text())
print(meta["model"], len(msgs), "messages,", len(acts), "actions")
```
+### Provenance
+
+`run-meta.json` carries a `provenance` block naming the exact revisions behind
+the names in the rest of the file, so a row on a leaderboard can be traced back
+to the code, corpus, and agent build that produced it:
+
+```json
+"provenance": {
+ "clawbench_version": "0.10.0",
+ "commit": "3f3599d...",
+ "branch": "main",
+ "dirty": false,
+ "corpus": {"suite": "v2", "path": "test-cases/v2", "revision": "62ee923..."},
+ "harness": {
+ "name": "openclaw",
+ "image_id": "sha256:...",
+ "agent_version": "2026.3.13",
+ "pinned_versions": {"openclaw": "2026.3.13"},
+ "pins_source": "dockerfile"
+ }
+}
+```
+
+`corpus.revision` is the last commit that touched that suite, so two runs with
+the same revision saw the same task text.
+
+`harness.pinned_versions` comes from the version pins in the harness Dockerfile
+— the agent and any plugins the image was built with. Both released versions
+(`opencode-ai@1.4.4`, `litellm[proxy]==1.77.3`) and pinned revisions
+(`pkg@github:owner/repo#
`, `pkg @ git+https://…@[`) count. A floating
+dist-tag like `@next` is deliberately **not** recorded: it names a moving
+target, so calling it a pin would be a false claim, and `agent_version` is
+`null` for a harness pinned that way.
+
+`pins_source` says whether those pins describe the image that actually ran:
+
+| Value | Meaning |
+|---|---|
+| `"dockerfile"` | The image was built from this checkout's Dockerfile during this run, so its pins are the versions that ran. |
+| `"unverified"` | The run reused an existing image (`--no-build`) that may predate the Dockerfile on disk. `pinned_versions` and `agent_version` are `null` — nothing is claimed. |
+
+`clawbench-batch` builds the image once and then runs every task with
+`--no-build`, so a batch run still reports `"dockerfile"`; a bare
+`clawbench-run --no-build` reports `"unverified"`.
+
+Every field is best-effort. A run from a PyPI install has no git checkout and
+reports `commit: null`; a task from an explicit `--cases-dir` reports its suite
+name but `revision: null`, because its history is not ClawBench's to claim.
+`dirty: null` means the lookup failed, which is not the same claim as `false`.
+Filter on these before comparing runs:
+
+```python
+same_code = {
+ run for run in runs
+ if run["provenance"]["commit"] == reference["provenance"]["commit"]
+ and run["provenance"]["dirty"] is False
+}
+```
+
## Recipes
**1. Agent SFT / distillation data.** `agent-messages.jsonl` from passing runs
diff --git a/llms.txt b/llms.txt
index 12260ea1..aca614f1 100644
--- a/llms.txt
+++ b/llms.txt
@@ -2,14 +2,14 @@
> ClawBench is an open-source benchmark for evaluating browser and computer-use agents on everyday tasks performed on live websites.
-ClawBench measures end-to-end task success while preserving replayable evidence from browser actions, screenshots, network requests, agent messages, and session recordings. The benchmark blocks only the final side-effecting submission request so tasks can run on real websites without completing purchases, bookings, or applications.
+ClawBench measures end-to-end task success while preserving replayable evidence from browser actions, screenshots, network requests, agent messages, and session recordings. The runtime intercepts requests matching each task’s evaluation schema; scoring distinguishes interception from judged task fulfillment. See the scoring documentation for the evaluation contract.
## Canonical resources
- [Project page](https://claw-bench.com): overview, task explorer, leaderboard, and evaluation details.
-- [GitHub repository](https://github.com/reacher-z/ClawBench): source code, task definitions, harnesses, and documentation.
+- [GitHub repository](https://github.com/TIGER-AI-Lab/ClawBench): source code, task definitions, harnesses, and documentation.
- [Research paper](https://arxiv.org/abs/2604.08523): *ClawBench: Can AI Agents Complete Everyday Online Tasks?*
-- [Citation metadata](https://github.com/reacher-z/ClawBench/blob/main/CITATION.cff): machine-readable citation record.
+- [Citation metadata](https://github.com/TIGER-AI-Lab/ClawBench/blob/main/CITATION.cff): machine-readable citation record.
- [Task and leaderboard Space](https://huggingface.co/spaces/TIGER-Lab/ClawBench): live evaluation results and task explorer.
- [Task dataset](https://huggingface.co/datasets/NAIL-Group/ClawBench): V1/V2 task definitions and evaluation metadata.
- [V1 execution traces](https://huggingface.co/datasets/NAIL-Group/ClawBenchV1Trace): replayable run artifacts.
@@ -17,11 +17,22 @@ ClawBench measures end-to-end task success while preserving replayable evidence
## Quick facts
-- V1: 153 tasks across 144 live websites and 15 life categories.
-- V2: 130 tasks with six first-class agent harnesses.
+- Shipping V1 corpus: 152 tasks across 143 live websites and 15 life categories.
+- Shipping V2 corpus: 129 tasks across 63 live websites.
+- The paper reports 153 V1 and 130 V2 tasks; two ASPCA tasks were removed after publication. Historical results may use the original corpus; report the corpus revision and denominator when comparing scores.
+- V1 Lite: 20 selected V1 tasks, not an additional independent corpus.
- Evidence layers: session replay, screenshots, HTTP traffic, browser actions, and agent messages.
- Supported harnesses include OpenClaw, OpenCode, Claude Code, Codex, Browser-Use, Hermes, Pi, and other compatible browser agents.
+## Evaluation and getting started
+
+- [Installation and quick start](https://github.com/TIGER-AI-Lab/ClawBench#quick-start): prerequisites, model configuration, and CLI entry points.
+- [Scoring contract](https://github.com/TIGER-AI-Lab/ClawBench/blob/main/eval/scoring.md): Stage 1 interception, Stage 2 judged fulfillment, and rubric definitions.
+- [Contributing](https://github.com/TIGER-AI-Lab/ClawBench/blob/main/CONTRIBUTING.md): task contributions and review guidance.
+- [Release history](https://github.com/TIGER-AI-Lab/ClawBench/blob/main/CHANGELOG.md): versioned changes.
+
+Consult the linked leaderboard for current scores, selecting the corpus, harness, rubric, and snapshot. Historical paper scores are not a claim about today's best result.
+
## Citation
-If you use ClawBench, cite the paper listed in [CITATION.cff](https://github.com/reacher-z/ClawBench/blob/main/CITATION.cff) and link to the canonical repository.
+If you use ClawBench, cite the paper listed in [CITATION.cff](https://github.com/TIGER-AI-Lab/ClawBench/blob/main/CITATION.cff) and link to the canonical repository.
diff --git a/pyproject.toml b/pyproject.toml
index 2b6a4d4c..b5e19dcf 100644
--- a/pyproject.toml
+++ b/pyproject.toml
@@ -18,8 +18,8 @@ dependencies = [
[project.urls]
Homepage = "https://claw-bench.com"
-Repository = "https://github.com/reacher-z/ClawBench"
-Issues = "https://github.com/reacher-z/ClawBench/issues"
+Repository = "https://github.com/TIGER-AI-Lab/ClawBench"
+Issues = "https://github.com/TIGER-AI-Lab/ClawBench/issues"
Paper = "https://arxiv.org/abs/2604.08523"
[project.scripts]
diff --git a/src/clawbench/eval/reproduce.py b/src/clawbench/eval/reproduce.py
index e947bd9c..d56dcd98 100644
--- a/src/clawbench/eval/reproduce.py
+++ b/src/clawbench/eval/reproduce.py
@@ -18,7 +18,9 @@
import argparse
import json
-import shutil
+import tempfile
+from collections.abc import Iterator
+from contextlib import contextmanager
import subprocess
import sys
from pathlib import Path
@@ -112,6 +114,21 @@ def verdict(
return ok, "\n".join(lines)
+@contextmanager
+def download_cache(work_dir: Path, keep_cache: bool) -> Iterator[Path]:
+ """Own only a unique child directory, never the caller's work directory."""
+ work_dir.mkdir(parents=True, exist_ok=True)
+ if keep_cache:
+ cache_dir = Path(tempfile.mkdtemp(prefix="clawbench-", dir=work_dir))
+ print(f" Cache retained at: {cache_dir}")
+ yield cache_dir
+ else:
+ with tempfile.TemporaryDirectory(prefix="clawbench-", dir=work_dir) as tmp:
+ cache_dir = Path(tmp)
+ print(f" Temporary cache: {cache_dir}")
+ yield cache_dir
+
+
def main() -> int:
p = argparse.ArgumentParser(
description=__doc__, formatter_class=argparse.RawDescriptionHelpFormatter
@@ -139,7 +156,7 @@ def main() -> int:
"--work-dir",
type=Path,
default=Path("./reproduce-cache"),
- help="Local dir for HF download (default ./reproduce-cache)",
+ help="Parent directory for an isolated download cache (default ./reproduce-cache)",
)
p.add_argument(
"--keep-cache",
@@ -157,66 +174,64 @@ def main() -> int:
return 2
published = PUBLISHED_V2_HERMES[args.model]
- args.work_dir.mkdir(parents=True, exist_ok=True)
- print(
- f"== Reproducing {args.model} (n={published[3]}, tolerance ±{args.tolerance}pp) ==\n"
- )
- print("[1/3] Download trace subset from HF ...")
- batch_dir = download(args.model, args.work_dir)
-
- # Need a model batch root; HF subset puts task dirs under
- # batch-aligned-...//batch-...// → walk down.
- candidates = [p for p in batch_dir.rglob("batch-*") if p.is_dir()]
- if not candidates:
- # No nested batch-* dir? The batch_dir itself is the root.
- candidates = [batch_dir]
- inner_batch = candidates[0]
- # If there's a model sub-dir inside, use it
- sub = [
- c
- for c in inner_batch.iterdir()
- if c.is_dir() and not c.name.startswith("batch-logs")
- ]
- if sub and any(
- (
- c / next(c.iterdir(), Path("/dev/null")) / "data" / "interception.json"
- ).exists()
- for c in sub
- ):
- inner_batch = sub[0]
- print(f" → batch root: {inner_batch}")
-
- print(f"\n[2/3] Re-judge with {args.judge_model} (rubric={args.rubric}) ...")
- summary = rescore(inner_batch, args.judge_model, args.rubric)
-
- n = summary["n_total"]
- observed_icpt = 100.0 * summary["n_intercepted"] / n if n else 0.0
- observed_lenient = 100.0 * summary.get("reward_pct_lenient", 0)
- observed_strict = 100.0 * summary.get("reward_pct_strict", 0)
- observed = (observed_icpt, observed_lenient, observed_strict, n)
-
- print("\n[3/3] Compare to published row ...")
- ok, table = verdict(observed, published, args.tolerance)
- print(table)
- print()
- if ok:
+ with download_cache(args.work_dir, args.keep_cache) as cache_dir:
print(
- f"✓ PASS — reproduction within ±{args.tolerance} pp of published numbers."
+ f"== Reproducing {args.model} (n={published[3]}, tolerance ±{args.tolerance}pp) ==\n"
)
- else:
- print(f"✗ FAIL — at least one metric deviates more than ±{args.tolerance} pp.")
- print(" Possible causes:")
- print(" - Different judge model (we use deepseek-v4-pro on OpenRouter).")
- print(
- " - Different rubric (our prompts in src/clawbench/runner/judge_llm.py)."
- )
- print(" - HF dataset rev drift — try `hf download --revision `.")
-
- if not args.keep_cache:
- shutil.rmtree(args.work_dir, ignore_errors=True)
- print(f" (deleted {args.work_dir}; pass --keep-cache to keep traces)")
-
- return 0 if ok else 1
+ print("[1/3] Download trace subset from HF ...")
+ batch_dir = download(args.model, cache_dir)
+
+ # Need a model batch root; HF subset puts task dirs under
+ # batch-aligned-...//batch-...// → walk down.
+ candidates = [p for p in batch_dir.rglob("batch-*") if p.is_dir()]
+ if not candidates:
+ # No nested batch-* dir? The batch_dir itself is the root.
+ candidates = [batch_dir]
+ inner_batch = candidates[0]
+ # If there's a model sub-dir inside, use it
+ sub = [
+ c
+ for c in inner_batch.iterdir()
+ if c.is_dir() and not c.name.startswith("batch-logs")
+ ]
+ if sub and any(
+ (
+ c / next(c.iterdir(), Path("/dev/null")) / "data" / "interception.json"
+ ).exists()
+ for c in sub
+ ):
+ inner_batch = sub[0]
+ print(f" → batch root: {inner_batch}")
+
+ print(f"\n[2/3] Re-judge with {args.judge_model} (rubric={args.rubric}) ...")
+ summary = rescore(inner_batch, args.judge_model, args.rubric)
+
+ n = summary["n_total"]
+ observed_icpt = 100.0 * summary["n_intercepted"] / n if n else 0.0
+ observed_lenient = 100.0 * summary.get("reward_pct_lenient", 0)
+ observed_strict = 100.0 * summary.get("reward_pct_strict", 0)
+ observed = (observed_icpt, observed_lenient, observed_strict, n)
+
+ print("\n[3/3] Compare to published row ...")
+ ok, table = verdict(observed, published, args.tolerance)
+ print(table)
+ print()
+ if ok:
+ print(
+ f"✓ PASS — reproduction within ±{args.tolerance} pp of published numbers."
+ )
+ else:
+ print(
+ f"✗ FAIL — at least one metric deviates more than ±{args.tolerance} pp."
+ )
+ print(" Possible causes:")
+ print(" - Different judge model (we use deepseek-v4-pro on OpenRouter).")
+ print(
+ " - Different rubric (our prompts in src/clawbench/runner/judge_llm.py)."
+ )
+ print(" - HF dataset rev drift — try `hf download --revision `.")
+
+ return 0 if ok else 1
if __name__ == "__main__":
diff --git a/src/clawbench/runner/run_support/docker.py b/src/clawbench/runner/run_support/docker.py
index 82a3ff32..b08e75de 100644
--- a/src/clawbench/runner/run_support/docker.py
+++ b/src/clawbench/runner/run_support/docker.py
@@ -24,6 +24,7 @@
engine,
harness_image,
)
+from clawbench.runner.run_support.provenance import IMAGE_BUILT_ENV
from clawbench.runner.run_support.usage import (
fetch_openrouter_pricing,
format_usage_status,
@@ -248,6 +249,11 @@ def docker_build(harness: str = DEFAULT_HARNESS) -> None:
_build_one(BASE_DOCKERFILE, BASE_IMAGE)
_build_one(_HARNESS_DOCKERFILES[harness], target_image)
+ # Record that this image really was built from the Dockerfile in this
+ # checkout, so run provenance can tell its version pins apart from a
+ # possibly-stale image reused via --no-build. clawbench-batch builds once
+ # here and its child runs inherit the environment.
+ os.environ[IMAGE_BUILT_ENV] = harness
console.print(f"[green]✓[/] Container image ready ({target_image})")
diff --git a/src/clawbench/runner/run_support/metadata.py b/src/clawbench/runner/run_support/metadata.py
index c7f44522..c816631d 100644
--- a/src/clawbench/runner/run_support/metadata.py
+++ b/src/clawbench/runner/run_support/metadata.py
@@ -17,6 +17,7 @@
harness_image,
)
from clawbench.runner.run_support.docker import container_engine_version, image_id
+from clawbench.runner.run_support.provenance import make_provenance
from clawbench.runner.run_support.task import normalize_extra_info
SECRET_CONFIG_RE = re.compile(
@@ -230,6 +231,7 @@ def make_run_meta(
temperature = model_cfg.get("temperature") if model_cfg else None
max_tokens = model_cfg.get("max_tokens") if model_cfg else None
+ runtime = _runtime_meta(harness)
meta = {
"test_case": case_name,
**metadata,
@@ -252,7 +254,14 @@ def make_run_meta(
"infra_flags": classification["infra_flags"],
"run_metrics": classification["metrics"],
"usage": classification["metrics"].get("usage"),
- "runtime": _runtime_meta(harness),
+ "runtime": runtime,
+ # Which code, corpus, and agent build produced this trace — the part
+ # that makes a published row reproducible rather than merely labelled.
+ "provenance": make_provenance(
+ harness=harness,
+ harness_image_id=runtime.get("harness_image_id"),
+ task_dir=task_dir,
+ ),
"browser_runtime": browser_runtime,
"task": _task_meta(
task=task,
diff --git a/src/clawbench/runner/run_support/provenance.py b/src/clawbench/runner/run_support/provenance.py
new file mode 100644
index 00000000..7bf7f0fd
--- /dev/null
+++ b/src/clawbench/runner/run_support/provenance.py
@@ -0,0 +1,249 @@
+"""Which exact code, corpus, and agent produced a run.
+
+`run-meta.json` already records *what* was run — task, model, harness name,
+image ids. It does not record the revisions those names resolved to, so two
+runs of "openclaw on v2" a month apart are indistinguishable in the artifact
+even when the agent, the corpus, and ClawBench itself all moved.
+
+This module fills that gap: the ClawBench version and commit, the corpus
+revision, and the agent versions pinned by the harness image. Every lookup is
+best-effort and returns ``None`` rather than raising — a PyPI install has no
+git repository, a container host may have no `git` at all, and a missing
+provenance field must never fail a run that otherwise succeeded.
+
+Results are cached because a batch run builds one of these per task and the
+answers cannot change inside a single process.
+"""
+
+from __future__ import annotations
+
+import os
+import re
+import subprocess
+from functools import lru_cache
+from importlib.metadata import PackageNotFoundError, version
+from pathlib import Path
+from typing import Any
+
+from clawbench.utils.paths import ASSET_ROOT, HARNESS_ROOT, SOURCE_ROOT
+
+_GIT_TIMEOUT_S = 10
+
+# Version pins declared in a harness Dockerfile. This reads the pins the image
+# was built from rather than asking the agent for its version, which would need
+# a running container.
+#
+# Both a released version and a pinned revision count, because a preview-stage
+# agent is usually pinned to a commit rather than a version:
+# npm pkg@1.2.3 · @scope/pkg@1.2.3 · pkg@github:o/r# · pkg@git+https://…#][
+# pip pkg==1.2.3 · pkg[extra]==1.2.3 · pkg @ git+https://…@][
+#
+# A floating dist-tag (`pkg@latest`, `pkg@next`) is deliberately NOT collected:
+# it names a moving target, so recording it as a pin would be a false claim.
+_VERSION_OR_REVISION = r"\d[\w.+-]*|(?:github:|git\+)[^\s\"']+"
+_NPM_PIN_RE = re.compile(
+ r"(? str | None:
+ """Run a git command, or return ``None`` if it could not run.
+
+ ``None`` means *the lookup failed*. A command that succeeded with no
+ output returns ``""`` — for ``git status --porcelain`` that empty string
+ is the meaningful answer "clean", so it must not be folded into ``None``.
+ """
+ try:
+ result = subprocess.run(
+ ["git", "-C", str(repo), *args],
+ capture_output=True,
+ text=True,
+ timeout=_GIT_TIMEOUT_S,
+ )
+ except (OSError, subprocess.SubprocessError):
+ return None
+ if result.returncode != 0:
+ return None
+ return result.stdout.strip()
+
+
+@lru_cache(maxsize=1)
+def clawbench_version() -> str | None:
+ try:
+ return version("clawbench-eval")
+ except PackageNotFoundError:
+ return None
+
+
+@lru_cache(maxsize=1)
+def _repo_root() -> Path | None:
+ """The ClawBench git checkout, when running from source rather than a wheel."""
+ if SOURCE_ROOT is None or not (SOURCE_ROOT / ".git").exists():
+ return None
+ return SOURCE_ROOT
+
+
+@lru_cache(maxsize=1)
+def clawbench_commit() -> dict[str, Any]:
+ """Commit, branch, and dirty state of the ClawBench checkout."""
+ repo = _repo_root()
+ if repo is None:
+ return {"commit": None, "branch": None, "dirty": None}
+ status = _git(repo, "status", "--porcelain")
+ return {
+ # An empty commit or branch would mean a successful lookup that said
+ # nothing, which is no more useful than a failed one.
+ "commit": _git(repo, "rev-parse", "HEAD") or None,
+ "branch": _git(repo, "rev-parse", "--abbrev-ref", "HEAD") or None,
+ # "" is a clean tree; None is a lookup that failed, which is not the
+ # same claim as clean.
+ "dirty": (status != "") if status is not None else None,
+ }
+
+
+def _relative_to_assets(path: Path) -> str | None:
+ try:
+ return path.resolve().relative_to(ASSET_ROOT.resolve()).as_posix()
+ except (OSError, ValueError):
+ return None
+
+
+@lru_cache(maxsize=8)
+def _corpus_commit(suite_path: str) -> str | None:
+ """Last commit that touched this corpus directory."""
+ repo = _repo_root()
+ if repo is None:
+ return None
+ return _git(repo, "log", "-1", "--format=%H", "--", suite_path) or None
+
+
+def corpus_meta(task_dir: Path | None) -> dict[str, Any]:
+ """Which corpus a task came from, and at which revision.
+
+ A task outside the bundled corpora (an explicit ``--cases-dir``) reports
+ its suite name but no revision — its history is not ClawBench's to claim.
+ """
+ if task_dir is None:
+ return {"suite": None, "path": None, "revision": None}
+ relative = _relative_to_assets(task_dir)
+ if relative is None:
+ return {"suite": task_dir.parent.name, "path": None, "revision": None}
+ # e.g. "test-cases/v2/v2-047-daily-life-personal-care-taskrabbit"
+ parts = relative.split("/")
+ suite_path = "/".join(parts[:2])
+ return {
+ "suite": parts[1] if len(parts) > 1 else parts[0],
+ "path": suite_path,
+ "revision": _corpus_commit(suite_path),
+ }
+
+
+@lru_cache(maxsize=32)
+def harness_pins(harness: str) -> dict[str, str]:
+ """Agent and plugin versions pinned by a harness Dockerfile.
+
+ Reads the pins the image was built from — `opencode-ai@1.4.4`,
+ `@playwright/mcp@0.0.70`, `pip install foo==1.2` — so a trace records the
+ exact agent build it used. Returns an empty mapping for a harness with no
+ version pins, or one whose Dockerfile cannot be read.
+ """
+ if harness in ("human", ""):
+ return {}
+ try:
+ from clawbench.runner.run_support.harness_registry import HARNESS_REGISTRY
+
+ dockerfile = HARNESS_REGISTRY.harness_dockerfiles.get(harness)
+ except (ImportError, ValueError):
+ dockerfile = None
+ if dockerfile is None:
+ candidates = sorted(HARNESS_ROOT.glob(f"{harness}/Dockerfile.*"))
+ dockerfile = candidates[0] if candidates else None
+ if dockerfile is None or not dockerfile.is_file():
+ return {}
+ try:
+ text = dockerfile.read_text(encoding="utf-8", errors="replace")
+ except OSError:
+ return {}
+
+ pins: dict[str, str] = {}
+ for line in text.splitlines():
+ if _PIN_SKIP_RE.match(line):
+ continue
+ for pattern in _PIN_PATTERNS:
+ for name, pinned in pattern.findall(line):
+ pins.setdefault(name.strip(), pinned)
+ return pins
+
+
+def _agent_version(harness: str, pins: dict[str, str]) -> str | None:
+ """The pin that is the agent itself, not one of its plugins.
+
+ Package names rarely equal the harness name exactly (`opencode` ships as
+ `opencode-ai`), so fall back to the pin whose package name contains it.
+ """
+ if harness in pins:
+ return pins[harness]
+ for name, pinned in pins.items():
+ if harness in name:
+ return pinned
+ return None
+
+
+def image_built_from_dockerfile(harness: str) -> bool:
+ """Whether the image this run uses was built from the Dockerfile we read.
+
+ With ``--no-build`` the container image can be arbitrarily older than the
+ Dockerfile on disk, so its pins are not evidence of what actually ran.
+ ``docker_build()`` records the harness it built; ``clawbench-batch`` builds
+ once and its children inherit that environment, so a batch run still
+ reports real pins while a bare ``--no-build`` run does not.
+ """
+ return os.environ.get(IMAGE_BUILT_ENV) == harness
+
+
+def harness_meta(harness: str, image_id: str | None) -> dict[str, Any]:
+ if not image_built_from_dockerfile(harness):
+ # Claim nothing rather than report a pin the running image may not have.
+ return {
+ "name": harness,
+ "image_id": image_id,
+ "pinned_versions": None,
+ "agent_version": None,
+ "pins_source": "unverified",
+ }
+ pins = harness_pins(harness)
+ return {
+ "name": harness,
+ "image_id": image_id,
+ "pinned_versions": pins or None,
+ # Named separately because it is the one a leaderboard row cites.
+ "agent_version": _agent_version(harness, pins),
+ "pins_source": "dockerfile",
+ }
+
+
+def make_provenance(
+ *,
+ harness: str,
+ harness_image_id: str | None,
+ task_dir: Path | None,
+) -> dict[str, Any]:
+ """The provenance block written into ``run-meta.json``."""
+ return {
+ "clawbench_version": clawbench_version(),
+ **clawbench_commit(),
+ "corpus": corpus_meta(task_dir),
+ "harness": harness_meta(harness, harness_image_id),
+ }
diff --git a/tests/test_provenance.py b/tests/test_provenance.py
new file mode 100644
index 00000000..0d5cabcb
--- /dev/null
+++ b/tests/test_provenance.py
@@ -0,0 +1,356 @@
+"""Run provenance: which code, corpus, and agent build produced a trace."""
+
+from __future__ import annotations
+
+import subprocess
+from pathlib import Path
+
+import pytest
+
+from clawbench.runner.run_support import provenance
+from clawbench.utils.paths import ASSET_ROOT
+
+
+@pytest.fixture(autouse=True)
+def _clear_provenance_caches() -> None:
+ for fn in (
+ provenance.clawbench_version,
+ provenance._repo_root,
+ provenance.clawbench_commit,
+ provenance._corpus_commit,
+ provenance.harness_pins,
+ ):
+ fn.cache_clear()
+
+
+@pytest.fixture
+def built_image(monkeypatch: pytest.MonkeyPatch):
+ """Mark a harness image as built from the Dockerfile in this checkout."""
+
+ def mark(harness: str) -> None:
+ monkeypatch.setenv(provenance.IMAGE_BUILT_ENV, harness)
+
+ return mark
+
+
+def _git_repo(path: Path) -> Path:
+ """A real git repository — mocking `_git` is what hid the dirty-flag bug."""
+ path.mkdir(parents=True, exist_ok=True)
+ for args in (
+ ["init", "-q", "."],
+ [
+ "-c",
+ "user.email=t@example.test",
+ "-c",
+ "user.name=t",
+ "commit",
+ "-q",
+ "--allow-empty",
+ "-m",
+ "init",
+ ],
+ ):
+ result = subprocess.run(
+ ["git", "-C", str(path), *args], capture_output=True, text=True
+ )
+ if result.returncode != 0:
+ pytest.skip(f"git unavailable: {result.stderr.strip()}")
+ return path
+
+
+# ---------------------------------------------------------------------------
+# Harness pins
+# ---------------------------------------------------------------------------
+
+
+@pytest.mark.parametrize(
+ ("harness", "package"),
+ [
+ ("openclaw", "openclaw"),
+ ("opencode", "opencode-ai"),
+ ("claude-code", "@anthropic-ai/claude-code"),
+ ("browser-use", "browser-use"),
+ ],
+)
+def test_bundled_harnesses_report_their_pinned_agent(
+ harness: str, package: str, built_image
+) -> None:
+ built_image(harness)
+ pins = provenance.harness_pins(harness)
+
+ assert package in pins, f"{harness} Dockerfile pin not detected: {pins}"
+ assert pins[package][0].isdigit()
+ assert provenance.harness_meta(harness, None)["agent_version"] == pins[package]
+
+
+def test_base_image_lines_are_not_mistaken_for_agent_pins() -> None:
+ """`FROM node:24-slim` and `COPY --from=...uv:0.11.6` are not agent pins."""
+ pins = provenance.harness_pins("opencode")
+
+ assert "node" not in pins
+ assert "uv" not in pins
+ assert pins["@playwright/mcp"] == "0.0.70"
+
+
+def test_harness_without_pins_reports_nothing_rather_than_guessing(
+ built_image,
+) -> None:
+ built_image("null")
+ meta = provenance.harness_meta("null", "sha256:abc")
+
+ assert meta["pinned_versions"] is None
+ assert meta["agent_version"] is None
+ assert meta["image_id"] == "sha256:abc"
+ # We did look; the Dockerfile simply pins nothing.
+ assert meta["pins_source"] == "dockerfile"
+
+
+def test_reused_image_claims_no_pins(monkeypatch: pytest.MonkeyPatch) -> None:
+ """--no-build can run an image far older than the Dockerfile on disk.
+
+ Its pins are then not evidence of what actually ran, so nothing is claimed.
+ """
+ monkeypatch.delenv(provenance.IMAGE_BUILT_ENV, raising=False)
+
+ meta = provenance.harness_meta("openclaw", "sha256:abc")
+
+ assert meta["pins_source"] == "unverified"
+ assert meta["pinned_versions"] is None
+ assert meta["agent_version"] is None
+ # The image id is still a fact about the run, and is still reported.
+ assert meta["image_id"] == "sha256:abc"
+
+
+def test_pins_are_claimed_for_the_built_harness_only(
+ monkeypatch: pytest.MonkeyPatch,
+) -> None:
+ monkeypatch.setenv(provenance.IMAGE_BUILT_ENV, "openclaw")
+
+ assert provenance.harness_meta("openclaw", None)["pins_source"] == "dockerfile"
+ assert provenance.harness_meta("codex", None)["pins_source"] == "unverified"
+
+
+def test_human_and_unknown_harnesses_are_safe() -> None:
+ assert provenance.harness_pins("human") == {}
+ assert provenance.harness_pins("not-a-harness") == {}
+
+
+def _demo_harness(tmp_path: Path, monkeypatch: pytest.MonkeyPatch, body: str) -> None:
+ harness_dir = tmp_path / "demo"
+ harness_dir.mkdir(exist_ok=True)
+ (harness_dir / "Dockerfile.demo").write_text(body, encoding="utf-8")
+ monkeypatch.setattr(provenance, "HARNESS_ROOT", tmp_path)
+ monkeypatch.setenv(provenance.IMAGE_BUILT_ENV, "demo")
+ provenance.harness_pins.cache_clear()
+
+
+def test_pip_pins_are_detected(tmp_path: Path, monkeypatch: pytest.MonkeyPatch) -> None:
+ _demo_harness(
+ tmp_path,
+ monkeypatch,
+ "FROM python:3.11-slim\nRUN pip install demo-agent==2.4.1 helper-plugin==0.9\n",
+ )
+
+ pins = provenance.harness_pins("demo")
+
+ assert pins == {"demo-agent": "2.4.1", "helper-plugin": "0.9"}
+ assert provenance.harness_meta("demo", None)["agent_version"] == "2.4.1"
+
+
+def test_pip_extras_do_not_drop_the_pin(
+ tmp_path: Path, monkeypatch: pytest.MonkeyPatch
+) -> None:
+ """`litellm[proxy]==1.77.3` used to match nothing, losing the pin silently."""
+ _demo_harness(
+ tmp_path,
+ monkeypatch,
+ 'FROM python:3.11-slim\nRUN pip install "litellm[proxy,extra]==1.77.3"\n',
+ )
+
+ assert provenance.harness_pins("demo") == {"litellm[proxy,extra]": "1.77.3"}
+
+
+@pytest.mark.parametrize(
+ ("line", "expected"),
+ [
+ (
+ "RUN npm install -g demo@github:acme/demo#a1b2c3d4e5f6a7b8",
+ {"demo": "github:acme/demo#a1b2c3d4e5f6a7b8"},
+ ),
+ (
+ "RUN npm install -g demo@git+https://github.com/acme/demo.git#v1.2.3",
+ {"demo": "git+https://github.com/acme/demo.git#v1.2.3"},
+ ),
+ (
+ 'RUN pip install "demo @ git+https://github.com/acme/demo@abc1234"',
+ {"demo": "git+https://github.com/acme/demo@abc1234"},
+ ),
+ ],
+)
+def test_revision_pins_are_detected(
+ tmp_path: Path,
+ monkeypatch: pytest.MonkeyPatch,
+ line: str,
+ expected: dict[str, str],
+) -> None:
+ """A preview-stage agent is pinned to a commit, not a released version."""
+ _demo_harness(tmp_path, monkeypatch, f"FROM python:3.11-slim\n{line}\n")
+
+ assert provenance.harness_pins("demo") == expected
+
+
+def test_floating_dist_tags_are_not_recorded_as_pins(
+ tmp_path: Path, monkeypatch: pytest.MonkeyPatch
+) -> None:
+ """`@next` names a moving target; calling it a pin would be a false claim."""
+ _demo_harness(
+ tmp_path, monkeypatch, "FROM node:24-slim\nRUN npm install -g demo@next\n"
+ )
+
+ assert provenance.harness_pins("demo") == {}
+ assert provenance.harness_meta("demo", None)["agent_version"] is None
+
+
+def test_url_userinfo_is_not_mistaken_for_a_pin(
+ tmp_path: Path, monkeypatch: pytest.MonkeyPatch
+) -> None:
+ _demo_harness(
+ tmp_path,
+ monkeypatch,
+ "FROM python:3.11-slim\nRUN curl -sf https://user@host.example/x\n",
+ )
+
+ assert provenance.harness_pins("demo") == {}
+
+
+# ---------------------------------------------------------------------------
+# Corpus revision
+# ---------------------------------------------------------------------------
+
+
+def test_bundled_task_reports_its_suite_and_path() -> None:
+ task_dir = ASSET_ROOT / "test-cases" / "v2" / "example-task"
+
+ corpus = provenance.corpus_meta(task_dir)
+
+ assert corpus["suite"] == "v2"
+ assert corpus["path"] == "test-cases/v2"
+
+
+def test_external_cases_dir_claims_no_revision(tmp_path: Path) -> None:
+ """An explicit --cases-dir has a history that is not ClawBench's to claim."""
+ corpus = provenance.corpus_meta(tmp_path / "my-suite" / "task-1")
+
+ assert corpus["suite"] == "my-suite"
+ assert corpus["path"] is None
+ assert corpus["revision"] is None
+
+
+def test_missing_task_dir_is_all_null() -> None:
+ assert provenance.corpus_meta(None) == {
+ "suite": None,
+ "path": None,
+ "revision": None,
+ }
+
+
+# ---------------------------------------------------------------------------
+# ClawBench revision
+# ---------------------------------------------------------------------------
+
+
+def test_commit_lookup_outside_a_checkout_is_null_not_a_failure(
+ monkeypatch: pytest.MonkeyPatch,
+) -> None:
+ """A PyPI install has no git repository; that must not fail a run."""
+ monkeypatch.setattr(provenance, "SOURCE_ROOT", None)
+ provenance._repo_root.cache_clear()
+ provenance.clawbench_commit.cache_clear()
+
+ assert provenance.clawbench_commit() == {
+ "commit": None,
+ "branch": None,
+ "dirty": None,
+ }
+
+
+def test_missing_git_binary_is_null_not_a_failure(
+ monkeypatch: pytest.MonkeyPatch,
+) -> None:
+ def no_git(*args: object, **kwargs: object) -> None:
+ raise FileNotFoundError("git")
+
+ monkeypatch.setattr(subprocess, "run", no_git)
+
+ assert provenance._git(Path("."), "rev-parse", "HEAD") is None
+
+
+def test_dirty_is_null_when_the_lookup_itself_failed(
+ monkeypatch: pytest.MonkeyPatch, tmp_path: Path
+) -> None:
+ """Unknown is not the same claim as clean."""
+ monkeypatch.setattr(provenance, "_repo_root", lambda: tmp_path)
+ monkeypatch.setattr(provenance, "_git", lambda *args: None)
+ provenance.clawbench_commit.cache_clear()
+
+ assert provenance.clawbench_commit()["dirty"] is None
+
+
+def test_git_distinguishes_empty_output_from_failure(tmp_path: Path) -> None:
+ """A clean `git status` succeeds with no output; that is an answer.
+
+ Folding "" into None here is what made `dirty` unable to ever be False.
+ """
+ repo = _git_repo(tmp_path / "repo")
+
+ assert provenance._git(repo, "status", "--porcelain") == ""
+ assert provenance._git(repo, "rev-parse", "HEAD")
+ assert provenance._git(repo, "not-a-git-command") is None
+
+
+def test_clean_and_dirty_trees_are_distinguished(
+ monkeypatch: pytest.MonkeyPatch, tmp_path: Path
+) -> None:
+ repo = _git_repo(tmp_path / "repo")
+ monkeypatch.setattr(provenance, "_repo_root", lambda: repo)
+
+ provenance.clawbench_commit.cache_clear()
+ clean = provenance.clawbench_commit()
+ assert clean["dirty"] is False
+ assert clean["commit"]
+
+ (repo / "changed.txt").write_text("edited", encoding="utf-8")
+ provenance.clawbench_commit.cache_clear()
+ assert provenance.clawbench_commit()["dirty"] is True
+
+
+# ---------------------------------------------------------------------------
+# The assembled block
+# ---------------------------------------------------------------------------
+
+
+def test_provenance_block_has_a_stable_shape(built_image) -> None:
+ built_image("openclaw")
+ block = provenance.make_provenance(
+ harness="openclaw",
+ harness_image_id="sha256:abc",
+ task_dir=ASSET_ROOT / "test-cases" / "v2" / "example-task",
+ )
+
+ assert set(block) == {
+ "clawbench_version",
+ "commit",
+ "branch",
+ "dirty",
+ "corpus",
+ "harness",
+ }
+ assert set(block["corpus"]) == {"suite", "path", "revision"}
+ assert set(block["harness"]) == {
+ "name",
+ "image_id",
+ "pinned_versions",
+ "agent_version",
+ "pins_source",
+ }
+ assert block["harness"]["name"] == "openclaw"
diff --git a/tests/test_reproduce_cache.py b/tests/test_reproduce_cache.py
new file mode 100644
index 00000000..562b11f3
--- /dev/null
+++ b/tests/test_reproduce_cache.py
@@ -0,0 +1,89 @@
+"""The reproduction CLI must never own its caller's entire work directory."""
+
+from pathlib import Path
+
+import pytest
+
+from clawbench.eval import reproduce
+
+
+@pytest.mark.parametrize(
+ "outcome", ["pass", "fail", "download_error", "judge_error", "interrupt"]
+)
+@pytest.mark.parametrize("keep_cache", [False, True])
+def test_cli_preserves_user_files_and_cleans_only_its_cache(
+ tmp_path: Path, monkeypatch: pytest.MonkeyPatch, outcome: str, keep_cache: bool
+) -> None:
+ work_dir = tmp_path / "existing-work"
+ work_dir.mkdir()
+ sentinel = work_dir / "unrelated.txt"
+ sentinel.write_text("user data")
+ previous_cache = work_dir / "clawbench-previous"
+ previous_cache.mkdir()
+ (previous_cache / "trace.json").write_text("previous run")
+ destinations = []
+
+ def download(model: str, dest: Path) -> Path:
+ destinations.append(dest)
+ assert dest.parent == work_dir
+ assert dest != previous_cache
+ (dest / "download.txt").write_text("owned download")
+ if outcome == "download_error":
+ raise SystemExit("download failed")
+ batch = dest / "traces"
+ batch.mkdir()
+ return batch
+
+ def rescore(batch_dir: Path, judge_model: str, rubric: str) -> dict:
+ if outcome == "judge_error":
+ raise RuntimeError("judge failed")
+ if outcome == "interrupt":
+ raise KeyboardInterrupt
+ return {
+ "n_total": 129,
+ "n_intercepted": 4 if outcome == "pass" else 129,
+ "reward_pct_lenient": 3 / 129,
+ "reward_pct_strict": 0,
+ }
+
+ monkeypatch.setattr(reproduce, "download", download)
+ monkeypatch.setattr(reproduce, "rescore", rescore)
+ argv = [
+ "clawbench-reproduce",
+ "--model",
+ "deepseek-v4-flash",
+ "--work-dir",
+ str(work_dir),
+ ]
+ if keep_cache:
+ argv.append("--keep-cache")
+ monkeypatch.setattr("sys.argv", argv)
+ errors = {
+ "download_error": SystemExit,
+ "judge_error": RuntimeError,
+ "interrupt": KeyboardInterrupt,
+ }
+ if outcome in errors:
+ with pytest.raises(errors[outcome]):
+ reproduce.main()
+ else:
+ assert reproduce.main() == (0 if outcome == "pass" else 1)
+
+ assert sentinel.read_text() == "user data"
+ assert (previous_cache / "trace.json").read_text() == "previous run"
+ assert len(destinations) == 1
+ assert destinations[0].exists() is keep_cache
+ if keep_cache:
+ assert (destinations[0] / "download.txt").read_text() == "owned download"
+
+
+def test_overlapping_invocations_have_independent_caches(tmp_path: Path) -> None:
+ with reproduce.download_cache(tmp_path, False) as first:
+ (first / "active.txt").write_text("first run")
+ with reproduce.download_cache(tmp_path, False) as second:
+ assert second != first
+ assert first.is_dir() and second.is_dir()
+ assert not second.exists()
+ assert (first / "active.txt").read_text() == "first run"
+ assert not first.exists()
+ assert tmp_path.is_dir()
diff --git a/tests/test_results_and_metadata.py b/tests/test_results_and_metadata.py
index f407ed76..a4b65d69 100644
--- a/tests/test_results_and_metadata.py
+++ b/tests/test_results_and_metadata.py
@@ -258,3 +258,11 @@ def test_run_metadata_redacts_model_and_judge_secrets(
assert meta["browser_runtime"]["cleanup_status"] == "released"
assert meta["usage"]["estimated_cost_usd"] == 0.0042
assert meta["run_metrics"]["usage"]["total_tokens"] == 123
+
+ provenance = meta["provenance"]
+ assert provenance["harness"]["name"] == "openclaw"
+ assert provenance["corpus"]["suite"] == "v1"
+ assert provenance["harness"]["image_id"] == meta["runtime"]["harness_image_id"]
+ # This run reused an existing image, so its pins are not claimed.
+ assert provenance["harness"]["pins_source"] == "unverified"
+ assert provenance["harness"]["agent_version"] is None
]