Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
3 changes: 3 additions & 0 deletions CHANGELOG.md
Original file line number Diff line number Diff line change
Expand Up @@ -8,6 +8,7 @@ and this project adheres to [Semantic Versioning](https://semver.org/).

## [Unreleased]
### Added
- `run-meta.json` now carries a `provenance` block: the ClawBench version, commit, branch, and dirty state; the corpus suite and the revision of the commit that last touched it; and the agent and plugin versions pinned by the harness Dockerfile, with a `pins_source` saying whether those pins describe the image that actually ran. Every field is best-effort and null outside a git checkout, so a run never fails on a missing one. See [`docs/trace-cookbook.md`](docs/trace-cookbook.md#provenance).
- Added `scripts/export_openeval.py`, an additive script exporting a batch's `rescore-summary.json` as an [EvalPort](https://github.com/adhabnr-ux/evalport) `ResultSet` Thanks to [@adhabnr-ux](https://github.com/adhabnr-ux).
- Added a `--browser-runtime kernel` mode to the Harbor adapter that runs each task against one Kernel cloud browser, exposing only a credential-free CDP bridge to the agent, and finalizes the replay and deletes the browser during verification.

Expand All @@ -18,6 +19,8 @@ and this project adheres to [Semantic Versioning](https://semver.org/).
- Changed the default Harbor version to `0.22.0`.

### Fixed
- Isolate `clawbench-reproduce` downloads in a per-invocation cache directory so cleanup preserves existing work-directory files and removes only owned downloads, including on failure.
- Align public discovery metadata with the canonical repository and shipping corpus, label historical V1 scores in both READMEs, and correct the v0.10.0 citation release date.
- Host-timeout container termination now uses the lazy container-engine resolver.
- Added host-side container and batch-job timeouts so a wedged run cannot stall a batch indefinitely.
- Fixed a judge-provider outage (or an unparseable judge reply) being recorded as an agent failure. `run.py` now exits 3 instead of 1 when the judge never renders a verdict, `batch.py` gives it its own `judge_inconclusive` bucket in `batch-summary.json` instead of folding it into `failed`, and `clawbench-rescore` now retries a cached `match: null` verdict even without `--force`.
Expand Down
2 changes: 1 addition & 1 deletion CITATION.cff
Original file line number Diff line number Diff line change
Expand Up @@ -6,7 +6,7 @@ version: "0.10.0"
license: Apache-2.0
url: "https://claw-bench.com"
repository-code: "https://github.com/TIGER-AI-Lab/ClawBench"
date-released: "2026-06-22"
date-released: "2026-08-30"
identifiers:
- type: swh
value: "swh:1:snp:4baa3fe53cbce2b1abfae44a21170ab079e351a3"
Expand Down
6 changes: 3 additions & 3 deletions CONTRIBUTING.md
Original file line number Diff line number Diff line change
Expand Up @@ -11,7 +11,7 @@
| Add a new model config (in `models/models.yaml`) | ~1 hour | Get your model on the leaderboard |
| Fix a flaky task (find a broken task, propose a fix) | ~20 min | Keeps the leaderboard fair |
| Translate docs into a new language | ~1 hour | Chinese / Japanese / Korean / Spanish welcomed |
| Report a bug via [issue template](https://github.com/reacher-z/ClawBench/issues/new/choose) | ~5 min | Helps us prioritize |
| Report a bug via [issue template](https://github.com/TIGER-AI-Lab/ClawBench/issues/new/choose) | ~5 min | Helps us prioritize |

## Recognition for contributors

Expand All @@ -22,7 +22,7 @@

## Good first issues

The issue tracker has a [`good first issue`](https://github.com/reacher-z/ClawBench/labels/good%20first%20issue) label for contributions sized at "30 minutes, no container experience required." Typical entries:
The issue tracker has a [`good first issue`](https://github.com/TIGER-AI-Lab/ClawBench/labels/good%20first%20issue) label for contributions sized at "30 minutes, no container experience required." Typical entries:

- Add a new test case for a site we don't yet cover (list in the issue description)
- Verify a flagged-flaky task still works
Expand Down Expand Up @@ -142,7 +142,7 @@ Released versions should not be modified by contributors other than maintainers.

## Reporting issues

Please use the [issue templates](https://github.com/reacher-z/ClawBench/issues/new/choose) to report bugs or propose new test cases.
Please use the [issue templates](https://github.com/TIGER-AI-Lab/ClawBench/issues/new/choose) to report bugs or propose new test cases.

## Community

Expand Down
6 changes: 3 additions & 3 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -78,13 +78,13 @@

<div align="center">

**ClawBench is an open-source benchmark that evaluates AI browser agents on everyday online tasks — booking travel, ordering food, applying for jobs, managing email — across live websites. V1 lives in `test-cases/v1/`, V2 in `test-cases/v2/`. It measures end-to-end task success with a 5-layer recording pipeline and an agentic evaluator that compares each run against human references. Top score to date: 33.3%.**
**ClawBench is an open-source benchmark that evaluates AI browser agents on everyday online tasks — booking travel, ordering food, applying for jobs, managing email — across live websites. V1 lives in `test-cases/v1/`, V2 in `test-cases/v2/`. It measures end-to-end task success with a 5-layer recording pipeline and an agentic evaluator that compares each run against human references. See the [live leaderboard](https://huggingface.co/spaces/TIGER-Lab/ClawBench) for scores by corpus, harness, and scoring rubric.**

<img src="assets/clawbench_logo.png" alt="ClawBench logo" width="320">

We asked frontier AI agents to do what people do every day --<br/>
order food, book travel, apply for jobs, write reviews, manage projects.<br/>
**Even the best agent only completes about 1 in 3.**
**In the historical V1 paper evaluation, the strongest evaluated agent completed about 1 in 3 tasks.**

---

Expand Down Expand Up @@ -746,7 +746,7 @@ Yes for a V1 benchmark signal: the tasks span 143 live websites and 15 life cate
<details>
<summary><b>What's the current top score?</b></summary>

33.3% — roughly one task in three — from the strongest frontier model we evaluated on V1. The majority of tasks still defeat every model we've tested; the headroom is real, and the benchmark is not saturated.
The [live leaderboard](https://huggingface.co/spaces/TIGER-Lab/ClawBench) reports scores by corpus, harness, and scoring rubric. The 33.3% result quoted above belongs to the historical V1 paper evaluation; it is not a current overall maximum. Compare results using the same corpus revision, metric, and denominator.

</details>

Expand Down
6 changes: 3 additions & 3 deletions docs/README.zh-CN.md
Original file line number Diff line number Diff line change
Expand Up @@ -26,13 +26,13 @@

<div align="center">

**ClawBench 是一个开源基准,用于评测 AI browser agent 在日常在线任务上的表现 —— 订酒店、点外卖、投简历、管理邮件 —— 全部在真实网站上进行。V1 位于 `test-cases/v1/`,V2 位于 `test-cases/v2/`。它通过 5 层录制管线和对照人工参考轨迹的 agentic evaluator 衡量端到端任务完成率。目前最高分:33.3%。**
**ClawBench 是一个开源基准,用于评测 AI browser agent 在日常在线任务上的表现 —— 订酒店、点外卖、投简历、管理邮件 —— 全部在真实网站上进行。V1 位于 `test-cases/v1/`,V2 位于 `test-cases/v2/`。它通过 5 层录制管线和对照人工参考轨迹的 agentic evaluator 衡量端到端任务完成率。按语料、harness 和评分规则划分的结果请见[实时榜单](https://huggingface.co/spaces/TIGER-Lab/ClawBench)。**

<img src="../assets/clawbench_logo.png" alt="ClawBench logo" width="320">

我们让前沿 AI 智能体去做人们每天都在做的事 --<br/>
点外卖、订酒店、投简历、写评价、管理项目。<br/>
**即使最强的模型,也只能完成其中约三分之一。**
**在历史 V1 论文评测中,表现最好的受测 agent 完成了约三分之一的任务。**

---

Expand Down Expand Up @@ -639,7 +639,7 @@ ClawBench 的定位:**真实消费级网站、日常任务、端到端录制**
<details>
<summary><b>目前最高分是多少?</b></summary>

33.3% —— 大约三分之一的任务完成率 —— 来自我们在 V1 上评测过的最强前沿模型。大多数任务仍能击败我们测试过的每一个模型;提升空间真实存在,基准尚未饱和。
请查看按语料、harness 和评分规则划分的[实时榜单](https://huggingface.co/spaces/TIGER-Lab/ClawBench)。上文的 33.3% 属于历史 V1 论文评测,并非当前所有结果中的最高分。比较成绩时,请使用相同的语料版本、指标和分母。

</details>

Expand Down
61 changes: 60 additions & 1 deletion docs/trace-cookbook.md
Original file line number Diff line number Diff line change
Expand Up @@ -27,7 +27,7 @@ Each run directory contains:
| `screenshots/*.png` | Timestamped PNG per action | Vision grounding, GUI datasets |
| `recording.mp4` | Full session video (H.264, 15 fps) | Qualitative analysis, demos |
| `interception.json` | The final blocked request | Outcome labels (Stage-1) |
| `run-meta.json` | Model, harness, task, timing | Joins and filtering |
| `run-meta.json` | Model, harness, task, timing, provenance | Joins and filtering |

Pull a single model or task without downloading everything:

Expand Down Expand Up @@ -56,6 +56,65 @@ outcome = json.loads((run / "interception.json").read_text())
print(meta["model"], len(msgs), "messages,", len(acts), "actions")
```

### Provenance

`run-meta.json` carries a `provenance` block naming the exact revisions behind
the names in the rest of the file, so a row on a leaderboard can be traced back
to the code, corpus, and agent build that produced it:

```json
"provenance": {
"clawbench_version": "0.10.0",
"commit": "3f3599d...",
"branch": "main",
"dirty": false,
"corpus": {"suite": "v2", "path": "test-cases/v2", "revision": "62ee923..."},
"harness": {
"name": "openclaw",
"image_id": "sha256:...",
"agent_version": "2026.3.13",
"pinned_versions": {"openclaw": "2026.3.13"},
"pins_source": "dockerfile"
}
}
```

`corpus.revision` is the last commit that touched that suite, so two runs with
the same revision saw the same task text.

`harness.pinned_versions` comes from the version pins in the harness Dockerfile
— the agent and any plugins the image was built with. Both released versions
(`opencode-ai@1.4.4`, `litellm[proxy]==1.77.3`) and pinned revisions
(`pkg@github:owner/repo#<sha>`, `pkg @ git+https://…@<ref>`) count. A floating
dist-tag like `@next` is deliberately **not** recorded: it names a moving
target, so calling it a pin would be a false claim, and `agent_version` is
`null` for a harness pinned that way.

`pins_source` says whether those pins describe the image that actually ran:

| Value | Meaning |
|---|---|
| `"dockerfile"` | The image was built from this checkout's Dockerfile during this run, so its pins are the versions that ran. |
| `"unverified"` | The run reused an existing image (`--no-build`) that may predate the Dockerfile on disk. `pinned_versions` and `agent_version` are `null` — nothing is claimed. |

`clawbench-batch` builds the image once and then runs every task with
`--no-build`, so a batch run still reports `"dockerfile"`; a bare
`clawbench-run --no-build` reports `"unverified"`.

Every field is best-effort. A run from a PyPI install has no git checkout and
reports `commit: null`; a task from an explicit `--cases-dir` reports its suite
name but `revision: null`, because its history is not ClawBench's to claim.
`dirty: null` means the lookup failed, which is not the same claim as `false`.
Filter on these before comparing runs:

```python
same_code = {
run for run in runs
if run["provenance"]["commit"] == reference["provenance"]["commit"]
and run["provenance"]["dirty"] is False
}
```

## Recipes

**1. Agent SFT / distillation data.** `agent-messages.jsonl` from passing runs
Expand Down
23 changes: 17 additions & 6 deletions llms.txt
Original file line number Diff line number Diff line change
Expand Up @@ -2,26 +2,37 @@

> ClawBench is an open-source benchmark for evaluating browser and computer-use agents on everyday tasks performed on live websites.

ClawBench measures end-to-end task success while preserving replayable evidence from browser actions, screenshots, network requests, agent messages, and session recordings. The benchmark blocks only the final side-effecting submission request so tasks can run on real websites without completing purchases, bookings, or applications.
ClawBench measures end-to-end task success while preserving replayable evidence from browser actions, screenshots, network requests, agent messages, and session recordings. The runtime intercepts requests matching each task’s evaluation schema; scoring distinguishes interception from judged task fulfillment. See the scoring documentation for the evaluation contract.

## Canonical resources

- [Project page](https://claw-bench.com): overview, task explorer, leaderboard, and evaluation details.
- [GitHub repository](https://github.com/reacher-z/ClawBench): source code, task definitions, harnesses, and documentation.
- [GitHub repository](https://github.com/TIGER-AI-Lab/ClawBench): source code, task definitions, harnesses, and documentation.
- [Research paper](https://arxiv.org/abs/2604.08523): *ClawBench: Can AI Agents Complete Everyday Online Tasks?*
- [Citation metadata](https://github.com/reacher-z/ClawBench/blob/main/CITATION.cff): machine-readable citation record.
- [Citation metadata](https://github.com/TIGER-AI-Lab/ClawBench/blob/main/CITATION.cff): machine-readable citation record.
- [Task and leaderboard Space](https://huggingface.co/spaces/TIGER-Lab/ClawBench): live evaluation results and task explorer.
- [Task dataset](https://huggingface.co/datasets/NAIL-Group/ClawBench): V1/V2 task definitions and evaluation metadata.
- [V1 execution traces](https://huggingface.co/datasets/NAIL-Group/ClawBenchV1Trace): replayable run artifacts.
- [V2 execution traces](https://huggingface.co/datasets/TIGER-Lab/ClawBenchV2Trace): rolling V2 run artifacts.

## Quick facts

- V1: 153 tasks across 144 live websites and 15 life categories.
- V2: 130 tasks with six first-class agent harnesses.
- Shipping V1 corpus: 152 tasks across 143 live websites and 15 life categories.
- Shipping V2 corpus: 129 tasks across 63 live websites.
- The paper reports 153 V1 and 130 V2 tasks; two ASPCA tasks were removed after publication. Historical results may use the original corpus; report the corpus revision and denominator when comparing scores.
- V1 Lite: 20 selected V1 tasks, not an additional independent corpus.
- Evidence layers: session replay, screenshots, HTTP traffic, browser actions, and agent messages.
- Supported harnesses include OpenClaw, OpenCode, Claude Code, Codex, Browser-Use, Hermes, Pi, and other compatible browser agents.

## Evaluation and getting started

- [Installation and quick start](https://github.com/TIGER-AI-Lab/ClawBench#quick-start): prerequisites, model configuration, and CLI entry points.
- [Scoring contract](https://github.com/TIGER-AI-Lab/ClawBench/blob/main/eval/scoring.md): Stage 1 interception, Stage 2 judged fulfillment, and rubric definitions.
- [Contributing](https://github.com/TIGER-AI-Lab/ClawBench/blob/main/CONTRIBUTING.md): task contributions and review guidance.
- [Release history](https://github.com/TIGER-AI-Lab/ClawBench/blob/main/CHANGELOG.md): versioned changes.

Consult the linked leaderboard for current scores, selecting the corpus, harness, rubric, and snapshot. Historical paper scores are not a claim about today's best result.

## Citation

If you use ClawBench, cite the paper listed in [CITATION.cff](https://github.com/reacher-z/ClawBench/blob/main/CITATION.cff) and link to the canonical repository.
If you use ClawBench, cite the paper listed in [CITATION.cff](https://github.com/TIGER-AI-Lab/ClawBench/blob/main/CITATION.cff) and link to the canonical repository.
4 changes: 2 additions & 2 deletions pyproject.toml
Original file line number Diff line number Diff line change
Expand Up @@ -18,8 +18,8 @@ dependencies = [

[project.urls]
Homepage = "https://claw-bench.com"
Repository = "https://github.com/reacher-z/ClawBench"
Issues = "https://github.com/reacher-z/ClawBench/issues"
Repository = "https://github.com/TIGER-AI-Lab/ClawBench"
Issues = "https://github.com/TIGER-AI-Lab/ClawBench/issues"
Paper = "https://arxiv.org/abs/2604.08523"

[project.scripts]
Expand Down
Loading
Loading