How well do LLMs actually know cheese?
CheeseBench is an open, browser-first benchmark suite that stress-tests AI models on five axes of dairy knowledge — from ranking a thousand cheeses by popularity to identifying them from photographs and estimating their market price. No toy questions, no multiple choice: models must produce, retrieve and reason, and the scoring is ruthless.
A model that cannot tell Comté from Taleggio is a model that is hallucinating.
| Test | What the model must do | Scoring |
|---|---|---|
| Popularity-1000 | List 1,000 cheeses sorted by worldwide popularity | Kendall τ-b rank correlation, top-50 precision, duplicate & hallucination rates |
| Vision-100 | Identify 100 cheeses from photographs alone | Free-text identification accuracy (alias-aware) |
| Origin-100 | Name the country of origin for 100 cheeses | Exact country match |
| Market Price-100 | Estimate retail price (USD/kg, Jan 2024) for 100 cheeses | Exponential log-error; hit = within ±25% |
| Wikipedia Tool-20 | Answer 20 questions using only Wikipedia search/summary tools | Task success + tool-call efficiency |
Ground truth comes from a curated dataset of 1,000 cheeses built from cross-Wiki
sources (scripts/build_dataset.py). Every run is stored, versioned and replayable.
┌─────────────────────────────┐ ┌──────────────────────────┐
│ Browser (React + Vite) │ │ Node server (zero-deps) │
│ │ JSON │ │
│ • test runners │ ─────► │ /api/runs CRUD │
│ • direct provider calls │ /api │ /api/leaderboard │
│ • keys in localStorage │ │ /api/fans sanitized │
└─────────────┬───────────────┘ │ SQLite (node:sqlite) │
│ HTTPS │ serves dist/ │
OpenAI · Anthropic · Google · └──────────────────────────┘
Grok · any OpenAI-compatible local
endpoint (Ollama, LM Studio, llama.cpp, vLLM)
- Frontend runs the benchmarks and calls model providers directly — your API keys never leave the browser (localStorage).
- Backend is a single-file, dependency-free Node server using the built-in
node:sqlitemodule. It stores runs, aggregates the leaderboard and keeps the fan club.
npm install
npm run dev # API on :8787 + Vite on :5173 (proxies /api)- Open the app → Settings → add a model (or point it at a local server and hit Test connection)
- Pick a test under Tests → Run
- Watch the leaderboard fill up at /leaderboard
Production:
npm start # build + serve everything from the Node server| Command | Description |
|---|---|
npm run dev |
API + web dev servers together |
npm run dev:api / npm run dev:web |
individually |
npm run build |
typecheck + production bundle |
npm start |
build, then serve app + API from one port |
npm test |
vitest suite (scoring, parsing, metrics) |
npm run lint |
oxlint |
| Method | Endpoint | Description |
|---|---|---|
GET |
/api/runs |
all stored runs |
POST |
/api/runs |
store a run result |
GET/DELETE |
/api/runs/:id |
fetch / delete one run |
DELETE |
/api/runs |
wipe the board |
GET |
/api/leaderboard |
per-model avg/best per test |
GET/POST |
/api/fans |
the fan club (sanitized, rate-limit-friendly) |
GET |
/api/health |
liveness + DB path |
Database lives at data/cheesebench.db (override with CHEESEBENCH_DB, port with PORT).
- OpenAI, Anthropic, Google, xAI (Grok) — direct from the browser
- Local / self-hosted — any OpenAI-compatible
/v1/chat/completionsendpoint (Ollama, LM Studio, llama.cpp, vLLM), with model discovery and connection testing - Per-model overrides: API key, max tokens, request concurrency (1–32)
No models ship enabled by default. You bring the models; we bring the questions.
server/index.js zero-dependency API + static server (node:sqlite)
src/lib/providers.ts provider adapters (OpenAI/Anthropic/Google/Grok/local)
src/lib/tests/ one runner + scorer per benchmark
src/lib/scoring/ Kendall τ-b, alias matching, price error, tool metrics
src/data/ cheeses.json (ground truth), cheeseBoard.json
src/pages/ landing, leaderboard, tests, runs, top1000, games, fans…
scripts/ dev launcher + Wikipedia dataset builder
This benchmark is scientifically rigorous the way a cheese board is peer-reviewed: tasted by everyone, trusted by no one. Scores are a snapshot of model behavior on curated dairy trivia and should not be used to draw conclusions about anything else. Please do not feed the results to your fondue.
CheeseBench — aged in the open. 🧀