Skip to content

Repository files navigation

🧀 CheeseBench

How well do LLMs actually know cheese?

CheeseBench is an open, browser-first benchmark suite that stress-tests AI models on five axes of dairy knowledge — from ranking a thousand cheeses by popularity to identifying them from photographs and estimating their market price. No toy questions, no multiple choice: models must produce, retrieve and reason, and the scoring is ruthless.

A model that cannot tell Comté from Taleggio is a model that is hallucinating.


The Tests

Test What the model must do Scoring
Popularity-1000 List 1,000 cheeses sorted by worldwide popularity Kendall τ-b rank correlation, top-50 precision, duplicate & hallucination rates
Vision-100 Identify 100 cheeses from photographs alone Free-text identification accuracy (alias-aware)
Origin-100 Name the country of origin for 100 cheeses Exact country match
Market Price-100 Estimate retail price (USD/kg, Jan 2024) for 100 cheeses Exponential log-error; hit = within ±25%
Wikipedia Tool-20 Answer 20 questions using only Wikipedia search/summary tools Task success + tool-call efficiency

Ground truth comes from a curated dataset of 1,000 cheeses built from cross-Wiki sources (scripts/build_dataset.py). Every run is stored, versioned and replayable.

Architecture

┌─────────────────────────────┐        ┌──────────────────────────┐
│  Browser (React + Vite)     │        │  Node server (zero-deps) │
│                             │  JSON  │                          │
│  • test runners             │ ─────► │  /api/runs     CRUD      │
│  • direct provider calls    │  /api  │  /api/leaderboard        │
│  • keys in localStorage     │        │  /api/fans     sanitized │
└─────────────┬───────────────┘        │  SQLite (node:sqlite)    │
              │ HTTPS                  │  serves dist/            │
   OpenAI · Anthropic · Google ·       └──────────────────────────┘
   Grok · any OpenAI-compatible local
   endpoint (Ollama, LM Studio, llama.cpp, vLLM)
  • Frontend runs the benchmarks and calls model providers directly — your API keys never leave the browser (localStorage).
  • Backend is a single-file, dependency-free Node server using the built-in node:sqlite module. It stores runs, aggregates the leaderboard and keeps the fan club.

Quickstart

npm install
npm run dev        # API on :8787 + Vite on :5173 (proxies /api)
  1. Open the app → Settings → add a model (or point it at a local server and hit Test connection)
  2. Pick a test under Tests → Run
  3. Watch the leaderboard fill up at /leaderboard

Production:

npm start          # build + serve everything from the Node server

Scripts

Command Description
npm run dev API + web dev servers together
npm run dev:api / npm run dev:web individually
npm run build typecheck + production bundle
npm start build, then serve app + API from one port
npm test vitest suite (scoring, parsing, metrics)
npm run lint oxlint

API

Method Endpoint Description
GET /api/runs all stored runs
POST /api/runs store a run result
GET/DELETE /api/runs/:id fetch / delete one run
DELETE /api/runs wipe the board
GET /api/leaderboard per-model avg/best per test
GET/POST /api/fans the fan club (sanitized, rate-limit-friendly)
GET /api/health liveness + DB path

Database lives at data/cheesebench.db (override with CHEESEBENCH_DB, port with PORT).

Model Support

  • OpenAI, Anthropic, Google, xAI (Grok) — direct from the browser
  • Local / self-hosted — any OpenAI-compatible /v1/chat/completions endpoint (Ollama, LM Studio, llama.cpp, vLLM), with model discovery and connection testing
  • Per-model overrides: API key, max tokens, request concurrency (1–32)

No models ship enabled by default. You bring the models; we bring the questions.

Project Layout

server/index.js      zero-dependency API + static server (node:sqlite)
src/lib/providers.ts provider adapters (OpenAI/Anthropic/Google/Grok/local)
src/lib/tests/       one runner + scorer per benchmark
src/lib/scoring/     Kendall τ-b, alias matching, price error, tool metrics
src/data/            cheeses.json (ground truth), cheeseBoard.json
src/pages/           landing, leaderboard, tests, runs, top1000, games, fans…
scripts/             dev launcher + Wikipedia dataset builder

Disclaimer

This benchmark is scientifically rigorous the way a cheese board is peer-reviewed: tasted by everyone, trusted by no one. Scores are a snapshot of model behavior on curated dairy trivia and should not be used to draw conclusions about anything else. Please do not feed the results to your fondue.


CheeseBench — aged in the open. 🧀

About

A cheesy benchmark

Resources

Stars

0 stars

Watchers

1 watching

Forks

Releases

Packages

Used by

Contributors

Languages