an open source, self-hosted alternative to typesafe's jev and its system one api. verdict serves the same wire protocol on your own hardware, so an existing jev client changes its base url and nothing else.
a small python server sits in front of your existing llama-server or llama-swap endpoint with any model of your choosing and serves a jev compatible API.
answer typed questions about a piece of text without generating any text.
one forward pass, read the probability the model puts on each declared option label, renormalise over them. the answer is a distribution over option ids, with the confidence signals needed to decide whether to act on it.
the shape of it, one question from the hacker news demo:
question: which of these links opens the story's comments?
options: A "229 comments" B "Hacker News" C "hide" D "past"
answer: a distribution over A-D, plus the option mass that landed on A-D
at all before renormalising
no tokens are sampled, so nothing here is random: there is no seed to set, because the sampler's rng never touches the numbers being read.
that is not the same as bit-identical, and the difference is measured. on a
quiet server the same prompt returns the same probabilities every time -- eight
models reproduced their scores exactly across separate runs. on a server under
load, 5 of 16 prompts came back different, by up to 3.2e-02, because
llama.cpp packs concurrent work into shared batches and matmul reduction order
follows batch shape. spec/SPEC.md section 12 carries the tolerance that
implies, and docs/EVALS.md carries the noise floor it puts on every number
measured here.
any gguf can be pointed at verdict, and these are the ones measured: jevbench's
231 public decisions through verdict's own /v1/systemone, each model read the
way it was trained (read as, see spec/SPEC.md 5.3), with no debiasing.
| model | read as | jevbench public | hard tier | hard ece | p50 latency |
|---|---|---|---|---|---|
| jev-omni:Q8_0 | jev-omni | 204/231 (0.883) | 0.775 | 0.093 | 0.52 s |
| gemma-4-12b-it:Q8_0 | chat | 201/231 (0.870) | 0.766 | 0.194 | 0.60 s |
| jevk5-4b:Q8_0 | semif | 199/231 (0.862) | 0.739 | 0.096 | 1.01 s |
| winnow-12b:Q8_0 | winnow | 198/231 (0.857) | 0.721 | 0.118 | 0.51 s |
| decider-4b:Q8_0 | decider-plain | 192/231 (0.831) | 0.658 | 0.201 | 1.46 s |
| clef-flash:Q8_0 | clef-flash | 190/231 (0.823) | 0.640 | 0.114 | 1.23 s |
| qwen3.5-4b:Q8_0 | chat | 182/231 (0.788) | 0.622 | 0.176 | 1.01 s |
| qwen3.5-9b:Q8_0 | chat | 182/231 (0.788) | 0.631 | 0.132 | 1.29 s |
| gemma-4-e4b-it:Q8_0 | chat | 180/231 (0.779) | 0.586 | 0.327 | 0.44 s |
| standardone-8b:Q8_0 | standardone-native | 179/231 (0.775) | 0.559 | 0.151 | 0.27 s |
| decider-2b:Q8_0 | decider-plain | 176/231 (0.762) | 0.577 | 0.221 | 0.66 s |
| gpt-oss-20b:Q8_0 | chat | 161/231 (0.697) | 0.513 | 0.320 | 0.68 s |
| granite-4.2-3b:Q8_0 | chat | 157/231 (0.680) | 0.451 | 0.446 | 0.42 s |
| gemma-4-e2b-it:Q8_0 | chat | 156/231 (0.675) | 0.396 | 0.536 | 0.29 s |
| qwen3.5-2b:Q8_0 | chat | 149/231 (0.645) | 0.460 | 0.229 | 0.64 s |
| minicpm5-2b:Q8_0 | chat | 144/231 (0.623) | 0.504 | 0.335 | 0.21 s |
| qwen3.5-0.8b:Q8_0 | chat | 131/231 (0.567) | 0.387 | 0.231 | 0.64 s |
public items only, on a strix halo igpu reached over a ~150 ms vpn, so these
are not comparable one-to-one with the jevbench board, which adds sealed items
and runs on datacenter gpus. latency includes that vpn, and rows measured
before verdict pooled its backend connections carry a round trip more than
current code does (docs/EVALS.md section 8). every number, with the weights it was measured on,
is in eval/results/export.json; docs/EVALS.md says what each measurement
does and does not show.
the same runs, all 231 items, best calibrated first.
| model | read as | accuracy | mean confidence | overconfidence | ece | answers at 0.9+ | right when confident | auroc |
|---|---|---|---|---|---|---|---|---|
| jevk5-4b:Q8_0 | semif | 0.862 | 0.875 | +0.013 | 0.030 | 65% | 0.960 | 0.848 |
| jev-omni:Q8_0 | jev-omni | 0.883 | 0.913 | +0.030 | 0.042 | 77% | 0.972 | 0.931 |
| clef-flash:Q8_0 | clef-flash | 0.823 | 0.843 | +0.020 | 0.059 | 61% | 0.979 | 0.885 |
| winnow-12b:Q8_0 | winnow | 0.857 | 0.919 | +0.062 | 0.062 | 81% | 0.947 | 0.916 |
| standardone-8b:Q8_0 | standardone-native | 0.775 | 0.802 | +0.027 | 0.064 | 42% | 0.990 | 0.890 |
| qwen3.5-9b:Q8_0 | chat | 0.788 | 0.853 | +0.065 | 0.065 | 62% | 0.944 | 0.849 |
| qwen3.5-4b:Q8_0 | chat | 0.788 | 0.848 | +0.060 | 0.076 | 56% | 0.930 | 0.805 |
| decider-4b:Q8_0 | decider-plain | 0.831 | 0.903 | +0.072 | 0.085 | 72% | 0.940 | 0.861 |
| qwen3.5-0.8b:Q8_0 | chat | 0.567 | 0.663 | +0.096 | 0.096 | 13% | 0.903 | 0.732 |
| decider-2b:Q8_0 | decider-plain | 0.762 | 0.865 | +0.103 | 0.103 | 61% | 0.907 | 0.831 |
| gemma-4-12b-it:Q8_0 | chat | 0.870 | 0.976 | +0.106 | 0.109 | 94% | 0.903 | 0.804 |
| qwen3.5-2b:Q8_0 | chat | 0.645 | 0.780 | +0.135 | 0.148 | 45% | 0.874 | 0.790 |
| gemma-4-e4b-it:Q8_0 | chat | 0.779 | 0.946 | +0.167 | 0.176 | 84% | 0.830 | 0.833 |
| gpt-oss-20b:Q8_0 | chat | 0.697 | 0.890 | +0.193 | 0.207 | 67% | 0.838 | 0.778 |
| granite-4.2-3b:Q8_0 | chat | 0.680 | 0.929 | +0.250 | 0.255 | 78% | 0.746 | 0.779 |
| minicpm5-2b:Q8_0 | chat | 0.623 | 0.891 | +0.268 | 0.273 | 66% | 0.745 | 0.746 |
| gemma-4-e2b-it:Q8_0 | chat | 0.675 | 0.965 | +0.289 | 0.289 | 89% | 0.717 | 0.769 |
overconfidence is mean confidence minus accuracy and ece is the expected
calibration error over ten bins: both say how far a model's stated confidence
sits from how often it is right, and a temperature fitted on your own data can
repair them. auroc is the chance a right answer carries more confidence than
a wrong one (0.5 says nothing, 1.0 separates them), which nothing fitted
afterwards can repair. the two columns between are what a caller gating at 0.9
gets: how much of the work is answered that confidently, and how often those
answers are right. docs/EVALS.md section 7 reads the table.
| jev (hosted) | verdict | |
|---|---|---|
| where it runs | typesafe's api | your hardware, including a phone |
| the model | theirs | any gguf you already have |
| your data | leaves the machine | does not |
| cost | per call | electricity |
| the wire protocol | POST /v1/systemone |
the same |
| accuracy | not measured here | measured, docs/EVALS.md |
the trade is real and stated plainly: a hosted service is somebody else's
problem to run and tune, and verdict makes you pick a model and live with what
docs/EVALS.md says about it. a 2.5b model answers in 374 ms, completes a
single-screen android goal 3/3 with no false completions, and fails a goal
that needs navigation 3/3.
a model asked to "reply with only the letter" still generates, still drifts, still needs parsing and retries. reading the logits at one position skips all of that. it is one short request instead of a generation, and the result comes with a health signal a generated answer does not have.
option mass is that signal: the raw probability that landed on the label tokens before renormalising. a model can rank options perfectly on an option mass of 1.7e-08, because renormalising a rounding error still ranks it. accuracy cannot see that; option mass can.
A-Za-z labels 52 options. past that the labels come from the model's own
vocabulary -- single-character letters it already has tokens for, thousands of
them -- so a longer list is still read at one position rather than bracketed
into an approximation. it is opt-in, --wide-alphabet.
it is not free, and what it costs depends on the model. measured with list length held fixed and every label unfamiliar, which is the worst case rather than the shipped one:
| ascii labels | all labels unfamiliar | |
|---|---|---|
| qwen3.5-9b | 16/16, mass 0.9996 | 16/16, mass 0.7780 |
| minicpm5-2b | 16/16, mass 0.9995 | 9/16, mass 0.6991 |
| granite-4.2-3b | 15/16, mass 0.9671 | 12/16, mass 0.3754 |
| qwen3.5-2b | 14/16, mass 0.9844 | 12/16, mass 0.3577 |
only the 9b model escapes an accuracy cost. in the configuration actually shipped the pinned 52 come first, so a 104 option list is half familiar and option mass stays at 0.92 to 0.99 on all four.
it also costs latency: an unfamiliar label is not in the top 64 candidates, so the readout has to widen to find it, and the median decision goes from 1033 ms to 2596 ms.
a model that cannot use its own alphabet is refused rather than served. the alphabet is verified the way a formatter is -- by scoring an unambiguous question labelled entirely from it -- and both the option mass and the answer have to hold up. of the four models measured, two are served and two refused, and the two served then pick the right element out of 80.
shortlist if you can; docs/EVALS.md section 2a has the rest, including three
cheaper ways to predict this that were measured and all failed.
spec/ |
the contract: prompt layout, question types, result object, golden fixtures |
python/llama_verdict/ |
reference client and a jev-compatible server |
demos/ |
agent loops driven entirely by typed decisions |
scripts/ |
probes, formatter derivation, profiling |
docs/ |
measured results and the decisions behind them |
spec/ is the source of truth. the definitive implementation will be rewritten
in rust, so the part that has to survive that is the spec and its fixtures, not
this client.
needs a llama-server or llama-swap with a chat model.
make # validate the spec and fixtures
make test # conformance tests, no model needed
make precommit # lint, build, testanything that needs a model reads its coordinates from the environment, and
the python code itself parses a .env -- the nearest one, searched upward
from the working directory, which is the repo root for every make target --
so the live setup is written down once rather than exported per shell:
cp .env.example .env # LLAMA_VERDICT_URL, LLAMA_VERDICT_MODEL, ...
make test-e2e # end to end, live and mock-backed suites
make serve # the jev endpoint, see belowmake test-e2e runs two suites. one drives a live llama-server and skips itself
without LLAMA_VERDICT_URL; the other points the same jev endpoint at
[the mock in the sibling ../fake-openai checkout (--llamacpp), which answers
/completion with a distribution the test programs, and needs no weights. that
second suite is where the cases real weights cannot be told to produce live: mass
landed off the option labels, a backend that answers 503 mid-decision.
.env is ignored and .env.example documents every variable. the same names
work as ordinary environment variables and beat the file, and every command
line flag beats both. nothing about the config depends on make: a bare
python3 -m llama_verdict.server reads the same .env.
neither does the toolchain. python/ is a uv project, so the endpoint and
the suites run with no make and no venv dance:
uv run python -m llama_verdict.server # the jev endpoint, same .env, from python/
uv run pytest # the suites, dev tools and alluv run installs the package editable and, with it, the dev group: pytest,
ruff, the jinja engine the derivation path needs, and the websocket client
the demo agent drives. uv.lock is committed, per NFR3.
a formatter is the small set of affixes that steer a given model to answer with a bare label. it is derived when a model is first used -- nothing to fit in advance and no per-model artifact to ship. the model's own chat template is rendered with sentinel messages and diffed, the result is checked against the option mass floor, and it is cached by the template's sha256 so it happens once per model rather than once per process.
that matters because a pinned table can only cover models someone thought to pin, which fails the case the library exists for: pointing it at an arbitrary gguf on a device.
# no --formatter: the affixes come from the model's own template
PYTHONPATH=python python3 -m llama_verdict.server \
--base-url "$LLAMA_VERDICT_URL" --model gemma-4-e4b-it:Q8_0 --port 8477
# formatter gemma-4-e4b-it:Q8_0 (mean option mass 1.0000, derived ...)first derivation costs about 30-50 s, nearly all of it resolving the 52 label
token ids; a cached template is ~300 ms. jinja2 is imported only on that
path, so a cached or pinned model stays stdlib-only.
spec/formatters/ holds four fitted tables and a synthetic reference.json.
those are golden regression fixtures, not the runtime source: conformance
asserts that deriving a pinned model today reproduces its affixes byte for
byte, so a change in the engine, the derivation or the template shows up as a
diff. make formatters refreshes them. a formatter fitted to one model does
not transfer to another -- measured, a mismatched opening puts about 1e-7 of
the mass on the labels.
some decision models, such as the decider family, are base models fine-tuned
on a plain text layout of their own, while their gguf still carries the base
model's chat template. deriving from that template builds a prompt they never
saw, so they are read with a named layout from spec/layouts/ instead.
which model needs which layout, and the layout's bytes, come from the model
registry, and verdict carries a committed copy of both: spec/models.json
names the readout of every model it recognises and spec/layouts/ holds the
layouts. a model is recognised by the registry key a llama-swap listing names
for it, else by the repository its weights were loaded from, which any
llama-server reports, else by its name, so it is read the right way with no
flag. a model verdict
does not recognise is derived from its own chat template. make layouts
refreshes both copies from the registry (then make fixtures, and review the
diff). a model the registry marks as answering through a head verdict cannot
apply is refused, since reading its label logits would measure its backbone.
--layout overrides the copy for the bound model. see spec/SPEC.md section
5.3.
a head that reads the one position verdict already reads can be applied
instead of refused. --layout jev-omni reads Jev-Omni that way: the model is
served by a llama-server in embedding mode (--embeddings --pooling last, or
--pooling none, with a batch large enough for the whole prompt), verdict
asks it for the final hidden state of the last prompt token and applies the
author's 256-way head to it. the head is fetched once from the revision the
layout pins and held to its checksum; --head uses a file already on disk.
such a readout has no option mass, so results report it as absent with a
no_option_mass flag, and the model is verified by answering a question
rather than by a mass floor. see spec/SPEC.md section 5.4.
clef-flash and clef go one step further and are the models read at more than
one position; clef has the same head design at its own size and is not yet
measured here. --layout clef-flash asks every question of a request in a single
prompt and decides them together through the author's joint head, which reads
the hidden state of every prompt position. it needs the model served with
--embeddings --pooling none and a batch as large as the context, and numpy
(pip install llama-verdict[joint-head]), which nothing else in the package
uses. the head (243 mb) is fetched once and held to its checksum, and the
output embedding rows it reads come by byte range from the author's weights;
--head and --head-rows point at files already on disk. every position's
state is a large reply, about 76 kb a token, so run verdict on the same host
as llama-server for this model: over a slow link a long state takes tens of
seconds. see spec/SPEC.md section 5.5.
verdict is an independent project. it is not affiliated with, endorsed by, or derived from typesafe, and jev and system one are their names, not ours.
typesafe's jev is a hosted decision model, called over an http api whose
decision endpoint is POST /v1/systemone. verdict serves that same wire shape,
so an existing jev client changes its base url and nothing else -- no
client code, no field renaming. see docs/JEV_API.md, which records the public
sources it was written from and the date they were read.
the relationship is one-directional and has three parts, worth separating:
| the protocol | a public api surface we implement. /v1/systemone, the choice / score / noul question types, and the response fields clients validate |
| the mechanism | ours. read the probability mass on declared option labels at one token position and renormalise. nothing about how jev works internally is known to us or claimed here |
| the conformance test | browser-use/jev-ultrafast, an unmodified third-party jev client, pointed at verdict with TYPESAFE_BASE_URL. it working is the only real evidence the wire shape is right |
that last one is why the compatibility matters to us at all: a protocol you implement from documentation is a guess until somebody else's client drives it unchanged.
compatibility is a surface, not a claim of equivalence. verdict is a local
readout over a gguf you already have. it makes no claim to match jev's
accuracy, calibration, latency or behaviour, and docs/EVALS.md measures what
it actually does rather than comparing against a hosted service we cannot
inspect.
the server is optional. the python client and the demos talk to it over the same protocol because that keeps one wire format in the project rather than two, and because the rust core is expected to replace this server rather than grow it.
make serve # reads .env: url, model, port, key
PYTHONPATH=python python3 -m llama_verdict.server \
--base-url "$LLAMA_VERDICT_URL" --model "$LLAMA_VERDICT_MODEL" --port 8477the two are the same call: every setting has an environment variable and a
flag, and the flag wins. VERDICT_API_KEY makes callers present
authorization: bearer <key> and leaves the endpoint open when unset.
the backend can also be named the way any openai-compatible client names it:
OPENAI_BASE_URL (or the older OPENAI_API_BASE) and OPENAI_MODEL.
verdict's own LLAMA_VERDICT_URL and LLAMA_VERDICT_MODEL win wherever they
are set, because OPENAI_BASE_URL is often exported for some other tool. a
backend that wants a bearer key gets LLAMA_VERDICT_API_KEY; OPENAI_API_KEY
is sent only to the backend the OPENAI_* url names, never to one named by
verdict's own setting or a flag, since it is usually a real key for somebody
else's service.
one endpoint serves every model its backend has. a request's model names
the backend model that answers it, and the endpoint derives and verifies that
model the first time it is asked for, which can take a minute. the jev aliases
(jev-latest, jev-preview) and a request that names no model are answered by
--model, which is optional: an endpoint started without one answers for
whichever model a request names and refuses a request that names none. so a
sweep over models runs against one endpoint that stays up:
VERDICT_ENDPOINT=http://host:8477 scripts/sweep_jevbench.sh OUT MODEL...
and scripts/sweep_models.py --endpoint http://host:8477 .... GET /v1/banner?model=<id> says how a model was read -- layout, weights, build --
which is what a result is evidence beside. --no-routing answers everything
with --model.
the demos below all expect it on 127.0.0.1:8477. the remaining flags describe
the bound model only: --formatter pins a table instead of deriving one;
--assistant-open spells out an assistant opening for a format whose
generation prompt ends before content begins (the ones verdict knows, gpt-oss
among them, need no flag); --layout reads it with a named layout instead of
the one verdict's copy of the registry names; --head points a layout that is
read through a decision head at a head file on disk, and --head-rows a joint
head at the embedding rows it reads.
agent loops where every decision is a typed question and nothing is generated.
# hacker news, our own cdp driver
./demos/browser_agent.py \
--goal "Read the top 5 comments on each of the top 3 stories on Hacker News." \
--require '[0-9]+\s*comments' \
--collect '^\s*[a-z0-9_-]{2,15} [0-9]+ (?:minute|hour|day)s? ago \|[^\n]*\n+[^\n]+' \
--collect-pages 3
# the same task through an unmodified browser-use/jev-ultrafast client
./demos/run_hn_demo.py --target-first --shortlist 26
# android, over mimic's accessibility surface
./demos/mimic_agent.py --goal "Open About phone and find the Android version." \
--require "Android version" --launch--require is how a run is SCORED: each pattern must actually appear on a
screen the agent reached. the agent's own DONE is an opinion, and one was
measured at 0.32 with a third of the goal outstanding.
that opinion does carry signal, though. on the device, across seven models,
true completions measure 0.81-0.98 and false ones 0.45-0.77, so the android
agent holds DONE, and BLOCKED with it, to --done-confidence 0.81 by default.
it is a default and not a constant because the gap is narrow and confidence
does not mean the same thing on every model: one true DONE was measured at
0.66 and cost that run its ending. --done-confidence 0 turns the gate off.
the browser agent leaves it off.
it holds a separate threshold from --min-confidence on purpose -- "is this
the right target" and "is the task finished" are different questions.
both agents ask one question a step: its options are every row or link, every
app and every move, and whether to stop is asked beside it in the same
request. the older form asked which target and then which operation, in two
calls, and is still there as --no-together. the one question is half the
calls. on android it was ahead on all three models it was run on, 6 runs of 9
against 2, and on the browser it halved the two head models' decision time
and lost nothing: docs/EVALS.md sections 5c and 6. every agent table before
those sections was taken with two questions.
on the browser no confidence separates a true DONE from a false one, so the
task says how many pages it needs: --collect-pages 3 does not offer DONE
until that many have been collected from. with it all three models measured
complete the task in every run.
scripts/sweep_models.py writes every step of every run to
/tmp/sweep-steps-<model>.log as it happens, and demos/reliability.py --log
does the same for one model. watch that file: most of what section 5c fixed
was visible in it within a minute and in the summaries not at all.
--collect is the task's OUTPUT. verdict decides where to look and never
produces the answer text, so whatever the agent navigated to is read off the
page rather than generated. scripts/page_text.py URL dumps exactly what the
agent sees, which is how to write one of these patterns.
docs/EVALS.md is the results document -- every measurement, the
instrument that produced it, and what it does not show. the cheap instruments
come first there for a reason: a benchmark ranks a model in a minute with no
browser and no device, and an agent run against a live site measures the model,
the harness, the network and a page that moves underneath it.
scripts/smoke_models.py --models qwen3.5-4b:Q8_0 minicpm5-2b:Q8_0docs/DEMOS.md carries the agent-loop narrative and the interventions that did
nothing.
GPL-3.0-or-later. the full text is in LICENSE, verbatim from the fsf.
worth knowing before building on it, because it is a strong copyleft and this
project is heading for a library: REQUIREMENTS.md FR7 describes a flutter
library loading a gguf through llama.cpp via ffi, and the plan's phases 5 and 6
build a dart core and an ffi plugin. under the GPL an application that links
that library must itself be GPL-compatible, which apache-2.0 -- the licence the
plan originally defaulted to -- would not have required. that is a deliberate
choice by the product owner and not an oversight; flagged here so nobody
discovers it at integration time.
THIRD_PARTY.md and the dependency licence check are not written yet. the
python client is stdlib only except jinja2, which is used on the derivation
path alone and is BSD-3-Clause.
calibrated probabilities, out of the box. gemma-4-e2b was measured reporting 1.000 confidence on wrong answers. gating on raw confidence is not safe until a calibration is fitted on your own data, and the docs say which model sizes are fit for which kinds of question rather than implying all of them are.