zig_eval is a registry-driven eval tool and library. It supports
deterministic matchers, model-graded checks, single-turn tool-calling evals,
file-backed multimodal evals, OpenAI-compatible chat-completions services, CLI
execution, and text or JSON report output.
services.json
-> eval definition JSON files
-> JSONL datasets
-> zig_eval list/run
-> raw run results
-> text or JSON reports
Services define the products or model endpoints to test. Use
examples/registry/services.json as a starting point.
Each service needs:
name: stable name used in reports and allowlistsbase_url: OpenAI-compatible chat-completions endpointdefault_model: model name sent in the requesttimeout_ms: request timeout value for the service config
api_key_env, provider, and system_prompt are optional. Omit
api_key_env for local or internal endpoints that do not require Bearer-token
authentication.
Services can also define retry behavior:
{
"retry": {
"max_attempts": 3,
"backoff_ms": 500,
"retry_on_status": [429, 500, 502, 503, 504]
}
}Eval definitions live under examples/registry/evals.
Each eval definition points to one dataset and one matcher config. Dataset paths are relative to the registry root when using the CLI or registry-root loaders.
{
"id": "smoke.reply_ok",
"group": "smoke",
"description": "Checks that a service can return a simple literal response.",
"dataset_path": "data/smoke/reply_ok/test.jsonl",
"split": "test",
"matcher": {
"kind": "exact_match",
"case_sensitive": true,
"trim_whitespace": true
},
"default_run_count": 1,
"service_allowlist": ["local-product", "product-staging"]
}service_allowlist is optional. If present, the runner only runs that eval
against the listed services.
Datasets are JSONL files under examples/registry/data. Each line is one eval
case with an input prompt and optional expected value.
{"id":"case-1","input":"Reply with exactly OK.","ideal":"OK"}
{"id":"case-2","input":"Reply with exactly READY.","ideal":"READY"}exact_match and includes use ideal. json_fields checks the service
output for configured root-level JSON fields.
Cases can attach files by registry-relative path. The default service client renders images as OpenAI-compatible image content blocks and appends text-like files as labeled context.
{"id":"case-1","input":"Summarize the attached product changelog.","ideal":"Mentions retry support and parallel execution.","attachments":[{"kind":"file","path":"assets/changelogs/release.md","mime_type":"text/markdown","label":"release notes"}]}
{"id":"case-2","input":"What object is shown?","ideal":"red mug","attachments":[{"kind":"image","path":"assets/images/red_mug.png","mime_type":"image/png","label":"reference image"}]}Built-in rendering supports PNG, JPEG, WebP, and UTF-8 text-like files such as
.txt, .md, .json, .jsonl, .csv, .zig, .py, .js, and .ts.
Other file types require a custom service adapter.
Use model_grade for quality checks where there is no single exact expected
answer. The runner first calls the selected product service, then sends the
original input, candidate output, optional ideal, and rubric to the judge
service.
The judge service is just another OpenAI-compatible service in services.json:
{
"name": "judge",
"base_url": "https://api.openai.com/v1/chat/completions",
"api_key_env": "OPENAI_API_KEY",
"default_model": "gpt-4.1-mini",
"system_prompt": "You are a strict eval judge. Return only the requested JSON.",
"timeout_ms": 30000
}The eval definition selects that judge service and defines the rubric:
{
"id": "quality.helpful_summary",
"group": "quality",
"description": "Grades whether a summary is useful, accurate, and concise.",
"dataset_path": "data/quality/helpful_summary/test.jsonl",
"split": "test",
"matcher": {
"kind": "model_grade",
"judge_service": "judge",
"judge_model": "gpt-4.1-mini",
"rubric": "Grade from 0 to 1. Passing answers are accurate, complete, and concise.",
"pass_score": 0.8
},
"default_run_count": 1,
"service_allowlist": ["local-product", "product-staging"]
}The judge must return JSON only:
{"score":0.9,"passed":true,"reason":"Covers the required points concisely."}Use service_allowlist to keep the eval targeted at product services. The
judge service can still be used for grading even when it is not in the
allowlist.
Use tool_call to validate that a product selects the expected OpenAI-style
tool and sends the expected root-level argument values. The eval definition
provides tool schemas:
{
"id": "tools.search_web",
"group": "tools",
"description": "Checks that the product chooses the search_web tool.",
"dataset_path": "data/tools/search_web/test.jsonl",
"split": "test",
"tools": [
{
"name": "search_web",
"description": "Search the web for current information.",
"parameters_json": "{\"type\":\"object\",\"properties\":{\"query\":{\"type\":\"string\"}},\"required\":[\"query\"]}"
}
],
"matcher": {
"kind": "tool_call"
},
"default_run_count": 1,
"service_allowlist": ["local-product"]
}Dataset cases provide the expected tool calls:
{"id":"case-1","input":"Search the web for the weather in Melbourne.","expected_tool_calls":[{"name":"search_web","arguments_json":"{\"query\":\"weather melbourne\"}"}]}Extra actual tool calls and extra actual argument fields are allowed, but every expected tool call must be present and every expected root-level argument value must match exactly.
List services and evals:
zig build run -- list --registry examples/registryRun evals with the default text report:
zig build run -- run --registry examples/registry --service local-productRun one eval and write aggregate JSON to stdout:
zig build run -- run --registry examples/registry --service local-product --eval smoke.reply_ok --format jsonRun a model-graded eval with a specific judge service:
zig build run -- run --registry examples/registry --service local-product --eval quality.helpful_summary --judge-service judgeRun a tool-calling eval:
zig build run -- run --registry examples/registry --service local-product --eval tools.search_webRun a multimodal file eval:
zig build run -- run --registry examples/registry --service local-product --eval multimodal.release_notesRun with bounded parallelism while limiting concurrent requests per service:
zig build run -- run --registry examples/registry --parallel 4 --max-inflight-per-service 2Supported flags:
--registry PATH: registry root, defaultexamples/registry--service NAME: run only one service--group GROUP: run only one eval group--eval ID: run only one eval id--judge-service NAME: override the judge service formodel_gradeevals--runs N: override each eval'sdefault_run_count--parallel N: number of worker threads, default1--max-inflight-per-service N: max concurrent requests per service, default1--format text|json: report format, defaulttext
run requires the selected service endpoint to be reachable and compatible
with OpenAI-style chat completions.
Text output prints progress lines during parallel runs. JSON output suppresses progress so stdout remains machine-readable.
The CLI is the easiest path, but library users can still call the same modules directly.
const std = @import("std");
const zig_eval = @import("zig_eval");
fn evaluateMatcher(
allocator: std.mem.Allocator,
matcher: zig_eval.matchers.MatcherConfig,
output: []const u8,
ideal: ?[]const u8,
tool_calls: ?[]const zig_eval.services.ToolCall,
expected_tool_calls: ?[]const zig_eval.registry.ExpectedToolCall,
) anyerror!zig_eval.runner.MatcherOutcome {
const outcome = try zig_eval.matchers.evaluate(
allocator,
matcher,
output,
ideal,
tool_calls,
expected_tool_calls,
);
return .{
.passed = outcome.passed,
.score = outcome.score,
.failure_reason = outcome.failure_reason,
};
}
pub fn runProductEvals(allocator: std.mem.Allocator) !void {
var registry_dir = try std.fs.cwd().openDir("examples/registry", .{});
defer registry_dir.close();
var loaded_services = try zig_eval.services.loadServices(
allocator,
registry_dir,
"services.json",
);
defer loaded_services.deinit();
var loaded_evals = try zig_eval.registry.loadRegistryEvalDefinitions(
allocator,
registry_dir,
);
defer loaded_evals.deinit();
var runner_result = try zig_eval.runner.runEvaluations(allocator, .{
.root_dir = registry_dir,
.services = loaded_services.items,
.evals = loaded_evals.items,
.matcher_evaluator = evaluateMatcher,
});
defer runner_result.deinit();
var reports = try zig_eval.reporting.aggregateRunResults(
allocator,
runner_result.runs,
);
defer reports.deinit();
var text_out = std.Io.Writer.Allocating.init(allocator);
defer text_out.deinit();
try zig_eval.reporting.formatEvalReports(&text_out.writer, reports.items);
var comparisons = try zig_eval.reporting.compareServicesToBaseline(
allocator,
reports.items,
"local-product",
);
defer comparisons.deinit();
}- Service calls must target an OpenAI-compatible chat-completions endpoint.
- JSON field matching checks root-level fields only.
- Model-graded evals require one extra judge model call per candidate output.
- Tool-calling evals validate tool selection and arguments only; tool execution and multi-turn tool-result loops are not implemented.
- The default multimodal renderer supports images and text-like files only.
- Streaming and formal p-value significance testing are out of scope for the current implementation.