Repository navigation
feat(kolibri1): serve the chat template and the kolibri1 parsers end to end - #3422
Conversation
The kolibri1 CPU arm reaches the model forward but not serving. This issue scopes the serving-completion unit: the kolibri1 reasoning parser (the plugin's Qwen3 grammar with the template's thinking switch), the kolibri1 tool-parser alias to the Hermes format, the template detection rows, and a jinja2-referenced rendering gate for the tokenizer-shipped chat template. Oracle: aleph-alpha-inference pin 049a6a7bd240. FOLLOWING_AGENTS_PROTOCOL Following-Agents-Protocol: true AI-Assisted: true Assisted-by: AGENT:zai/glm-5.3-flash [maki]
…to end The kolibri1 CPU arm reached the model forward but not serving: no kolibri1 reasoning parser, no kolibri1 tool-parser name, and detection resolved think_auto/hermes off the Kolibri template. The model-author plugin (aleph-alpha-inference @ 049a6a7bd240) is the serving oracle: its reasoning parser is the Qwen3 grammar with the starting state derived the way the template switches thinking (reasoning_effort wins and only "none" disables; else a literal enable_thinking false does), and its tool parser IS the Hermes class under the kolibri1 name. - kolibri1 reasoning adapter (reasoning.py:36-61 ported): per-request state from chat_template_kwargs, over the shared Qwen3 engine. - kolibri1 tool parser: alias to HermesToolParser (init.py:50-54). - Detection rows on the template's no-reasoning sentence, ahead of the generic <think> and hermes rows; registry pins 12->13 and 42->43. - Renderer parity: minja's tojson dumped insertion order; jinja2's default policy sorts keys. Fixed adapter-side with a child-scope sorted-dump tojson; no vendor change. - Gate: byte-exact rendering vs CPython jinja2 references on 15 scenarios (fixtures/kolibri1_chat_template_references.json), parser unit tests on the plugin's switch truth table, alias equivalence with hermes. Red-first: detection and registry cases failed before the change. Evidence: docs/bench-evidence/kolibri1-serve-20261008.md. FOLLOWING_AGENTS_PROTOCOL Following-Agents-Protocol: true AI-Assisted: true Assisted-by: AGENT:zai/glm-5.3-flash [maki]
The serving unit landed in the previous commit; this records the dated resolution on the canonical issue and refreshes the derived index. FOLLOWING_AGENTS_PROTOCOL Following-Agents-Protocol: true AI-Assisted: true Assisted-by: AGENT:zai/glm-5.3-flash [maki]
ef97514 to
4aade35
Compare
localai-org-maint-bot
left a comment
There was a problem hiding this comment.
@mudler @richiejp Two blockers from source review of 4aade352b1b052226c5035423c1cafc38f90475a:
-
P1: the global
tojsonoverride uses the wrong serving reference.src/vllm/entrypoints/chat_template.cppnow sorts keys for every model. The pinned Transformers 5.14.1 renderer overrides Jinja's filter withtojson(..., sort_keys=False)and installs that filter in_compile_jinja_template: https://github.com/huggingface/transformers/blob/v5.14.1/src/transformers/utils/chat_template_utils.py#L481 . Thus plain Jinja's default is not the Transformers serving behavior. For an insertion-ordered object{"z":1,"a":2}, the reference preservesz,a; this override forcesa,z, changing prompt bytes and tokens across models. The new fixture generator uses plainjinja2.Environment, so it certifies the same incorrect reference. Regenerate through the pinned Transformers renderer and preserve its default/options before changing this shared seam; include unsorted nested tool schemas and Unicode/HTML characters in the comparison. -
P2: the registered template test depends on a private absolute model path.
tests/vllm/entrypoints/test_kolibri1_chat_template.cpploads/mnt/models/Aleph-Alpha/Kolibri-1/tokenizer_config.jsonin both rendering and detection cases. The fixture includes expected strings but not that input template.LoadChatTemplateFromConfigthrows when the file is absent, so a clean checkout cannot run this test using the committed fixtures. Commit the pinned template input and load it from the fixture directory (or explicitly provision a pinned artifact); the generator's hard-coded output under/tmp/vllm-kolibri-servealso needs a portable destination.
Validation: inspected the complete changed parser/renderer/test code and the pinned Transformers filter; traced the missing-file exception in the current base. Native tests and mutation checks were not run because this environment lacks Python/CMake/a C++ compiler. This is a findings review, not an acceptance or performance gate.
…ma order The PR localai-org#3422 review falsified the global tojson override that sorted object keys for every model. Measured first, against the pinned transformers 5.14.1 renderer (.agents/oracles/transformers.md) through render_jinja_template with a probe over an insertion-ordered {"z": 1, "a": 2}, nested unsorted tool schemas and Unicode/HTML leaves: the pinned renderer's tojson override (utils/chat_template_utils.py:481) is json.dumps with sort_keys=False, ensure_ascii=False and NO HTML escaping, so the serving reference keeps insertion key order and raw UTF-8 -- plain Jinja's builtin (sorted, HTML-escaped) is not the serving behavior, and the old sorted-dump override rendered prompt bytes no serving reference produces. chat_template.cpp replaces the sorted-dump override with a port of the pinned filter's FULL signature (ensure_ascii/indent/separators/ sort_keys, CPython json.dumps semantics; minja's builtin matches the default but rejects three of the four options). BuildTools now mirrors the measured pinned vLLM tool shape ([tool.model_dump() ...] at online_renderer.py:178: fixed field order, description/parameters present as null when absent). FunctionDefinition::parameters becomes nlohmann::ordered_json (nlohmann::json sorts object keys at parse, losing the request document's order that the pinned renderer preserves into the prompt), with an ordered from_json overload and RestoreToolSchemaOrder, which re-reads tools from an order-preserving body parse. Every chat entry point calls it after its regular parse: api_server handle_chat_ completions, the C ABI ParseChatRequest (vllm_chat and vllm_chat_stream), and run_batch. Key-lookup readers (the tool parsers) are behavior-identical; step3.cpp's one pointer type follows the field. Gates: full tree builds (881 targets); test_kolibri1_chat_template 61, test_chat_template 204, test_reasoning_kolibri1 43, test_tool_parser_kolibri1 19, test_reasoning_parser_detect 75, test_tool_parser_detect 361, test_reasoning_qwen3 164, test_openai_tool_parsers 64, test_kolibri1 27/234, test_kolibri1_decode_bench (anchor 109726), test_openai_serving 1365, test_openai_api_server 1517, test_openai_conformance 252, test_parser_engine_assembly 5038, test_openai_run_batch 16, and the parameters-reading tool-parser suites all green; check-tree-compiles 161/161 in scope; check-agent-record OK. FOLLOWING_AGENTS_PROTOCOL Following-Agents-Protocol: true AI-Assisted: true Assisted-by: AGENT:mistral/mistral-large-4 [maki]
…inned renderer Review blockers P1/P2 on the fixture side. The old generator rendered with plain jinja2 (sorted keys, HTML-escaped -- not the serving reference) and hard-coded its output under /tmp/vllm-kolibri-serve; the test loaded /mnt/models/Aleph-Alpha/Kolibri-1/ tokenizer_config.json in both the rendering and detection cases. The generator now drives the PINNED transformers 5.14.1 renderer (render_jinja_template, the function apply_chat_template delegates to; it asserts transformers==5.14.1 and records the pin, jinja2's version, the template sha256 and the capture path in the provenance), reads the template from the committed fixture config, mirrors pinned vLLM's _postprocess_messages (assistant tool-call arguments parsed to dicts) and model_dump tool shape, and writes in-tree next to itself. The reference fixture grows from 15 to 20 scenarios: the 15 original plus with_tools_unsorted_unicode and with_tools_unsorted_unicode_thinking_off (unsorted nested tool schemas + Unicode/HTML), with_tools_minimal_no_ description (pins "description": null), assistant_tool_call_unsorted_ arguments (the template's second tojson site, tool_call.arguments | tojson, unsorted Unicode/HTML arguments), and unicode_html_user_message. The template input is committed at tests/fixtures/kolibri1-chat- template-tokenizer_config.json (the checkpoint's chat_template, sha256 9ba35d4bd6baa26b66aa75d03a922dfee98b16bb1fa37481b195d247267b0f97 in the fixture's fixture_provenance). The test loads the template and the references from KOLIBRI1_TEMPLATE_FIXTURE_DIR (tests/fixtures at compile time); no /mnt path remains in any load path -- proven by running the test from / with both fixtures copied to /tmp/p2-proof: 61/61 PASS with the model directory absent. Red-first under the corrected reference (unmodified adapter, 20-case fixture): 6 of 20 rendering scenarios failed -- with_tools, with_tools_thinking_off, with_tools_unsorted_unicode, with_tools_unsorted_unicode_thinking_off, with_tools_minimal_no_ description, assistant_tool_call_unsorted_arguments (34/40 assertions). After the adapter fix: 61/61 PASS. FOLLOWING_AGENTS_PROTOCOL Following-Agents-Protocol: true AI-Assisted: true Assisted-by: AGENT:mistral/mistral-large-4 [maki]
…plates
The tojson fix touches a seam every model's rendering depends on, so
the non-kolibri behavior is pinned explicitly (+8 assertions, 204
total): a non-kolibri tool template (the Hermes/Qwen-style tool branch
AND the real Qwen3.5 fixture template) renders an insertion-ordered,
Unicode/HTML tool byte-identical to the pinned renderer; the four
tojson options (indent=2, sort_keys=True, ensure_ascii=True,
separators=(',', ':')) are pinned byte-for-byte against the pin; and
an already-alphabetical schema renders byte-identical before/after the
repair, so the change alters only what the pinned renderer actually
differs on. Expected bytes were generated by the pinned transformers
5.14.1 renderer itself (render_jinja_template, 2026-10-09).
FOLLOWING_AGENTS_PROTOCOL
Following-Agents-Protocol: true
AI-Assisted: true
Assisted-by: AGENT:mistral/mistral-large-4 [maki]
…ssue, spec, anchors Records the repair of both review blockers on the serving-completion issue ISSUE-LOCAL-01M4EF3R0H2H3FN5NA0NAB0S62 (the tracker forbids reopening via update, so the repair is recorded in the Resolution): docs/bench-evidence/kolibri1-serve-20261008.md gains a "Review repair" section -- what the pinned transformers 5.14.1 renderer actually does (MEASURED: insertion order, raw UTF-8, no HTML escaping, plus the four tojson options and the model_dump tool shape), what the old sorted-dump override did, the red-first evidence (6 of 20 scenarios failed under the corrected reference), the fixture regeneration, the portability fix and the gate numbers -- and marks the two sections that recorded the falsified decision SUPERSEDED. The spec's SERVING COMPLETION paragraph in .agents/specs/kolibri-1-cpu.md is corrected the same way. Three engine-matrix.md api_server.cpp anchors shifted by the +5-line RestoreToolSchemaOrder call are repaired (1354->1359, 1365->1370, 1614->1619); check-agent-record reports ANCHOR-ROT=0. FOLLOWING_AGENTS_PROTOCOL Following-Agents-Protocol: true AI-Assisted: true Assisted-by: AGENT:mistral/mistral-large-4 [maki]
The review-repair commits are pushed; the PR body gained the "Review repair" section after that push, and a body edit alone does not rerun the guard jobs. This empty commit retriggers CI so the trailer/style guards read the body that will become the squash commit message. FOLLOWING_AGENTS_PROTOCOL Following-Agents-Protocol: true AI-Assisted: true Assisted-by: AGENT:mistral/mistral-large-4 [maki]
…oreToolSchemaOrder The PR localai-org#3422 repair wired RestoreToolSchemaOrder into RunBatch::DispatchChat as RestoreToolSchemaOrder(body, request), but the seam's first parameter is the body TEXT (const std::string&, protocol.h:605) and body is the parsed nlohmann::json object. The call compiles only through nlohmann's implicit get<std::string>() conversion, which throws json.exception.type_error.302 on every object body, and DispatchChat has no try/catch around it — so every object-body /v1/chat/completions line through RunBatch throws. Clean-head evidence at d93d9d4: test_openai_run_batch 3 of 7 cases throw (test_run_batch.cpp:451/:507/:601); the repair's recorded 'run_batch 16/16' gate counted assertions and missed the throwing cases (stale binary), so the PR's CI lane would be red. ISSUE-LOCAL-01M4FR20MES4HQVBRWJBJ2AVCN. A first in-flow attempt (body.dump()) is wrong and was reverted: body is the key-sorted nlohmann::json, so its dump is already sorted and the order restoration reads sorted text — the crash goes away but the seam no-ops and the request document's tool-schema key order is lost on the batch path. The original request text IS available: RunLine parses request_json and drops it. RunLine now re-serializes the chat body (the line's top-level "body" member) from an order-preserving nlohmann::ordered_json parse of the original line text and threads it through DispatchChat(custom_id, body, body_json); DispatchChat calls RestoreToolSchemaOrder(body_json, request). ordered_json::dump keeps insertion order, so the re-read restores the request document's schema key order into the prompt, matching the api_server and C-ABI entry points, which already pass their raw body strings. The shared seam's signature is unchanged. Red-first: at the committed throwing call, 4 of 8 test_openai_run_batch cases throw type_error.302 (the 3 pre-existing plus the new F1 case) while assertions read 16/16; under the body.dump() mutation the 3 old cases pass but the new F1 case fails both key-order assertions (seam no-op); at this fix 8/8 cases, 89/89 assertions pass. FOLLOWING_AGENTS_PROTOCOL Following-Agents-Protocol: true AI-Assisted: true Assisted-by: AGENT:mistral/mistral-large-4 [maki]
…nd api_server entry points Scoped re-review F1 of PR localai-org#3422 (ISSUE-LOCAL-01M4FR1D7RH7CN5X2R2CAR6N61): RestoreToolSchemaOrder was detected by NO committed test — a no-op passed every gate, and the batch call site was not merely unguarded but crashing (ISSUE-LOCAL-01M4FR20MES4HQVBRWJBJ2AVCN). Two entry-point tests now close the gap, one per affected dispatch: - tests/vllm/entrypoints/openai/test_run_batch.cpp: a new case drives RunBatch::RunLine (the batch entry point) with a RAW JSONL line whose nested chat body is non-alphabetical at every level (tool wrapper, function, and parameters schema with zeta before alpha). The raw string is load-bearing: BatchLine() re-serializes through the key-sorting nlohmann::json and would destroy the order under test. The case asserts (a) no exception and a 200 row — pinning the crash fix — and (b) the prompt seam receives the schema in the request document's key order and not the key-sorted form, captured via a CapturingToolsPrompt seam over the parsed tools (the capture pattern the api_server harness uses). Mutation-proven in both directions: with the committed throwing call restored the case throws type_error.302; with the call's argument mutated to body.dump() it fails both key-order assertions (the sorted dump no-ops the seam). - tests/vllm/entrypoints/openai/test_api_server.cpp: a new case posts an unsorted-schema tools request through the PRODUCTION /v1/chat/completions dispatch (handle_chat_completions) with the real Qwen3.8 fixture template and asserts the rendered prompt bytes carry the document order. Mutation-proven: with RestoreToolSchemaOrder no-op'd at api_server.cpp:400 the case fails both key-order assertions. Gates at this head: test_openai_run_batch 8/8 cases / 89/89 assertions; test_openai_api_server 104/104 cases / 1521/1521 assertions (1517 pre-existing + 4 new). No existing assertion weakened: both additions are new TEST_CASEs; the run_batch file's pre-existing cases are untouched. FOLLOWING_AGENTS_PROTOCOL Following-Agents-Protocol: true AI-Assisted: true Assisted-by: AGENT:mistral/mistral-large-4 [maki]
…ec, evidence The scoped re-review of PR localai-org#3422 (ISSUE-LOCAL-01M4FR1D7RH7CN5X2R2CAR6N61) filed two findings; the F1 half landed with the F1 tests in the previous commit. This commit lands F2 and the records the repair invalidates. F2 (mislabeled guard case): the 'byte-stable for already-ordered schemas' case in tests/vllm/entrypoints/test_chat_template.cpp used an AlphabeticalTool fixture that was NOT fully alphabetical (the tool wrapper is insertion-ordered non-alphabetically at type/function and name/description/parameters, and the schema was type/properties/required), so under a sorted tojson mutation the case went red too and the 'alphabetical-schema byte-identity before/after' claim was not demonstrated. The fixture is now genuinely alphabetical at every level (schema properties < required < type; single-property schema needs no order) and the case renders the schema alone ({{ tools[0].function.parameters | tojson }}) because the tool wrapper's field order is the pinned model_dump order, which is NOT alphabetical — so the byte-identity claim is now real and the comment says so. Mutation-proven: with the tojson default mutated to sorted keys (the override the review repair removed) the insertion-order case goes RED while this byte-stable case stays GREEN. No existing assertion was weakened: the full-template exact-bytes rendering remains pinned by the insertion-order case (UnsortedUnicodeTool) and the WeatherTool cases; test_chat_template 45/45 cases, 204/204 assertions. Records: ISSUE-LOCAL-01M4FR20MES4HQVBRWJBJ2AVCN (the confirmed run_batch crash, rewritten — the first in-flow draft prescribed body.dump(), which silences the crash but no-ops the order restoration because the dump is already key-sorted; the landed fix threads the original body text through) and ISSUE-LOCAL-01M4FR1D7RH7CN5X2R2CAR6N61 (F1+F2) are closed with dated resolution evidence. The row spec's ## Now records the re-review repair and its gates; docs/bench-evidence/kolibri1-serve-20261008.md gains a 'Re-review repair' section that corrects the falsified 'test_openai_run_batch 16' gate line (a stale binary: 3 of 7 cases threw type_error.302 at the repair head while assertions read 16/16). Gates at this head: test_openai_run_batch 8/8 cases / 89/89 assertions, test_openai_api_server 104/104 / 1521/1521, test_openai_serving 48/48 / 1365/1365, test_chat_template 45/45 / 204/204, test_kolibri1_chat_template 4/4 / 61/61, test_reasoning_kolibri1 7/7 / 43/43, test_tool_parser_kolibri1 3/3 / 19/19, test_reasoning_parser_detect 8/8 / 75/75, test_tool_parser_detect 16/16 / 361/361, test_kolibri1 27/27 / 234/234, test_kolibri1_decode_bench 1/1 / 2/2 (anchor chain ends 109726) — all PASS; W3 not rerun (no forward change). check-agent-record OK (ANCHOR-ROT=0); agent-issue-index refreshed; agent-preflight.sh --staged exits 0 (five argument-requiring checkers SKIP in this environment, unchanged by this diff). FOLLOWING_AGENTS_PROTOCOL Following-Agents-Protocol: true AI-Assisted: true Assisted-by: AGENT:mistral/mistral-large-4 [maki]
The re-review repair commits are pushed; the PR body gained the "Re-review repair" section after that push, and a body edit alone does not rerun the guard jobs. This empty commit retriggers CI so the trailer/style guards read the body that will become the squash commit message. FOLLOWING_AGENTS_PROTOCOL Following-Agents-Protocol: true AI-Assisted: true Assisted-by: AGENT:mistral/mistral-large-4 [maki]
…ma order The PR #3422 review falsified the global tojson override that sorted object keys for every model. Measured first, against the pinned transformers 5.14.1 renderer (.agents/oracles/transformers.md) through render_jinja_template with a probe over an insertion-ordered {"z": 1, "a": 2}, nested unsorted tool schemas and Unicode/HTML leaves: the pinned renderer's tojson override (utils/chat_template_utils.py:481) is json.dumps with sort_keys=False, ensure_ascii=False and NO HTML escaping, so the serving reference keeps insertion key order and raw UTF-8 -- plain Jinja's builtin (sorted, HTML-escaped) is not the serving behavior, and the old sorted-dump override rendered prompt bytes no serving reference produces. chat_template.cpp replaces the sorted-dump override with a port of the pinned filter's FULL signature (ensure_ascii/indent/separators/ sort_keys, CPython json.dumps semantics; minja's builtin matches the default but rejects three of the four options). BuildTools now mirrors the measured pinned vLLM tool shape ([tool.model_dump() ...] at online_renderer.py:178: fixed field order, description/parameters present as null when absent). FunctionDefinition::parameters becomes nlohmann::ordered_json (nlohmann::json sorts object keys at parse, losing the request document's order that the pinned renderer preserves into the prompt), with an ordered from_json overload and RestoreToolSchemaOrder, which re-reads tools from an order-preserving body parse. Every chat entry point calls it after its regular parse: api_server handle_chat_ completions, the C ABI ParseChatRequest (vllm_chat and vllm_chat_stream), and run_batch. Key-lookup readers (the tool parsers) are behavior-identical; step3.cpp's one pointer type follows the field. Gates: full tree builds (881 targets); test_kolibri1_chat_template 61, test_chat_template 204, test_reasoning_kolibri1 43, test_tool_parser_kolibri1 19, test_reasoning_parser_detect 75, test_tool_parser_detect 361, test_reasoning_qwen3 164, test_openai_tool_parsers 64, test_kolibri1 27/234, test_kolibri1_decode_bench (anchor 109726), test_openai_serving 1365, test_openai_api_server 1517, test_openai_conformance 252, test_parser_engine_assembly 5038, test_openai_run_batch 16, and the parameters-reading tool-parser suites all green; check-tree-compiles 161/161 in scope; check-agent-record OK. FOLLOWING_AGENTS_PROTOCOL Following-Agents-Protocol: true AI-Assisted: true Assisted-by: AGENT:mistral/mistral-large-4 [maki]
… anchors Records the repair of both review blockers on the serving-completion issue ISSUE-LOCAL-01M4EF3R0H2H3FN5NA0NAB0S62 (the tracker forbids reopening via update, so the repair is recorded in the Resolution): docs/bench-evidence/kolibri1-serve-20261008.md gains a "Review repair" section -- what the pinned transformers 5.14.1 renderer actually does (MEASURED: insertion order, raw UTF-8, no HTML escaping, plus the four tojson options and the model_dump tool shape), what the old sorted-dump override did, the red-first evidence (6 of 20 scenarios failed under the corrected reference), the fixture regeneration, the portability fix and the gate numbers -- and marks the two sections that recorded the falsified decision SUPERSEDED. The spec's SERVING COMPLETION paragraph in .agents/specs/kolibri-1-cpu.md is corrected the same way. Three engine-matrix.md api_server.cpp anchors shifted by the +5-line RestoreToolSchemaOrder call are repaired (1354->1359, 1365->1370, 1614->1619); check-agent-record reports ANCHOR-ROT=0. FOLLOWING_AGENTS_PROTOCOL Following-Agents-Protocol: true AI-Assisted: true Assisted-by: AGENT:mistral/mistral-large-4 [maki]
…oreToolSchemaOrder The PR #3422 repair wired RestoreToolSchemaOrder into RunBatch::DispatchChat as RestoreToolSchemaOrder(body, request), but the seam's first parameter is the body TEXT (const std::string&, protocol.h:605) and body is the parsed nlohmann::json object. The call compiles only through nlohmann's implicit get<std::string>() conversion, which throws json.exception.type_error.302 on every object body, and DispatchChat has no try/catch around it — so every object-body /v1/chat/completions line through RunBatch throws. Clean-head evidence at d93d9d4: test_openai_run_batch 3 of 7 cases throw (test_run_batch.cpp:451/:507/:601); the repair's recorded 'run_batch 16/16' gate counted assertions and missed the throwing cases (stale binary), so the PR's CI lane would be red. ISSUE-LOCAL-01M4FR20MES4HQVBRWJBJ2AVCN. A first in-flow attempt (body.dump()) is wrong and was reverted: body is the key-sorted nlohmann::json, so its dump is already sorted and the order restoration reads sorted text — the crash goes away but the seam no-ops and the request document's tool-schema key order is lost on the batch path. The original request text IS available: RunLine parses request_json and drops it. RunLine now re-serializes the chat body (the line's top-level "body" member) from an order-preserving nlohmann::ordered_json parse of the original line text and threads it through DispatchChat(custom_id, body, body_json); DispatchChat calls RestoreToolSchemaOrder(body_json, request). ordered_json::dump keeps insertion order, so the re-read restores the request document's schema key order into the prompt, matching the api_server and C-ABI entry points, which already pass their raw body strings. The shared seam's signature is unchanged. Red-first: at the committed throwing call, 4 of 8 test_openai_run_batch cases throw type_error.302 (the 3 pre-existing plus the new F1 case) while assertions read 16/16; under the body.dump() mutation the 3 old cases pass but the new F1 case fails both key-order assertions (seam no-op); at this fix 8/8 cases, 89/89 assertions pass. FOLLOWING_AGENTS_PROTOCOL Following-Agents-Protocol: true AI-Assisted: true Assisted-by: AGENT:mistral/mistral-large-4 [maki]
…nd api_server entry points Scoped re-review F1 of PR #3422 (ISSUE-LOCAL-01M4FR1D7RH7CN5X2R2CAR6N61): RestoreToolSchemaOrder was detected by NO committed test — a no-op passed every gate, and the batch call site was not merely unguarded but crashing (ISSUE-LOCAL-01M4FR20MES4HQVBRWJBJ2AVCN). Two entry-point tests now close the gap, one per affected dispatch: - tests/vllm/entrypoints/openai/test_run_batch.cpp: a new case drives RunBatch::RunLine (the batch entry point) with a RAW JSONL line whose nested chat body is non-alphabetical at every level (tool wrapper, function, and parameters schema with zeta before alpha). The raw string is load-bearing: BatchLine() re-serializes through the key-sorting nlohmann::json and would destroy the order under test. The case asserts (a) no exception and a 200 row — pinning the crash fix — and (b) the prompt seam receives the schema in the request document's key order and not the key-sorted form, captured via a CapturingToolsPrompt seam over the parsed tools (the capture pattern the api_server harness uses). Mutation-proven in both directions: with the committed throwing call restored the case throws type_error.302; with the call's argument mutated to body.dump() it fails both key-order assertions (the sorted dump no-ops the seam). - tests/vllm/entrypoints/openai/test_api_server.cpp: a new case posts an unsorted-schema tools request through the PRODUCTION /v1/chat/completions dispatch (handle_chat_completions) with the real Qwen3.8 fixture template and asserts the rendered prompt bytes carry the document order. Mutation-proven: with RestoreToolSchemaOrder no-op'd at api_server.cpp:400 the case fails both key-order assertions. Gates at this head: test_openai_run_batch 8/8 cases / 89/89 assertions; test_openai_api_server 104/104 cases / 1521/1521 assertions (1517 pre-existing + 4 new). No existing assertion weakened: both additions are new TEST_CASEs; the run_batch file's pre-existing cases are untouched. FOLLOWING_AGENTS_PROTOCOL Following-Agents-Protocol: true AI-Assisted: true Assisted-by: AGENT:mistral/mistral-large-4 [maki]
What
Serving completion for the kolibri1 CPU arm: the OpenAI chat path now renders
the checkpoint's chat template and parses kolibri1 reasoning and tool calls.
Review-repaired (see "Review repair" below) after the PR review blocked the
first landing on two findings against the shared
tojsonseam and thetemplate test's model-path dependency.
kolibri1reasoning parser (src/vllm/entrypoints/openai/reasoning_parsers/kolibri1.{h,cpp}): the plugin'sreasoning.py:36-61ported — the Qwen3 engine grammar with the starting state derived per request fromchat_template_kwargs(reasoning_effortwins and only"none"disables; else a literalenable_thinking: falsedoes).kolibri1tool parser: an alias toHermesToolParser, exactly the plugin's registration (__init__.py:50-54).<think>/ hermes rows; registry count pins 12→13 (reasoning) and 42→43 (tool).src/vllm/entrypoints/chat_template.cpp:tojsonis a port of the pinned transformers 5.14.1 renderer's filter (utils/chat_template_utils.py:481:sort_keys=False,ensure_ascii=False, no HTML escaping) — insertion key order, raw UTF-8 — replacing the first landing's sorted-dump override, which the review falsified (see "Review repair").FunctionDefinition::parametersisnlohmann::ordered_json, and every chat entry point (api_server, the C ABI,run_batch) re-readstoolsfrom an order-preserving body parse (RestoreToolSchemaOrder).test_reasoning_kolibri1(43),test_tool_parser_kolibri1(19),test_kolibri1_chat_template(61: byte-exact vs the pinned transformers renderer on 20 scenarios,tests/fixtures/kolibri1_chat_template_references.json),test_chat_template(204: 196 pre-existing + 8 non-kolibri tojson guards).Why
Upstream vLLM implements nothing for
Kolibri1ForCausalLM; the model-authorplugin (
aleph-alpha-inference@049a6a7bd240, the row's serving oracle per.agents/oracles/aleph-alpha-inference.md) is where the behavior is defined,and its serving recipe (
--reasoning-parser kolibri1 --tool-call-parser kolibri1) could not be served from this tree. ClosesISSUE-LOCAL-01M4EF3R0H2H3FN5NA0NAB0S62; spec
## Nowupdated in the samebranch.
Review repair (2026-10-09, review blockers P1/P2)
The review (localai-org-maint-bot) blocked the first landing with two findings.
P1 — the global tojson override used the wrong serving reference. The
first landing sorted object keys for every model, because its references were
captured with plain jinja2. Measured against the PINNED transformers 5.14.1
renderer (
.agents/oracles/transformers.md, the version the pinned vLLMenvironment resolves; venv with jinja2 3.1.6), through
transformers.utils.chat_template_utils.render_jinja_template(the functionapply_chat_templatedelegates to, whose_compile_jinja_templateinstallsthe override) with a probe over an insertion-ordered
{"z": 1, "a": 2},nested unsorted tool schemas and Unicode/HTML leaves:
{"z": 1, "a": 2}— INSERTION order (sort_keys=False), raw UTF-8(
ensure_ascii=False), NO HTML escaping (Jinja's builtin sorts keys andescapes HTML — exactly what the transformers override exists to avoid);
indent(int or string;indent=0/-1is newline with no spaces;empty containers stay
{}/[]),sort_keys=True(recursive sort),ensure_ascii=True(\uXXXXwith surrogate pairs),separators=(item, key)— all four accepted; plain jinja2 rejects three of them and sorts/escapes
by default.
Probe command (reproducible):
python3 -m venv /tmp/tfprobe-venv && /tmp/tfprobe-venv/bin/pip install transformers==5.14.1 jinja2==3.1.6, thenrender a probe template through the pinned renderer —
/tmp/tfprobe-venv/bin/python /tmp/probe_tojson.pyrenders{{ obj | tojson }},{{ tool | tojson }},{{ s | tojson }}and the four option forms throughrender_jinja_templateand contrasts each with plain jinja2.Fix:
chat_template.cpp'stojsonis now a port of the pinned filter's fullsignature (CPython
json.dumpssemantics) instead of the sorted-dumpoverride;
BuildToolsmirrors the measured pinned vLLM tool shape(
[tool.model_dump() for tool in request.tools]atonline_renderer.py:178:fixed field order,
description/parameterspresent as null when absent);tool schemas keep the request document's key order end to end
(
ordered_jsonparameters+RestoreToolSchemaOrderat every chat entrypoint).
P2 — the registered template test depended on a private absolute model
path. The template input is now committed at
tests/fixtures/kolibri1-chat-template-tokenizer_config.json(thecheckpoint's
chat_template, sha2569ba35d4bd6baa26b66aa75d03a922dfee98b16bb1fa37481b195d247267b0f97in thefixture's
fixture_provenance); the test loads it (and the references) fromKOLIBRI1_TEMPLATE_FIXTURE_DIRwith no/mntpath, and the generator writesin-tree. Proven by running the test from
/with both fixtures copied to/tmp/p2-proof: 61/61 PASS with the model directory absent.The kolibri1 references were REGENERATED through the pinned renderer (the
generator asserts
transformers==5.14.1): 20 scenarios (was 15), addingunsorted nested tool schemas, Unicode/HTML content, a description-less tool
(pins
"description": null), unsorted historical tool-call arguments (thetemplate's second tojson site,
tool_call.arguments | tojson), and aUnicode/HTML user message. Red-first under the corrected reference: 6 of 20
scenarios failed before the adapter fix (
with_tools,with_tools_thinking_off,with_tools_unsorted_unicode,with_tools_unsorted_unicode_thinking_off,with_tools_minimal_no_description,assistant_tool_call_unsorted_arguments).Non-kolibri guard:
test_chat_templatepins the shared seam on non-kolibritemplates (+8 assertions): insertion-ordered tojson byte-exact vs the pin on
the Hermes/Qwen-style tool branch AND the real Qwen3.5 fixture template, the
four options byte-exact, and an already-alphabetical schema byte-identical
before/after the repair.
Re-review repair (2026-10-09, scoped re-review F1/F2 + the operator's clean-head check)
The operator's clean-head re-verification at the review-repair head
d93d9d492found the repair itself RED, and the scoped re-review filed twofindings. Both are fixed on this branch; the "Review repair" gate list above
is corrected here:
test_openai_run_batch16 was a stale binary — atd93d9d492the suite fails 3 of 7 cases, all throwing[json.exception.type_error.302] type must be string, but is object, whilethe doctest summary still reads
assertions: 16 | 16 passed, which is whatthe gate report counted. ISSUE-LOCAL-01M4FR20MES4HQVBRWJBJ2AVCN.
The crash. The repair wired
RestoreToolSchemaOrder(body, request)intoRunBatch::DispatchChat(run_batch.cpp:110), but the seam's first parameteris the body TEXT and
bodyis the parsednlohmann::jsonobject; theimplicit
get<std::string>()conversion throws on every object body, andDispatchChathas no try/catch around the call, so every object-body chatline through
RunBatchthrows. The other two call sites pass the raw bodystring and were already correct (api_server.cpp:400, vllm_c.cpp:518).
The fix (commit
a6679a82a). The original request text IS available:RunLineparsesrequest_jsonand dropped the text.RunLinenowre-serializes the chat body (the line's top-level
bodymember) from anorder-preserving
nlohmann::ordered_jsonparse of the original line text andthreads it through
DispatchChat(custom_id, body, body_json);DispatchChatcalls
RestoreToolSchemaOrder(body_json, request).ordered_json::dumpkeeps insertion order, so the re-read restores the request document's schema
key order into the prompt, matching the other entry points. A first in-flow
attempt (
body.dump()) was REVERTED:bodyis the key-sortednlohmann::json, so its dump is already sorted and the order restorationreads sorted text — the crash goes away but the seam no-ops.
F1 — the ordered-parameters seam was detected by no committed test
(ISSUE-LOCAL-01M4FR1D7RH7CN5X2R2CAR6N61, commit
2f535c002a). Twoentry-point tests now pin it:
test_run_batch.cpp, new case "run_batch: a tools line keeps thedocument's schema key order in the prompt": drives
RunBatch::RunLinewitha RAW JSONL line whose nested chat body is non-alphabetical at every level
(the raw string is load-bearing —
BatchLine()would re-serialize throughthe key-sorting
nlohmann::jsonand destroy the order under test).Asserts (a) no exception + a 200 row and (b) the prompt seam receives the
schema in the request document's key order, not the sorted form.
test_api_server.cpp, new case "api_server: an unsorted tool schema keepsits document key order in the rendered prompt": posts an unsorted-schema
tools request through the PRODUCTION
/v1/chat/completionsdispatch withthe real Qwen3.8 fixture template and asserts the rendered prompt bytes.
Red-first / mutation evidence: with the committed throwing call restored,
4 of 8
test_openai_run_batchcases throw type_error.302 (the 3 pre-existingto
body.dump()the 3 old cases pass but the new case fails BOTH key-orderassertions (the sorted dump no-ops the seam); with
RestoreToolSchemaOrderno-op'd at api_server.cpp:400 the api_server case fails both key-order
assertions. At the fix:
test_openai_run_batch8/8 cases / 89/89 assertions,test_openai_api_server104/104 cases / 1521/1521 assertions.F2 — the "byte-stable" guard's fixture was not alphabetical (same
issue). The
AlphabeticalToolfixture was insertion-orderednon-alphabetically, so the case's before/after byte-identity claim was not
demonstrated. The fixture is now genuinely alphabetical at every level
(
properties<required<type; single-property schema) and the caserenders the schema alone (
{{ tools[0].function.parameters | tojson }})because the tool wrapper's
model_dumpfield order is not alphabetical.Mutation-proven: with the
tojsondefault mutated to sorted keys (theoverride the repair removed) the insertion-order case goes RED while the
byte-stable case stays GREEN.
test_chat_template45/45 cases, 204/204assertions.
Gates re-verified at the fixed head (CPU-only, this host): the battery in
"How to verify" with
test_openai_run_batch8/8 cases / 89/89 assertions(7 pre-existing + the new F1 case) and
test_openai_api_server1521 (1517pre-existing + 4 new); everything else unchanged and green; W3 not rerun (no
forward change).
scripts/check-agent-record.pyOK (ANCHOR-ROT=0);scripts/agent-preflight.sh --stagedexits 0. Evidence:docs/bench-evidence/kolibri1-serve-20261008.md("Re-review repair").How to verify
Observed on this branch (CPU-only, this host):
test_reasoning_kolibri143,test_tool_parser_kolibri119,test_kolibri1_chat_template61,test_chat_template204 (196 pre-existing + 8 new guard assertions),test_reasoning_parser_detect75,test_tool_parser_detect361,test_reasoning_qwen3164,test_openai_tool_parsers64 — all PASS.test_kolibri127/234,test_kolibri1_decode_bench(default config, anchor109726),
test_kolibri1_dequant10,test_kolibri1_dequant_cache52,test_kolibri1_moe_glue21,test_kolibri1_w21608 — all PASS.Serving/protocol suites green:
test_openai_serving1365,test_openai_api_server1521 (1517 pre-existing + 4 new),test_openai_conformance252,test_parser_engine_assembly5038,test_openai_api_server_dots3_mm_forward16499,test_openai_run_batch8/8 cases / 89/89 assertions (7 pre-existing + the new F1 case; the repair
head's "16" was a stale binary — see "Re-review repair"),
test_input_batch232; the parameters-reading tool-parser suites(
test_deepseek_v32,test_glm47,test_minimax_m2_tool,test_tool_parser_step3,test_tool_parser_step3p5,test_tool_parser_qwen3_coder,test_tool_parser_minicpm5,test_tool_parser_hy_v3,test_tool_parser_poolside_v1,test_tool_parser_gemma4) all PASS.test_capi710/711 — the one failure iscapi v27(gliner fixture load), the pre-existing host failure recorded inthe evidence doc. W3 (
test_kolibri1_w3) NOT rerun: no forward change; thedecode bench's own gate is green. Full tree builds (881 targets);
scripts/check-tree-compiles.py161/161 in scope;scripts/check-agent-record.pyOK (ANCHOR-ROT=0);scripts/agent-preflight.sh --stagedexits 0. Red-first evidence, themeasurement, plugin
file:lineanchors and the reference-capture recipe arein
docs/bench-evidence/kolibri1-serve-20261008.md("Review repair").Out of scope: the plugin gateability measurement (
vllm serveon a GPUlease, owed by the oracle file),
chat_template_kwargsthreading for theOTHER engine-backed reasoning parsers (the pre-existing W4 note), the fp8 KV
cache / 1M-context serving recipe, and the pre-existing
test_capiv27 hostfailure.
Issue
ISSUE-LOCAL-01M4EF3R0H2H3FN5NA0NAB0S62 (row
MODEL-TEXT-kolibri-1-kolibri1-for-causal-lm), closed with dated resolution in the same branch; the review repair is recorded in its Resolution (the tracker forbids reopening via update).ISSUE-LOCAL-01M4FR20MES4HQVBRWJBJ2AVCN (row
MODEL-TEXT-kolibri-1-kolibri1-for-causal-lm), closed with dated resolution in the same branch: the confirmed run_batch crash (every object-body chat line threw json.exception.type_error.302 at the repair head), its fix (the original body text threaded through RunLine -> DispatchChat), and the red-first/mutation evidence.ISSUE-LOCAL-01M4FR1D7RH7CN5X2R2CAR6N61 (row
MODEL-TEXT-kolibri-1-kolibri1-for-causal-lm), closed with dated resolution in the same branch: the scoped re-review findings F1 (the ordered-parameters seam detected by no committed test — now pinned at the run_batch and api_server entry points) and F2 (the "byte-stable" guard's non-alphabetical fixture — now genuinely alphabetical at every level).FOLLOWING_AGENTS_PROTOCOL
Following-Agents-Protocol: true
AI-Assisted: true
Assisted-by: AGENT:mistral/mistral-large-4 [maki]