Skip to content

feat(kolibri1): serve the chat template and the kolibri1 parsers end to end - #3422

Merged
lu-zero merged 12 commits into
localai-org:mainfrom
lu-zero:row/kolibri-serve
Oct 9, 2026
Merged

lu-zero merged 12 commits into
localai-org:mainfrom
lu-zero:row/kolibri-serve

Conversation

@lu-zero

@lu-zero lu-zero commented Oct 8, 2026 •

Copy link
Copy Markdown
Collaborator

What

Serving completion for the kolibri1 CPU arm: the OpenAI chat path now renders
the checkpoint's chat template and parses kolibri1 reasoning and tool calls.
Review-repaired (see "Review repair" below) after the PR review blocked the
first landing on two findings against the shared tojson seam and the
template test's model-path dependency.

  • kolibri1 reasoning parser (src/vllm/entrypoints/openai/reasoning_parsers/kolibri1.{h,cpp}): the plugin's reasoning.py:36-61 ported — the Qwen3 engine grammar with the starting state derived per request from chat_template_kwargs (reasoning_effort wins and only "none" disables; else a literal enable_thinking: false does).
  • kolibri1 tool parser: an alias to HermesToolParser, exactly the plugin's registration (__init__.py:50-54).
  • Detection rows for both tables keyed on the template's no-reasoning sentence, ahead of the generic <think> / hermes rows; registry count pins 12→13 (reasoning) and 42→43 (tool).
  • Renderer parity in src/vllm/entrypoints/chat_template.cpp: tojson is a port of the pinned transformers 5.14.1 renderer's filter (utils/chat_template_utils.py:481: sort_keys=False, ensure_ascii=False, no HTML escaping) — insertion key order, raw UTF-8 — replacing the first landing's sorted-dump override, which the review falsified (see "Review repair").
  • Tool schemas keep the request document's key order end to end: FunctionDefinition::parameters is nlohmann::ordered_json, and every chat entry point (api_server, the C ABI, run_batch) re-reads tools from an order-preserving body parse (RestoreToolSchemaOrder).
  • Gates: test_reasoning_kolibri1 (43), test_tool_parser_kolibri1 (19), test_kolibri1_chat_template (61: byte-exact vs the pinned transformers renderer on 20 scenarios, tests/fixtures/kolibri1_chat_template_references.json), test_chat_template (204: 196 pre-existing + 8 non-kolibri tojson guards).

Why

Upstream vLLM implements nothing for Kolibri1ForCausalLM; the model-author
plugin (aleph-alpha-inference @ 049a6a7bd240, the row's serving oracle per
.agents/oracles/aleph-alpha-inference.md) is where the behavior is defined,
and its serving recipe (--reasoning-parser kolibri1 --tool-call-parser kolibri1) could not be served from this tree. Closes
ISSUE-LOCAL-01M4EF3R0H2H3FN5NA0NAB0S62; spec ## Now updated in the same
branch.

Review repair (2026-10-09, review blockers P1/P2)

The review (localai-org-maint-bot) blocked the first landing with two findings.

P1 — the global tojson override used the wrong serving reference. The
first landing sorted object keys for every model, because its references were
captured with plain jinja2. Measured against the PINNED transformers 5.14.1
renderer (.agents/oracles/transformers.md, the version the pinned vLLM
environment resolves; venv with jinja2 3.1.6), through
transformers.utils.chat_template_utils.render_jinja_template (the function
apply_chat_template delegates to, whose _compile_jinja_template installs
the override) with a probe over an insertion-ordered {"z": 1, "a": 2},
nested unsorted tool schemas and Unicode/HTML leaves:

  • default: {"z": 1, "a": 2} — INSERTION order (sort_keys=False), raw UTF-8
    (ensure_ascii=False), NO HTML escaping (Jinja's builtin sorts keys and
    escapes HTML — exactly what the transformers override exists to avoid);
  • options: indent (int or string; indent=0/-1 is newline with no spaces;
    empty containers stay {}/[]), sort_keys=True (recursive sort),
    ensure_ascii=True (\uXXXX with surrogate pairs), separators=(item, key)
    — all four accepted; plain jinja2 rejects three of them and sorts/escapes
    by default.

Probe command (reproducible): python3 -m venv /tmp/tfprobe-venv && /tmp/tfprobe-venv/bin/pip install transformers==5.14.1 jinja2==3.1.6, then
render a probe template through the pinned renderer —
/tmp/tfprobe-venv/bin/python /tmp/probe_tojson.py renders {{ obj | tojson }},
{{ tool | tojson }}, {{ s | tojson }} and the four option forms through
render_jinja_template and contrasts each with plain jinja2.

Fix: chat_template.cpp's tojson is now a port of the pinned filter's full
signature (CPython json.dumps semantics) instead of the sorted-dump
override; BuildTools mirrors the measured pinned vLLM tool shape
([tool.model_dump() for tool in request.tools] at online_renderer.py:178:
fixed field order, description/parameters present as null when absent);
tool schemas keep the request document's key order end to end
(ordered_json parameters + RestoreToolSchemaOrder at every chat entry
point).

P2 — the registered template test depended on a private absolute model
path.
The template input is now committed at
tests/fixtures/kolibri1-chat-template-tokenizer_config.json (the
checkpoint's chat_template, sha256
9ba35d4bd6baa26b66aa75d03a922dfee98b16bb1fa37481b195d247267b0f97 in the
fixture's fixture_provenance); the test loads it (and the references) from
KOLIBRI1_TEMPLATE_FIXTURE_DIR with no /mnt path, and the generator writes
in-tree. Proven by running the test from / with both fixtures copied to
/tmp/p2-proof: 61/61 PASS with the model directory absent.

The kolibri1 references were REGENERATED through the pinned renderer (the
generator asserts transformers==5.14.1): 20 scenarios (was 15), adding
unsorted nested tool schemas, Unicode/HTML content, a description-less tool
(pins "description": null), unsorted historical tool-call arguments (the
template's second tojson site, tool_call.arguments | tojson), and a
Unicode/HTML user message. Red-first under the corrected reference: 6 of 20
scenarios failed before the adapter fix (with_tools,
with_tools_thinking_off, with_tools_unsorted_unicode,
with_tools_unsorted_unicode_thinking_off, with_tools_minimal_no_description,
assistant_tool_call_unsorted_arguments).

Non-kolibri guard: test_chat_template pins the shared seam on non-kolibri
templates (+8 assertions): insertion-ordered tojson byte-exact vs the pin on
the Hermes/Qwen-style tool branch AND the real Qwen3.5 fixture template, the
four options byte-exact, and an already-alphabetical schema byte-identical
before/after the repair.

Re-review repair (2026-10-09, scoped re-review F1/F2 + the operator's clean-head check)

The operator's clean-head re-verification at the review-repair head
d93d9d492 found the repair itself RED, and the scoped re-review filed two
findings. Both are fixed on this branch; the "Review repair" gate list above
is corrected here: test_openai_run_batch 16 was a stale binary — at
d93d9d492 the suite fails 3 of 7 cases, all throwing
[json.exception.type_error.302] type must be string, but is object, while
the doctest summary still reads assertions: 16 | 16 passed, which is what
the gate report counted. ISSUE-LOCAL-01M4FR20MES4HQVBRWJBJ2AVCN.

The crash. The repair wired RestoreToolSchemaOrder(body, request) into
RunBatch::DispatchChat (run_batch.cpp:110), but the seam's first parameter
is the body TEXT and body is the parsed nlohmann::json object; the
implicit get<std::string>() conversion throws on every object body, and
DispatchChat has no try/catch around the call, so every object-body chat
line through RunBatch throws. The other two call sites pass the raw body
string and were already correct (api_server.cpp:400, vllm_c.cpp:518).

The fix (commit a6679a82a). The original request text IS available:
RunLine parses request_json and dropped the text. RunLine now
re-serializes the chat body (the line's top-level body member) from an
order-preserving nlohmann::ordered_json parse of the original line text and
threads it through DispatchChat(custom_id, body, body_json); DispatchChat
calls RestoreToolSchemaOrder(body_json, request). ordered_json::dump
keeps insertion order, so the re-read restores the request document's schema
key order into the prompt, matching the other entry points. A first in-flow
attempt (body.dump()) was REVERTED: body is the key-sorted
nlohmann::json, so its dump is already sorted and the order restoration
reads sorted text — the crash goes away but the seam no-ops.

F1 — the ordered-parameters seam was detected by no committed test
(ISSUE-LOCAL-01M4FR1D7RH7CN5X2R2CAR6N61, commit 2f535c002a). Two
entry-point tests now pin it:

  • test_run_batch.cpp, new case "run_batch: a tools line keeps the
    document's schema key order in the prompt": drives RunBatch::RunLine with
    a RAW JSONL line whose nested chat body is non-alphabetical at every level
    (the raw string is load-bearing — BatchLine() would re-serialize through
    the key-sorting nlohmann::json and destroy the order under test).
    Asserts (a) no exception + a 200 row and (b) the prompt seam receives the
    schema in the request document's key order, not the sorted form.
  • test_api_server.cpp, new case "api_server: an unsorted tool schema keeps
    its document key order in the rendered prompt": posts an unsorted-schema
    tools request through the PRODUCTION /v1/chat/completions dispatch with
    the real Qwen3.8 fixture template and asserts the rendered prompt bytes.

Red-first / mutation evidence: with the committed throwing call restored,
4 of 8 test_openai_run_batch cases throw type_error.302 (the 3 pre-existing

  • the new case) while assertions read 16/16; with the call's argument mutated
    to body.dump() the 3 old cases pass but the new case fails BOTH key-order
    assertions (the sorted dump no-ops the seam); with RestoreToolSchemaOrder
    no-op'd at api_server.cpp:400 the api_server case fails both key-order
    assertions. At the fix: test_openai_run_batch 8/8 cases / 89/89 assertions,
    test_openai_api_server 104/104 cases / 1521/1521 assertions.

F2 — the "byte-stable" guard's fixture was not alphabetical (same
issue). The AlphabeticalTool fixture was insertion-ordered
non-alphabetically, so the case's before/after byte-identity claim was not
demonstrated. The fixture is now genuinely alphabetical at every level
(properties < required < type; single-property schema) and the case
renders the schema alone ({{ tools[0].function.parameters | tojson }})
because the tool wrapper's model_dump field order is not alphabetical.
Mutation-proven: with the tojson default mutated to sorted keys (the
override the repair removed) the insertion-order case goes RED while the
byte-stable case stays GREEN. test_chat_template 45/45 cases, 204/204
assertions.

Gates re-verified at the fixed head (CPU-only, this host): the battery in
"How to verify" with test_openai_run_batch 8/8 cases / 89/89 assertions
(7 pre-existing + the new F1 case) and test_openai_api_server 1521 (1517
pre-existing + 4 new); everything else unchanged and green; W3 not rerun (no
forward change). scripts/check-agent-record.py OK (ANCHOR-ROT=0);
scripts/agent-preflight.sh --staged exits 0. Evidence:
docs/bench-evidence/kolibri1-serve-20261008.md ("Re-review repair").

How to verify

cmake -G Ninja -DCMAKE_BUILD_TYPE=Release -DVLLM_CPP_CUDA=OFF /tmp/vllm-kolibri-serve -S . -B /tmp/build-kolibri-serve
ninja -C /tmp/build-kolibri-serve test_reasoning_kolibri1 test_tool_parser_kolibri1 test_kolibri1_chat_template test_reasoning_parser_detect test_tool_parser_detect test_reasoning_qwen3 test_openai_tool_parsers test_chat_template
ctest --test-dir /tmp/build-kolibri-serve -R kolibri

Observed on this branch (CPU-only, this host): test_reasoning_kolibri1 43,
test_tool_parser_kolibri1 19, test_kolibri1_chat_template 61,
test_chat_template 204 (196 pre-existing + 8 new guard assertions),
test_reasoning_parser_detect 75, test_tool_parser_detect 361,
test_reasoning_qwen3 164, test_openai_tool_parsers 64 — all PASS.
test_kolibri1 27/234, test_kolibri1_decode_bench (default config, anchor
109726), test_kolibri1_dequant 10, test_kolibri1_dequant_cache 52,
test_kolibri1_moe_glue 21, test_kolibri1_w2 1608 — all PASS.
Serving/protocol suites green: test_openai_serving 1365,
test_openai_api_server 1521 (1517 pre-existing + 4 new), test_openai_conformance 252,
test_parser_engine_assembly 5038,
test_openai_api_server_dots3_mm_forward 16499, test_openai_run_batch
8/8 cases / 89/89 assertions (7 pre-existing + the new F1 case; the repair
head's "16" was a stale binary — see "Re-review repair"), test_input_batch 232; the parameters-reading tool-parser suites
(test_deepseek_v32, test_glm47, test_minimax_m2_tool,
test_tool_parser_step3, test_tool_parser_step3p5,
test_tool_parser_qwen3_coder, test_tool_parser_minicpm5,
test_tool_parser_hy_v3, test_tool_parser_poolside_v1,
test_tool_parser_gemma4) all PASS. test_capi 710/711 — the one failure is
capi v27 (gliner fixture load), the pre-existing host failure recorded in
the evidence doc. W3 (test_kolibri1_w3) NOT rerun: no forward change; the
decode bench's own gate is green. Full tree builds (881 targets);
scripts/check-tree-compiles.py 161/161 in scope;
scripts/check-agent-record.py OK (ANCHOR-ROT=0);
scripts/agent-preflight.sh --staged exits 0. Red-first evidence, the
measurement, plugin file:line anchors and the reference-capture recipe are
in docs/bench-evidence/kolibri1-serve-20261008.md ("Review repair").

Out of scope: the plugin gateability measurement (vllm serve on a GPU
lease, owed by the oracle file), chat_template_kwargs threading for the
OTHER engine-backed reasoning parsers (the pre-existing W4 note), the fp8 KV
cache / 1M-context serving recipe, and the pre-existing test_capi v27 host
failure.

Issue

ISSUE-LOCAL-01M4EF3R0H2H3FN5NA0NAB0S62 (row MODEL-TEXT-kolibri-1-kolibri1-for-causal-lm), closed with dated resolution in the same branch; the review repair is recorded in its Resolution (the tracker forbids reopening via update).

ISSUE-LOCAL-01M4FR20MES4HQVBRWJBJ2AVCN (row MODEL-TEXT-kolibri-1-kolibri1-for-causal-lm), closed with dated resolution in the same branch: the confirmed run_batch crash (every object-body chat line threw json.exception.type_error.302 at the repair head), its fix (the original body text threaded through RunLine -> DispatchChat), and the red-first/mutation evidence.

ISSUE-LOCAL-01M4FR1D7RH7CN5X2R2CAR6N61 (row MODEL-TEXT-kolibri-1-kolibri1-for-causal-lm), closed with dated resolution in the same branch: the scoped re-review findings F1 (the ordered-parameters seam detected by no committed test — now pinned at the run_batch and api_server entry points) and F2 (the "byte-stable" guard's non-alphabetical fixture — now genuinely alphabetical at every level).

FOLLOWING_AGENTS_PROTOCOL

Following-Agents-Protocol: true
AI-Assisted: true
Assisted-by: AGENT:mistral/mistral-large-4 [maki]

The kolibri1 CPU arm reaches the model forward but not serving. This
issue scopes the serving-completion unit: the kolibri1 reasoning parser
(the plugin's Qwen3 grammar with the template's thinking switch), the
kolibri1 tool-parser alias to the Hermes format, the template detection
rows, and a jinja2-referenced rendering gate for the tokenizer-shipped
chat template. Oracle: aleph-alpha-inference pin 049a6a7bd240.

FOLLOWING_AGENTS_PROTOCOL

Following-Agents-Protocol: true
AI-Assisted: true
Assisted-by: AGENT:zai/glm-5.3-flash [maki]
…to end

The kolibri1 CPU arm reached the model forward but not serving: no
kolibri1 reasoning parser, no kolibri1 tool-parser name, and detection
resolved think_auto/hermes off the Kolibri template. The model-author
plugin (aleph-alpha-inference @ 049a6a7bd240) is the serving oracle:
its reasoning parser is the Qwen3 grammar with the starting state
derived the way the template switches thinking (reasoning_effort wins
and only "none" disables; else a literal enable_thinking false does),
and its tool parser IS the Hermes class under the kolibri1 name.

- kolibri1 reasoning adapter (reasoning.py:36-61 ported): per-request
  state from chat_template_kwargs, over the shared Qwen3 engine.
- kolibri1 tool parser: alias to HermesToolParser (init.py:50-54).
- Detection rows on the template's no-reasoning sentence, ahead of the
  generic <think> and hermes rows; registry pins 12->13 and 42->43.
- Renderer parity: minja's tojson dumped insertion order; jinja2's
  default policy sorts keys. Fixed adapter-side with a child-scope
  sorted-dump tojson; no vendor change.
- Gate: byte-exact rendering vs CPython jinja2 references on 15
  scenarios (fixtures/kolibri1_chat_template_references.json), parser
  unit tests on the plugin's switch truth table, alias equivalence with
  hermes. Red-first: detection and registry cases failed before the
  change. Evidence: docs/bench-evidence/kolibri1-serve-20261008.md.

FOLLOWING_AGENTS_PROTOCOL

Following-Agents-Protocol: true
AI-Assisted: true
Assisted-by: AGENT:zai/glm-5.3-flash [maki]
The serving unit landed in the previous commit; this records the dated
resolution on the canonical issue and refreshes the derived index.

FOLLOWING_AGENTS_PROTOCOL

Following-Agents-Protocol: true
AI-Assisted: true
Assisted-by: AGENT:zai/glm-5.3-flash [maki]
@lu-zero
lu-zero force-pushed the row/kolibri-serve branch from ef97514 to 4aade35 Compare October 8, 2026 23:39

@localai-org-maint-bot localai-org-maint-bot left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

@mudler @richiejp Two blockers from source review of 4aade352b1b052226c5035423c1cafc38f90475a:

  1. P1: the global tojson override uses the wrong serving reference. src/vllm/entrypoints/chat_template.cpp now sorts keys for every model. The pinned Transformers 5.14.1 renderer overrides Jinja's filter with tojson(..., sort_keys=False) and installs that filter in _compile_jinja_template: https://github.com/huggingface/transformers/blob/v5.14.1/src/transformers/utils/chat_template_utils.py#L481 . Thus plain Jinja's default is not the Transformers serving behavior. For an insertion-ordered object {"z":1,"a":2}, the reference preserves z,a; this override forces a,z, changing prompt bytes and tokens across models. The new fixture generator uses plain jinja2.Environment, so it certifies the same incorrect reference. Regenerate through the pinned Transformers renderer and preserve its default/options before changing this shared seam; include unsorted nested tool schemas and Unicode/HTML characters in the comparison.

  2. P2: the registered template test depends on a private absolute model path. tests/vllm/entrypoints/test_kolibri1_chat_template.cpp loads /mnt/models/Aleph-Alpha/Kolibri-1/tokenizer_config.json in both rendering and detection cases. The fixture includes expected strings but not that input template. LoadChatTemplateFromConfig throws when the file is absent, so a clean checkout cannot run this test using the committed fixtures. Commit the pinned template input and load it from the fixture directory (or explicitly provision a pinned artifact); the generator's hard-coded output under /tmp/vllm-kolibri-serve also needs a portable destination.

Validation: inspected the complete changed parser/renderer/test code and the pinned Transformers filter; traced the missing-file exception in the current base. Native tests and mutation checks were not run because this environment lacks Python/CMake/a C++ compiler. This is a findings review, not an acceptance or performance gate.

…ma order

The PR localai-org#3422 review falsified the global tojson override that sorted
object keys for every model. Measured first, against the pinned
transformers 5.14.1 renderer (.agents/oracles/transformers.md) through
render_jinja_template with a probe over an insertion-ordered
{"z": 1, "a": 2}, nested unsorted tool schemas and Unicode/HTML
leaves: the pinned renderer's tojson override
(utils/chat_template_utils.py:481) is json.dumps with sort_keys=False,
ensure_ascii=False and NO HTML escaping, so the serving reference keeps
insertion key order and raw UTF-8 -- plain Jinja's builtin (sorted,
HTML-escaped) is not the serving behavior, and the old sorted-dump
override rendered prompt bytes no serving reference produces.

chat_template.cpp replaces the sorted-dump override with a port of the
pinned filter's FULL signature (ensure_ascii/indent/separators/
sort_keys, CPython json.dumps semantics; minja's builtin matches the
default but rejects three of the four options). BuildTools now mirrors
the measured pinned vLLM tool shape ([tool.model_dump() ...] at
online_renderer.py:178: fixed field order, description/parameters
present as null when absent).

FunctionDefinition::parameters becomes nlohmann::ordered_json
(nlohmann::json sorts object keys at parse, losing the request
document's order that the pinned renderer preserves into the prompt),
with an ordered from_json overload and RestoreToolSchemaOrder, which
re-reads tools from an order-preserving body parse. Every chat entry
point calls it after its regular parse: api_server handle_chat_
completions, the C ABI ParseChatRequest (vllm_chat and
vllm_chat_stream), and run_batch. Key-lookup readers (the tool parsers)
are behavior-identical; step3.cpp's one pointer type follows the field.

Gates: full tree builds (881 targets); test_kolibri1_chat_template 61,
test_chat_template 204, test_reasoning_kolibri1 43,
test_tool_parser_kolibri1 19, test_reasoning_parser_detect 75,
test_tool_parser_detect 361, test_reasoning_qwen3 164,
test_openai_tool_parsers 64, test_kolibri1 27/234,
test_kolibri1_decode_bench (anchor 109726), test_openai_serving 1365,
test_openai_api_server 1517, test_openai_conformance 252,
test_parser_engine_assembly 5038, test_openai_run_batch 16, and the
parameters-reading tool-parser suites all green; check-tree-compiles
161/161 in scope; check-agent-record OK.

FOLLOWING_AGENTS_PROTOCOL

Following-Agents-Protocol: true
AI-Assisted: true
Assisted-by: AGENT:mistral/mistral-large-4 [maki]
…inned renderer

Review blockers P1/P2 on the fixture side. The old generator rendered
with plain jinja2 (sorted keys, HTML-escaped -- not the serving
reference) and hard-coded its output under /tmp/vllm-kolibri-serve;
the test loaded /mnt/models/Aleph-Alpha/Kolibri-1/
tokenizer_config.json in both the rendering and detection cases.

The generator now drives the PINNED transformers 5.14.1 renderer
(render_jinja_template, the function apply_chat_template delegates to;
it asserts transformers==5.14.1 and records the pin, jinja2's version,
the template sha256 and the capture path in the provenance), reads the
template from the committed fixture config, mirrors pinned vLLM's
_postprocess_messages (assistant tool-call arguments parsed to dicts)
and model_dump tool shape, and writes in-tree next to itself. The
reference fixture grows from 15 to 20 scenarios: the 15 original plus
with_tools_unsorted_unicode and with_tools_unsorted_unicode_thinking_off
(unsorted nested tool schemas + Unicode/HTML), with_tools_minimal_no_
description (pins "description": null), assistant_tool_call_unsorted_
arguments (the template's second tojson site, tool_call.arguments |
tojson, unsorted Unicode/HTML arguments), and unicode_html_user_message.

The template input is committed at tests/fixtures/kolibri1-chat-
template-tokenizer_config.json (the checkpoint's chat_template, sha256
9ba35d4bd6baa26b66aa75d03a922dfee98b16bb1fa37481b195d247267b0f97 in the
fixture's fixture_provenance). The test loads the template and the
references from KOLIBRI1_TEMPLATE_FIXTURE_DIR (tests/fixtures at
compile time); no /mnt path remains in any load path -- proven by
running the test from / with both fixtures copied to /tmp/p2-proof:
61/61 PASS with the model directory absent.

Red-first under the corrected reference (unmodified adapter, 20-case
fixture): 6 of 20 rendering scenarios failed -- with_tools,
with_tools_thinking_off, with_tools_unsorted_unicode,
with_tools_unsorted_unicode_thinking_off, with_tools_minimal_no_
description, assistant_tool_call_unsorted_arguments (34/40 assertions).
After the adapter fix: 61/61 PASS.

FOLLOWING_AGENTS_PROTOCOL

Following-Agents-Protocol: true
AI-Assisted: true
Assisted-by: AGENT:mistral/mistral-large-4 [maki]
…plates

The tojson fix touches a seam every model's rendering depends on, so
the non-kolibri behavior is pinned explicitly (+8 assertions, 204
total): a non-kolibri tool template (the Hermes/Qwen-style tool branch
AND the real Qwen3.5 fixture template) renders an insertion-ordered,
Unicode/HTML tool byte-identical to the pinned renderer; the four
tojson options (indent=2, sort_keys=True, ensure_ascii=True,
separators=(',', ':')) are pinned byte-for-byte against the pin; and
an already-alphabetical schema renders byte-identical before/after the
repair, so the change alters only what the pinned renderer actually
differs on. Expected bytes were generated by the pinned transformers
5.14.1 renderer itself (render_jinja_template, 2026-10-09).

FOLLOWING_AGENTS_PROTOCOL

Following-Agents-Protocol: true
AI-Assisted: true
Assisted-by: AGENT:mistral/mistral-large-4 [maki]
…ssue, spec, anchors

Records the repair of both review blockers on the serving-completion
issue ISSUE-LOCAL-01M4EF3R0H2H3FN5NA0NAB0S62 (the tracker forbids
reopening via update, so the repair is recorded in the Resolution):
docs/bench-evidence/kolibri1-serve-20261008.md gains a "Review
repair" section -- what the pinned transformers 5.14.1 renderer
actually does (MEASURED: insertion order, raw UTF-8, no HTML escaping,
plus the four tojson options and the model_dump tool shape), what the
old sorted-dump override did, the red-first evidence (6 of 20 scenarios
failed under the corrected reference), the fixture regeneration, the
portability fix and the gate numbers -- and marks the two sections that
recorded the falsified decision SUPERSEDED. The spec's SERVING
COMPLETION paragraph in .agents/specs/kolibri-1-cpu.md is corrected the
same way. Three engine-matrix.md api_server.cpp anchors shifted by the
+5-line RestoreToolSchemaOrder call are repaired (1354->1359,
1365->1370, 1614->1619); check-agent-record reports ANCHOR-ROT=0.

FOLLOWING_AGENTS_PROTOCOL

Following-Agents-Protocol: true
AI-Assisted: true
Assisted-by: AGENT:mistral/mistral-large-4 [maki]
The review-repair commits are pushed; the PR body gained the "Review
repair" section after that push, and a body edit alone does not rerun
the guard jobs. This empty commit retriggers CI so the trailer/style
guards read the body that will become the squash commit message.

FOLLOWING_AGENTS_PROTOCOL

Following-Agents-Protocol: true
AI-Assisted: true
Assisted-by: AGENT:mistral/mistral-large-4 [maki]
…oreToolSchemaOrder

The PR localai-org#3422 repair wired RestoreToolSchemaOrder into RunBatch::DispatchChat
as RestoreToolSchemaOrder(body, request), but the seam's first parameter is
the body TEXT (const std::string&, protocol.h:605) and body is the parsed
nlohmann::json object. The call compiles only through nlohmann's implicit
get<std::string>() conversion, which throws json.exception.type_error.302 on
every object body, and DispatchChat has no try/catch around it — so every
object-body /v1/chat/completions line through RunBatch throws. Clean-head
evidence at d93d9d4: test_openai_run_batch 3 of 7 cases throw
(test_run_batch.cpp:451/:507/:601); the repair's recorded 'run_batch 16/16'
gate counted assertions and missed the throwing cases (stale binary), so the
PR's CI lane would be red. ISSUE-LOCAL-01M4FR20MES4HQVBRWJBJ2AVCN.

A first in-flow attempt (body.dump()) is wrong and was reverted: body is the
key-sorted nlohmann::json, so its dump is already sorted and the order
restoration reads sorted text — the crash goes away but the seam no-ops and
the request document's tool-schema key order is lost on the batch path.

The original request text IS available: RunLine parses request_json and drops
it. RunLine now re-serializes the chat body (the line's top-level "body"
member) from an order-preserving nlohmann::ordered_json parse of the original
line text and threads it through DispatchChat(custom_id, body, body_json);
DispatchChat calls RestoreToolSchemaOrder(body_json, request).
ordered_json::dump keeps insertion order, so the re-read restores the request
document's schema key order into the prompt, matching the api_server and C-ABI
entry points, which already pass their raw body strings. The shared seam's
signature is unchanged.

Red-first: at the committed throwing call, 4 of 8 test_openai_run_batch cases
throw type_error.302 (the 3 pre-existing plus the new F1 case) while
assertions read 16/16; under the body.dump() mutation the 3 old cases pass but
the new F1 case fails both key-order assertions (seam no-op); at this fix
8/8 cases, 89/89 assertions pass.

FOLLOWING_AGENTS_PROTOCOL

Following-Agents-Protocol: true
AI-Assisted: true
Assisted-by: AGENT:mistral/mistral-large-4 [maki]
…nd api_server entry points

Scoped re-review F1 of PR localai-org#3422 (ISSUE-LOCAL-01M4FR1D7RH7CN5X2R2CAR6N61):
RestoreToolSchemaOrder was detected by NO committed test — a no-op passed
every gate, and the batch call site was not merely unguarded but crashing
(ISSUE-LOCAL-01M4FR20MES4HQVBRWJBJ2AVCN). Two entry-point tests now close the
gap, one per affected dispatch:

- tests/vllm/entrypoints/openai/test_run_batch.cpp: a new case drives
  RunBatch::RunLine (the batch entry point) with a RAW JSONL line whose nested
  chat body is non-alphabetical at every level (tool wrapper, function, and
  parameters schema with zeta before alpha). The raw string is load-bearing:
  BatchLine() re-serializes through the key-sorting nlohmann::json and would
  destroy the order under test. The case asserts (a) no exception and a 200
  row — pinning the crash fix — and (b) the prompt seam receives the schema in
  the request document's key order and not the key-sorted form, captured via a
  CapturingToolsPrompt seam over the parsed tools (the capture pattern the
  api_server harness uses). Mutation-proven in both directions: with the
  committed throwing call restored the case throws type_error.302; with the
  call's argument mutated to body.dump() it fails both key-order assertions
  (the sorted dump no-ops the seam).
- tests/vllm/entrypoints/openai/test_api_server.cpp: a new case posts an
  unsorted-schema tools request through the PRODUCTION
  /v1/chat/completions dispatch (handle_chat_completions) with the real
  Qwen3.8 fixture template and asserts the rendered prompt bytes carry the
  document order. Mutation-proven: with RestoreToolSchemaOrder no-op'd at
  api_server.cpp:400 the case fails both key-order assertions.

Gates at this head: test_openai_run_batch 8/8 cases / 89/89 assertions;
test_openai_api_server 104/104 cases / 1521/1521 assertions (1517 pre-existing
+ 4 new). No existing assertion weakened: both additions are new TEST_CASEs;
the run_batch file's pre-existing cases are untouched.

FOLLOWING_AGENTS_PROTOCOL

Following-Agents-Protocol: true
AI-Assisted: true
Assisted-by: AGENT:mistral/mistral-large-4 [maki]
…ec, evidence

The scoped re-review of PR localai-org#3422 (ISSUE-LOCAL-01M4FR1D7RH7CN5X2R2CAR6N61)
filed two findings; the F1 half landed with the F1 tests in the previous
commit. This commit lands F2 and the records the repair invalidates.

F2 (mislabeled guard case): the 'byte-stable for already-ordered schemas'
case in tests/vllm/entrypoints/test_chat_template.cpp used an
AlphabeticalTool fixture that was NOT fully alphabetical (the tool wrapper
is insertion-ordered non-alphabetically at type/function and
name/description/parameters, and the schema was type/properties/required),
so under a sorted tojson mutation the case went red too and the
'alphabetical-schema byte-identity before/after' claim was not demonstrated.
The fixture is now genuinely alphabetical at every level (schema
properties < required < type; single-property schema needs no order) and
the case renders the schema alone ({{ tools[0].function.parameters | tojson
}}) because the tool wrapper's field order is the pinned model_dump order,
which is NOT alphabetical — so the byte-identity claim is now real and the
comment says so. Mutation-proven: with the tojson default mutated to sorted
keys (the override the review repair removed) the insertion-order case goes
RED while this byte-stable case stays GREEN. No existing assertion was
weakened: the full-template exact-bytes rendering remains pinned by the
insertion-order case (UnsortedUnicodeTool) and the WeatherTool cases;
test_chat_template 45/45 cases, 204/204 assertions.

Records: ISSUE-LOCAL-01M4FR20MES4HQVBRWJBJ2AVCN (the confirmed run_batch
crash, rewritten — the first in-flow draft prescribed body.dump(), which
silences the crash but no-ops the order restoration because the dump is
already key-sorted; the landed fix threads the original body text through)
and ISSUE-LOCAL-01M4FR1D7RH7CN5X2R2CAR6N61 (F1+F2) are closed with dated
resolution evidence. The row spec's ## Now records the re-review repair and
its gates; docs/bench-evidence/kolibri1-serve-20261008.md gains a
'Re-review repair' section that corrects the falsified
'test_openai_run_batch 16' gate line (a stale binary: 3 of 7 cases threw
type_error.302 at the repair head while assertions read 16/16).

Gates at this head: test_openai_run_batch 8/8 cases / 89/89 assertions,
test_openai_api_server 104/104 / 1521/1521, test_openai_serving 48/48 /
1365/1365, test_chat_template 45/45 / 204/204, test_kolibri1_chat_template
4/4 / 61/61, test_reasoning_kolibri1 7/7 / 43/43, test_tool_parser_kolibri1
3/3 / 19/19, test_reasoning_parser_detect 8/8 / 75/75,
test_tool_parser_detect 16/16 / 361/361, test_kolibri1 27/27 / 234/234,
test_kolibri1_decode_bench 1/1 / 2/2 (anchor chain ends 109726) — all PASS;
W3 not rerun (no forward change). check-agent-record OK (ANCHOR-ROT=0);
agent-issue-index refreshed; agent-preflight.sh --staged exits 0 (five
argument-requiring checkers SKIP in this environment, unchanged by this
diff).

FOLLOWING_AGENTS_PROTOCOL

Following-Agents-Protocol: true
AI-Assisted: true
Assisted-by: AGENT:mistral/mistral-large-4 [maki]
The re-review repair commits are pushed; the PR body gained the
"Re-review repair" section after that push, and a body edit alone does
not rerun the guard jobs. This empty commit retriggers CI so the
trailer/style guards read the body that will become the squash commit
message.

FOLLOWING_AGENTS_PROTOCOL

Following-Agents-Protocol: true
AI-Assisted: true
Assisted-by: AGENT:mistral/mistral-large-4 [maki]
@lu-zero
lu-zero merged commit 5ce71e4 into localai-org:main Oct 9, 2026
25 of 31 checks passed
lu-zero added a commit that referenced this pull request Oct 9, 2026
…ma order

The PR #3422 review falsified the global tojson override that sorted
object keys for every model. Measured first, against the pinned
transformers 5.14.1 renderer (.agents/oracles/transformers.md) through
render_jinja_template with a probe over an insertion-ordered
{"z": 1, "a": 2}, nested unsorted tool schemas and Unicode/HTML
leaves: the pinned renderer's tojson override
(utils/chat_template_utils.py:481) is json.dumps with sort_keys=False,
ensure_ascii=False and NO HTML escaping, so the serving reference keeps
insertion key order and raw UTF-8 -- plain Jinja's builtin (sorted,
HTML-escaped) is not the serving behavior, and the old sorted-dump
override rendered prompt bytes no serving reference produces.

chat_template.cpp replaces the sorted-dump override with a port of the
pinned filter's FULL signature (ensure_ascii/indent/separators/
sort_keys, CPython json.dumps semantics; minja's builtin matches the
default but rejects three of the four options). BuildTools now mirrors
the measured pinned vLLM tool shape ([tool.model_dump() ...] at
online_renderer.py:178: fixed field order, description/parameters
present as null when absent).

FunctionDefinition::parameters becomes nlohmann::ordered_json
(nlohmann::json sorts object keys at parse, losing the request
document's order that the pinned renderer preserves into the prompt),
with an ordered from_json overload and RestoreToolSchemaOrder, which
re-reads tools from an order-preserving body parse. Every chat entry
point calls it after its regular parse: api_server handle_chat_
completions, the C ABI ParseChatRequest (vllm_chat and
vllm_chat_stream), and run_batch. Key-lookup readers (the tool parsers)
are behavior-identical; step3.cpp's one pointer type follows the field.

Gates: full tree builds (881 targets); test_kolibri1_chat_template 61,
test_chat_template 204, test_reasoning_kolibri1 43,
test_tool_parser_kolibri1 19, test_reasoning_parser_detect 75,
test_tool_parser_detect 361, test_reasoning_qwen3 164,
test_openai_tool_parsers 64, test_kolibri1 27/234,
test_kolibri1_decode_bench (anchor 109726), test_openai_serving 1365,
test_openai_api_server 1517, test_openai_conformance 252,
test_parser_engine_assembly 5038, test_openai_run_batch 16, and the
parameters-reading tool-parser suites all green; check-tree-compiles
161/161 in scope; check-agent-record OK.

FOLLOWING_AGENTS_PROTOCOL

Following-Agents-Protocol: true
AI-Assisted: true
Assisted-by: AGENT:mistral/mistral-large-4 [maki]
lu-zero added a commit that referenced this pull request Oct 9, 2026
… anchors

Records the repair of both review blockers on the serving-completion
issue ISSUE-LOCAL-01M4EF3R0H2H3FN5NA0NAB0S62 (the tracker forbids
reopening via update, so the repair is recorded in the Resolution):
docs/bench-evidence/kolibri1-serve-20261008.md gains a "Review
repair" section -- what the pinned transformers 5.14.1 renderer
actually does (MEASURED: insertion order, raw UTF-8, no HTML escaping,
plus the four tojson options and the model_dump tool shape), what the
old sorted-dump override did, the red-first evidence (6 of 20 scenarios
failed under the corrected reference), the fixture regeneration, the
portability fix and the gate numbers -- and marks the two sections that
recorded the falsified decision SUPERSEDED. The spec's SERVING
COMPLETION paragraph in .agents/specs/kolibri-1-cpu.md is corrected the
same way. Three engine-matrix.md api_server.cpp anchors shifted by the
+5-line RestoreToolSchemaOrder call are repaired (1354->1359,
1365->1370, 1614->1619); check-agent-record reports ANCHOR-ROT=0.

FOLLOWING_AGENTS_PROTOCOL

Following-Agents-Protocol: true
AI-Assisted: true
Assisted-by: AGENT:mistral/mistral-large-4 [maki]
lu-zero added a commit that referenced this pull request Oct 9, 2026
…oreToolSchemaOrder

The PR #3422 repair wired RestoreToolSchemaOrder into RunBatch::DispatchChat
as RestoreToolSchemaOrder(body, request), but the seam's first parameter is
the body TEXT (const std::string&, protocol.h:605) and body is the parsed
nlohmann::json object. The call compiles only through nlohmann's implicit
get<std::string>() conversion, which throws json.exception.type_error.302 on
every object body, and DispatchChat has no try/catch around it — so every
object-body /v1/chat/completions line through RunBatch throws. Clean-head
evidence at d93d9d4: test_openai_run_batch 3 of 7 cases throw
(test_run_batch.cpp:451/:507/:601); the repair's recorded 'run_batch 16/16'
gate counted assertions and missed the throwing cases (stale binary), so the
PR's CI lane would be red. ISSUE-LOCAL-01M4FR20MES4HQVBRWJBJ2AVCN.

A first in-flow attempt (body.dump()) is wrong and was reverted: body is the
key-sorted nlohmann::json, so its dump is already sorted and the order
restoration reads sorted text — the crash goes away but the seam no-ops and
the request document's tool-schema key order is lost on the batch path.

The original request text IS available: RunLine parses request_json and drops
it. RunLine now re-serializes the chat body (the line's top-level "body"
member) from an order-preserving nlohmann::ordered_json parse of the original
line text and threads it through DispatchChat(custom_id, body, body_json);
DispatchChat calls RestoreToolSchemaOrder(body_json, request).
ordered_json::dump keeps insertion order, so the re-read restores the request
document's schema key order into the prompt, matching the api_server and C-ABI
entry points, which already pass their raw body strings. The shared seam's
signature is unchanged.

Red-first: at the committed throwing call, 4 of 8 test_openai_run_batch cases
throw type_error.302 (the 3 pre-existing plus the new F1 case) while
assertions read 16/16; under the body.dump() mutation the 3 old cases pass but
the new F1 case fails both key-order assertions (seam no-op); at this fix
8/8 cases, 89/89 assertions pass.

FOLLOWING_AGENTS_PROTOCOL

Following-Agents-Protocol: true
AI-Assisted: true
Assisted-by: AGENT:mistral/mistral-large-4 [maki]
lu-zero added a commit that referenced this pull request Oct 9, 2026
…nd api_server entry points

Scoped re-review F1 of PR #3422 (ISSUE-LOCAL-01M4FR1D7RH7CN5X2R2CAR6N61):
RestoreToolSchemaOrder was detected by NO committed test — a no-op passed
every gate, and the batch call site was not merely unguarded but crashing
(ISSUE-LOCAL-01M4FR20MES4HQVBRWJBJ2AVCN). Two entry-point tests now close the
gap, one per affected dispatch:

- tests/vllm/entrypoints/openai/test_run_batch.cpp: a new case drives
  RunBatch::RunLine (the batch entry point) with a RAW JSONL line whose nested
  chat body is non-alphabetical at every level (tool wrapper, function, and
  parameters schema with zeta before alpha). The raw string is load-bearing:
  BatchLine() re-serializes through the key-sorting nlohmann::json and would
  destroy the order under test. The case asserts (a) no exception and a 200
  row — pinning the crash fix — and (b) the prompt seam receives the schema in
  the request document's key order and not the key-sorted form, captured via a
  CapturingToolsPrompt seam over the parsed tools (the capture pattern the
  api_server harness uses). Mutation-proven in both directions: with the
  committed throwing call restored the case throws type_error.302; with the
  call's argument mutated to body.dump() it fails both key-order assertions
  (the sorted dump no-ops the seam).
- tests/vllm/entrypoints/openai/test_api_server.cpp: a new case posts an
  unsorted-schema tools request through the PRODUCTION
  /v1/chat/completions dispatch (handle_chat_completions) with the real
  Qwen3.8 fixture template and asserts the rendered prompt bytes carry the
  document order. Mutation-proven: with RestoreToolSchemaOrder no-op'd at
  api_server.cpp:400 the case fails both key-order assertions.

Gates at this head: test_openai_run_batch 8/8 cases / 89/89 assertions;
test_openai_api_server 104/104 cases / 1521/1521 assertions (1517 pre-existing
+ 4 new). No existing assertion weakened: both additions are new TEST_CASEs;
the run_batch file's pre-existing cases are untouched.

FOLLOWING_AGENTS_PROTOCOL

Following-Agents-Protocol: true
AI-Assisted: true
Assisted-by: AGENT:mistral/mistral-large-4 [maki]
@lu-zero
lu-zero deleted the row/kolibri-serve branch October 9, 2026 11:09
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants