Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
6 changes: 3 additions & 3 deletions .agents/engine-matrix.md

Large diffs are not rendered by default.

Original file line number Diff line number Diff line change
@@ -0,0 +1,23 @@
ID: ISSUE-LOCAL-01M4EF3R0H2H3FN5NA0NAB0S62
Title: kolibri1 serving completion: chat template rendering plus kolibri1 reasoning and tool-call parsers
Row: MODEL-TEXT-kolibri-1-kolibri1-for-causal-lm
State: CLOSED
Kind: enhancement
GitHub: -
Mirror: PENDING
Availability: FULL
Created: 2026-10-08
Updated: 2026-10-09
Closed: 2026-10-08

## Problem

The kolibri1 CPU arm reaches the model forward but not serving: the OpenAI chat path has no kolibri1 reasoning parser, no kolibri1 tool-parser alias, and no template-detection rows, so the model-author plugin's serving recipe (aleph-alpha-inference pin 049a6a7bd240: reasoning parser kolibri1 = Qwen3 grammar with the template's thinking switch derived from chat_template_kwargs reasoning_effort/enable_thinking; tool parser kolibri1 = the Hermes <tool_call> format; chat template shipped in tokenizer_config.json) cannot be served end to end. Scope: port the thinking_enabled switch and the engine-backed kolibri1 reasoning adapter, alias the kolibri1 tool parser to hermes, add template-marker detection rows, gate template rendering against jinja2 reference outputs on kolibri1 goldens. No tokenizer change, no behavior change to other models' parsers.

## Resolution

Landed on row/kolibri-serve 2026-10-08 (commits 3c0fcc164, 4aade352b); see docs/bench-evidence/kolibri1-serve-20261008.md.

REOPENED IN EFFECT BY REVIEW 2026-10-09: the PR #3422 review (localai-org-maint-bot) blocked the landing with two findings against the serving-completion change. P1: the adapter's global tojson override sorted object keys for every model, but the pinned transformers 5.14.1 renderer overrides Jinja's tojson with sort_keys=False / ensure_ascii=False / no HTML escaping (utils/chat_template_utils.py:481), so the override rendered prompt bytes no serving reference produces; the reference fixtures had been captured with plain jinja2, which certifies the same incorrect reference (sorted keys, HTML escaped). P2: tests/vllm/entrypoints/test_kolibri1_chat_template.cpp loaded /mnt/models/Aleph-Alpha/Kolibri-1/tokenizer_config.json in both the rendering and detection cases, so a clean checkout could not run the gate, and the fixture generator hard-coded its output under /tmp/vllm-kolibri-serve.

REPAIR LANDED on row/kolibri-serve 2026-10-09 (review-repair commits): the pinned renderer was MEASURED first (probe through transformers 5.14.1 render_jinja_template over an insertion-ordered {"z":1,"a":2}, nested unsorted tool schemas, Unicode/HTML leaves): the default keeps insertion key order, raw UTF-8, no HTML escaping. chat_template.cpp's tojson is now a port of the pinned filter's full signature (ensure_ascii/indent/separators/sort_keys, CPython json.dumps semantics) instead of the sorted-dump override; FunctionDefinition::parameters is order-preserving (ordered_json) with RestoreToolSchemaOrder re-reading tools from an order-preserving body parse at every chat entry point (api_server, C ABI, run_batch), and BuildTools mirrors pinned vLLM's model_dump tool shape (description/parameters present as null when absent). The kolibri1 references were REGENERATED through the pinned renderer (20 scenarios, including unsorted nested tool schemas, Unicode/HTML content, a description-less tool and unsorted historical tool-call arguments) by a generator that asserts transformers==5.14.1 and writes in-tree; the template input is committed at tests/fixtures/kolibri1-chat-template-tokenizer_config.json (sha256 9ba35d4bd6baa26b66aa75d03a922dfee98b16bb1fa37481b195d247267b0f97) and the test loads it from the fixture directory with no /mnt path. Red-first: 6 of 20 scenarios failed under the corrected reference before the fix (with_tools, with_tools_thinking_off, with_tools_unsorted_unicode, with_tools_unsorted_unicode_thinking_off, with_tools_minimal_no_description, assistant_tool_call_unsorted_arguments). Gates after the repair: test_reasoning_kolibri1 43, test_tool_parser_kolibri1 19, test_kolibri1_chat_template 61, test_chat_template 204 (196 pre-existing plus 8 new non-kolibri tojson guard assertions), test_reasoning_parser_detect 75, test_tool_parser_detect 361, test_reasoning_qwen3 164, test_openai_tool_parsers 64, test_kolibri1 27/234, test_kolibri1_decode_bench anchor 109726, all green; serving/protocol suites (test_openai_serving 1365, test_openai_api_server 1517, test_openai_conformance 252, test_parser_engine_assembly 5038, test_openai_run_batch 16) and the parameters-reading tool-parser suites green; test_capi v27 (gliner fixture load) remains the pre-existing host failure recorded in the evidence doc. W3 not rerun: no forward change. Evidence: docs/bench-evidence/kolibri1-serve-20261008.md "Review repair".
Original file line number Diff line number Diff line change
@@ -0,0 +1,29 @@
ID: ISSUE-LOCAL-01M4FR1D7RH7CN5X2R2CAR6N61
Title: Scoped re-review of PR #3422's repair: F1 the ordered-parameters seam (RestoreToolSchemaOrder) is detected by no committed test; F2 the 'byte-stable for already-ordered schemas' guard uses a non-alphabetical fixture
Row: MODEL-TEXT-kolibri-1-kolibri1-for-causal-lm
State: CLOSED
Kind: bug
GitHub: -
Mirror: PENDING
Availability: FULL
Created: 2026-10-09
Updated: 2026-10-09
Closed: 2026-10-09

## Problem

Two findings from the scoped re-review of the PR #3422 review repair (branch row/kolibri-serve, head d93d9d492). F1 (coverage gap): RestoreToolSchemaOrder — the ordered-parameters seam called at src/vllm/entrypoints/openai/api_server.cpp:400 (handle_chat_completions), src/capi/vllm_c.cpp:518 (ParseChatRequest) and src/vllm/entrypoints/openai/run_batch.cpp:110 (DispatchChat) — is detected by NO committed test: a no-op passes every gate. The kolibri1 chat-template gate drives apply_chat_template directly over C++-constructed tools and never crosses the entry-point parse, so the seam that re-reads tools order-preserving after the key-sorting nlohmann::json parse is unguarded. OWED: a test that posts an unsorted-schema tools request through a chat entry point and asserts the rendered prompt bytes keep the request document's key order. F2 (mislabeled guard case): the 'byte-stable for already-ordered schemas' case in tests/vllm/entrypoints/test_chat_template.cpp uses an AlphabeticalTool fixture that is NOT fully alphabetical — the rendered tool is insertion-ordered non-alphabetically at type/function, name/description/parameters and type/properties/required — so under a sorted mutation the case goes red too and the 'alphabetical-schema byte-identity before/after' sub-claim is not demonstrated. OWED: make the fixture genuinely alphabetical at every level (or relabel the case to what it actually pins; prefer genuinely alphabetical so the byte-identity-before/after claim is real) and correct the comment.

The same re-review surfaced a CONFIRMED crash the repair introduced on the batch path (RestoreToolSchemaOrder(body, request) throws type_error.302 on every object body) — filed separately as ISSUE-LOCAL-01M4FR20MES4HQVBRWJBJ2AVCN, since the crash and the coverage gap are distinct defects with distinct fixes.

## Resolution

Both findings landed on row/kolibri-serve 2026-10-09 (test commit 2f535c002a1b15c8717a53a5e8103258ad5fd229, landed with the F2 guard and
these records in the follow-up records commit — see commit bodies).

F1 — two entry-point tests now detect a no-op RestoreToolSchemaOrder, one per affected entry point:
- tests/vllm/entrypoints/openai/test_run_batch.cpp, new case 'run_batch: a tools line keeps the document's schema key order in the prompt': drives RunBatch::RunLine (the batch entry point) with a RAW JSONL line whose nested chat body is non-alphabetical at every level (built as a raw string — BatchLine() would re-serialize through the key-sorting nlohmann::json and destroy the order under test). It asserts (a) no exception and a 200 row — the crash fix, ISSUE-LOCAL-01M4FR20MES4HQVBRWJBJ2AVCN — and (b) the prompt seam receives the schema in the request document's key order (captured via a CapturingToolsPrompt seam over the parsed tools, the same capture pattern the api_server harness uses), and NOT the key-sorted form. Mutation-proven in both directions: with the committed throwing call restored the case throws type_error.302 (RED); with the call's argument mutated to body.dump() — the reverted first repair attempt — the case fails both key-order assertions (the sorted dump no-ops the seam); at the fix it passes.
- tests/vllm/entrypoints/openai/test_api_server.cpp, new case 'api_server: an unsorted tool schema keeps its document key order in the rendered prompt': posts an unsorted-schema tools request through the PRODUCTION /v1/chat/completions dispatch (handle_chat_completions) with the real Qwen3.8 fixture template and asserts the rendered prompt bytes carry the document order. Mutation-proven: with RestoreToolSchemaOrder no-op'd at api_server.cpp:400 the case fails both key-order assertions (RED); restored, green.
Gates: test_openai_run_batch 8/8 cases / 89/89 assertions; test_openai_api_server 104/104 cases / 1521/1521 assertions (1517 pre-existing + 4 new).

F2 — the 'byte-stable for already-ordered schemas' case (tests/vllm/entrypoints/test_chat_template.cpp) now uses a genuinely alphabetical fixture at every level: the schema is {"properties": {"city": ...}, "required": [...], "type": "object"} (properties < required < type; the single-property schema needs no order), and the case renders the SCHEMA alone ({{ tools[0].function.parameters | tojson }}) because the tool wrapper's field order (type/function, name/description/parameters) is the pinned model_dump order, which is not alphabetical — so the byte-identity claim is now real and the comment says so. Mutation-proven: with the tojson default mutated to sorted keys (the override the repair removed), the insertion-order case goes RED while this byte-stable case stays GREEN — exactly the before/after agreement it claims to pin. Gate: test_chat_template 45/45 cases / 204/204 assertions.
Original file line number Diff line number Diff line change
@@ -0,0 +1,81 @@
ID: ISSUE-LOCAL-01M4FR20MES4HQVBRWJBJ2AVCN
Title: run_batch: RestoreToolSchemaOrder(body, request) passes a nlohmann::json where a string is expected — every chat batch line throws json.exception.type_error.302 (test_openai_run_batch red at the PR #3422 repair head)
Row: MODEL-TEXT-kolibri-1-kolibri1-for-causal-lm
State: CLOSED
Kind: bug
GitHub: -
Mirror: PENDING
Availability: FULL
Created: 2026-10-09
Updated: 2026-10-09
Closed: 2026-10-09

## Problem

Found by the operator's clean-head re-verification of PR #3422 (branch
row/kolibri-serve, head d93d9d492), after the scoped re-review F1 finding
(ISSUE-LOCAL-01M4FR1D7RH7CN5X2R2CAR6N61) showed the ordered-parameters seam
was detected by no committed test. The repair commit 5983c6918 wired
RestoreToolSchemaOrder into RunBatch::DispatchChat as
RestoreToolSchemaOrder(body, request) (src/vllm/entrypoints/openai/run_batch.cpp:110),
but the function's first parameter is const std::string&
(include/vllm/entrypoints/openai/protocol.h:605) and body is a
const nlohmann::json&. The call compiles only through nlohmann's implicit
operator ValueType() -> get<std::string>(), which throws
json.exception.type_error.302 ('type must be string, but is object') for every
object body — and DispatchChat has no try/catch around the call, so EVERY
/v1/chat/completions line dispatched through RunBatch throws an uncaught
exception. Operator's clean-head evidence at d93d9d492:
/tmp/build-kolibri-serve/tests/test_openai_run_batch reports 'test cases: 7 |
4 passed | 3 failed', 'assertions: 16 | 16 passed', Status: FAILURE — the three
chat-dispatch cases (test_run_batch.cpp:451, :507, :601 at that head) all throw
type_error.302. The repair's recorded gate 'test_openai_run_batch 16'
(docs/bench-evidence/kolibri1-serve-20261008.md "Review repair"; the PR #3422
body; commit 5983c6918's message) counted the ASSERTION total (16/16) and
missed the three throwing test cases — a stale binary, so the PR's CI lane
(build-test-cpu runs the full ctest) would be red. The other two call sites
pass the raw body string and are correct (api_server.cpp:400
handle_chat_completions(request_body); vllm_c.cpp:518
ParseChatRequest(request_json)).

A first in-flow repair attempt (passing body.dump()) is WRONG and was
reverted: body is the key-sorted nlohmann::json (std::map), so its dump is
already alphabetically sorted and the order restoration re-reads sorted text —
the crash goes away but the seam no-ops, and the request document's tool-schema
key order is lost on the batch path (the pinned renderer dumps schemas into
the prompt with sort_keys=False, so the prompt bytes would silently change
order). Mutation-proven below.

## Resolution

Landed on row/kolibri-serve 2026-10-09 (fix commit a6679a82a3e6edc7bbc71a09585a5944bcd91810; test commit
2f535c002a1b15c8717a53a5e8103258ad5fd229 — see commit bodies). The original request text IS available:
RunBatch::RunLine(const std::string& request_json) parses it at run_batch.cpp:159
and drops the text. The fix threads it through: RunLine re-serializes the chat
body — the line's top-level "body" member — from an ORDER-PRESERVING
(nlohmann::ordered_json) parse of the original line text, and passes it to
DispatchChat(custom_id, body, body_json); DispatchChat calls
RestoreToolSchemaOrder(body_json, request). ordered_json::dump keeps insertion
order, so body_json carries the request document's key order and the re-read
restores it into request.tools, exactly as the api_server and C-ABI entry
points already did with their raw body strings. The shared seam's signature is
unchanged.

Red-first and mutation evidence (all at /tmp/build-kolibri-serve, this host):
- Committed throwing call (head d93d9d492 plus the new F1 test): 4 of 8 cases
throw type_error.302 — the 3 pre-existing cases (test_run_batch.cpp:484/:540/:634
in the edited file) plus the new F1 case; assertions read '16 | 16 passed'
while 4 cases fail — the stale-binary signature.
- body.dump() mutation (the reverted first attempt): the 3 pre-existing cases
go GREEN (crash silenced) but the new F1 case fails BOTH key-order assertions
— the prompt carries the key-sorted schema, proving the seam no-ops.
- The fix: test_openai_run_batch 8/8 cases, 89/89 assertions, Status: SUCCESS;
the F1 case's prompt carries the document-ordered schema.
Full battery re-verified at the fixed head: test_openai_api_server 1521,
test_openai_serving 1365, test_chat_template 204, test_kolibri1_chat_template
61, test_reasoning_kolibri1 43, test_tool_parser_kolibri1 19,
test_reasoning_parser_detect 75, test_tool_parser_detect 361, test_kolibri1
27/234, test_kolibri1_decode_bench anchor 109726 — all PASS; W3 not rerun (no
forward change). The repair's 'test_openai_run_batch 16' gate line in
docs/bench-evidence/kolibri1-serve-20261008.md is falsified by this evidence and
corrected in that document's re-review section.
Loading
Loading