Repository navigation
Conversation
localai-org-maint-bot
left a comment
Collaborator
There was a problem hiding this comment.
@mudler @richiejp Two evidence corrections needed at 6f45aa7d9a4d4a5ea02aa274cbf329211b49a455 before treating this as the decode-lever decision:
- P2: whole-run attribution is presented as decode attribution. Section 2 explicitly says the 60-forward profile includes prefill and uses 35.24 s as its denominator, but then applies its 50%/25%/17% shares to a ~315 ms decode step. Section 1 itself reports 16.84 s prefill versus 19.65 s decode, so prefill is a substantial part of that profile. Different GEMM shapes and cold-cache behavior can give those phases different bottlenecks. Keep the table labeled whole-run, and capture separate prefill/decode counters before using these percentages to choose a decode lever.
- P2: the experiments do not establish that all bit-exact code levers are exhausted. The two-shape, single-thread fp8-direct microbenchmark falsifies that implementation in that regime. It does not prove that every GEMM scheduling/layout/threading change is impossible, nor does equality with the existing uncached decoder prove the decoder cannot improve. The spec and outcome currently turn these limited observations into a universal ceiling and declare cache budget the only remaining lever. Narrow those conclusions to the tested candidates and leave unmeasured alternatives open, consistent with AGENTS.md's no-ceiling rule. Preserve the negative measurements and budget table.
I reviewed all three changed files and checked the reported phase denominators. I did not rerun the native benchmark on the author's hardware; this is an evidence review, not acceptance of its performance claims. No default cache-budget change is proposed by this review.
lu-zero
added a commit
to lu-zero/vllm.cpp
that referenced
this pull request
Oct 9, 2026
…; split prefill/decode attribution Records-only repair of the blocking review on PR localai-org#3423 (two corrections, no product code, no new performance claim, W3 not re-run: the host is contended and no code changed). Correction 1, attribution: the whole-run stage table in docs/bench-evidence/kolibri1-perf-next-lever-20261008.md is now labeled WHOLE-RUN (60 forwards, 35.24 s denominator, prefill included) and its shares are no longer applied to a decode step; the old "per decode step ~315 ms: ~50% GEMV, ~25% decodes, ~17% attention" reading is replaced by a correction note. New §2b captures the phase split from two fresh bench legs at the production config (VLLM_CPP_CPU_THREADS=8 VT_KOLIBRI1_DEQUANT_CACHE_MB=16384, VT_KOLIBRI1_PROFILE=1, verified quiet): wall 35.79/35.56 s, prefill 15.93/15.68 s, decode 19.87/19.89 s, 3.171/3.168 tok/s, chains exact, counters byte-identical. The profiler cannot split by phase itself, but the bench's forward structure is deterministic (forward localai-org#1 prefill t=128, forwards localai-org#2..localai-org#64 decode t=1, reports every 10 forwards), so R60-R10 is exactly the 50 pure decode forwards localai-org#11..localai-org#60: the decode-phase table is exact (linear_gemm 59%, attention 27%+3%, lm_head 5%, steady-state cold misses 3%, ~302-304 ms per step profiled, 315 ms per step bench window; counters hits=61717 misses=783 evictions=783 decode_calls=62500, 15.7 misses/step steady state). The prefill-phase numbers are derived (R10 minus the nine warming decode forwards) and labeled as upper bounds; the exact prefill/decode counter split is not measurable with this profiler and is not claimed. 95.8% of the run's misses occur in the first 10 forwards: the ~26% whole-run cold-miss share is prefill/warmup-concentrated and was never decode evidence. Correction 2, ceiling wording: the universal claims are narrowed to the tested candidates everywhere they appeared (the evidence doc, the spec's NEXT-LEVER UNIT paragraph, the issue Resolution, and the PR body). The negative measurements (fp8-direct GEMV bit-exact 8x slower, two shapes single thread; no bit-exact change to the landed GEMM inner loop) and the budget table are preserved as-is; the text now says what was tested and leaves unmeasured alternatives open per AGENTS.md's no-ceiling rule: attention (~30% of the measured decode step, never adjudicated), GEMM scheduling/layout/threading variants, decoder improvements, and prefill-path levers. The cache budget remains the next DECISION (a measured policy knob), not the only remaining lever. Files: docs/bench-evidence/kolibri1-perf-next-lever-20261008.md, .agents/specs/kolibri-1-cpu.md, .agents/issues/MODEL-TEXT-kolibri-1-kolibri1-for-causal-lm/ISSUE-LOCAL-01M4EEX40G7NH571CR5GQY4CA1.md. FOLLOWING_AGENTS_PROTOCOL Following-Agents-Protocol: true AI-Assisted: true Assisted-by: AGENT:mistral/mistral-large-4 [maki]
The combined NEON-plus-cache run leaves moe_glue and linear_gemm at roughly 0.25 s each per forward and nothing has attributed the remaining decode wall by measurement on the current tree. File the profile-first issue that names the lever from a fresh attribution before any code lands. FOLLOWING_AGENTS_PROTOCOL Following-Agents-Protocol: true AI-Assisted: true Assisted-by: AGENT:zai/glm-5.3-flash [maki]
…the rest The baseline reproduces on bf7b654 at 3.17-3.21 decode tok/s (8 threads, 16 GiB), with the anchor, chain, and cache counters byte-identical to the records. The profile attribution corrects the recorded moe_glue reading: the scope nests the routed experts' GEMM and dequant, so its separable residual is small, and the wall is linear_gemm ~50 percent (at the order-preserving ALU roof), cold-miss whole-matrix decodes ~26 percent, attention ~17 percent. Two candidate levers are falsified by measurement. An fp8-direct GEMV (dequantize the packed weights in-register, halve the bytes) is bit-exact in a scratch micro-benchmark over the real expert shapes and 8x SLOWER, because the order-preserving contract forces one mul and one add per element per lane and the kernel is ALU-bound, not bandwidth-bound. The same roof measurement closes the bit-exact GEMM inner-loop family: fmla and bf16 dot break the row's contract. The one open lever is the dequant-cache production budget, which is the developer's policy decision, so this unit delivers numbers, not a default: OFF 1.23-1.26 tok/s, 8 GiB 2.99-3.13, 16 GiB 3.17-3.21, 32 GiB 3.36-3.37, with evictions 18321/13264/5043 and identical chains and counters per budget; thread scaling 4/8/16 = 2.11/3.2/3.9 tok/s. The W3 gate is green on this tree at 16 GiB: 900/900, ARGMAX 141/145 (4 near-tie, 0 hard), fingerprints 2.18646/26763, counters byte-identical, 92.7 GiB RSS. Resolves ISSUE-LOCAL-01M4EEX40G7NH571CR5GQY4CA1. FOLLOWING_AGENTS_PROTOCOL Following-Agents-Protocol: true AI-Assisted: true Assisted-by: AGENT:zai/glm-5.3-flash [maki]
…; split prefill/decode attribution Records-only repair of the blocking review on PR localai-org#3423 (two corrections, no product code, no new performance claim, W3 not re-run: the host is contended and no code changed). Correction 1, attribution: the whole-run stage table in docs/bench-evidence/kolibri1-perf-next-lever-20261008.md is now labeled WHOLE-RUN (60 forwards, 35.24 s denominator, prefill included) and its shares are no longer applied to a decode step; the old "per decode step ~315 ms: ~50% GEMV, ~25% decodes, ~17% attention" reading is replaced by a correction note. New §2b captures the phase split from two fresh bench legs at the production config (VLLM_CPP_CPU_THREADS=8 VT_KOLIBRI1_DEQUANT_CACHE_MB=16384, VT_KOLIBRI1_PROFILE=1, verified quiet): wall 35.79/35.56 s, prefill 15.93/15.68 s, decode 19.87/19.89 s, 3.171/3.168 tok/s, chains exact, counters byte-identical. The profiler cannot split by phase itself, but the bench's forward structure is deterministic (forward localai-org#1 prefill t=128, forwards localai-org#2..localai-org#64 decode t=1, reports every 10 forwards), so R60-R10 is exactly the 50 pure decode forwards localai-org#11..localai-org#60: the decode-phase table is exact (linear_gemm 59%, attention 27%+3%, lm_head 5%, steady-state cold misses 3%, ~302-304 ms per step profiled, 315 ms per step bench window; counters hits=61717 misses=783 evictions=783 decode_calls=62500, 15.7 misses/step steady state). The prefill-phase numbers are derived (R10 minus the nine warming decode forwards) and labeled as upper bounds; the exact prefill/decode counter split is not measurable with this profiler and is not claimed. 95.8% of the run's misses occur in the first 10 forwards: the ~26% whole-run cold-miss share is prefill/warmup-concentrated and was never decode evidence. Correction 2, ceiling wording: the universal claims are narrowed to the tested candidates everywhere they appeared (the evidence doc, the spec's NEXT-LEVER UNIT paragraph, the issue Resolution, and the PR body). The negative measurements (fp8-direct GEMV bit-exact 8x slower, two shapes single thread; no bit-exact change to the landed GEMM inner loop) and the budget table are preserved as-is; the text now says what was tested and leaves unmeasured alternatives open per AGENTS.md's no-ceiling rule: attention (~30% of the measured decode step, never adjudicated), GEMM scheduling/layout/threading variants, decoder improvements, and prefill-path levers. The cache budget remains the next DECISION (a measured policy knob), not the only remaining lever. Files: docs/bench-evidence/kolibri1-perf-next-lever-20261008.md, .agents/specs/kolibri-1-cpu.md, .agents/issues/MODEL-TEXT-kolibri-1-kolibri1-for-causal-lm/ISSUE-LOCAL-01M4EEX40G7NH571CR5GQY4CA1.md. FOLLOWING_AGENTS_PROTOCOL Following-Agents-Protocol: true AI-Assisted: true Assisted-by: AGENT:mistral/mistral-large-4 [maki]
The PR body was edited after creation (the 2026-10-09 review-repair section), which does not retrigger CI by itself; this empty commit does. No tree change. FOLLOWING_AGENTS_PROTOCOL Following-Agents-Protocol: true AI-Assisted: true Assisted-by: AGENT:mistral/mistral-large-4 [maki]
lu-zero
force-pushed
the
row/kolibri-perf-next
branch
from
October 9, 2026 10:58
f19d433 to
4c285f2
Compare
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
record(kolibri1): narrow the next-lever evidence to tested candidates; split prefill/decode attribution
What does this change do, in one paragraph?
It resolves the Kolibri-1 CPU next-performance-lever unit as a measured proposal and lands no product code. The TESTED next-lever candidates are closed by direct measurement: the remaining wall is attributed per stage and split into prefill and decode phases, two candidate families are falsified by direct measurement, and the unmeasured alternatives are named and left open per AGENTS.md's no-ceiling rule. The next DECISION is the dequant-cache production budget (
VT_KOLIBRI1_DEQUANT_CACHE_MB, currently default 0) — a measured policy knob that belongs to the developer, not the only remaining lever. The unit delivers the corrected attribution, the phase split, the falsifications with their numbers, the budget-sensitivity table (0/8/16/32 GiB), thread scaling (4/8/16), and the re-run correctness gates, so the budget default can be chosen from evidence.Why is it needed now?
Production stood at ~3.1 decode tok/s (8 threads, 16 GiB) after two landed levers, and the recorded stage profile attributed a ~0.25 s/forward "moe_glue residual" that suggested recoverable per-expert work remained (docs/bench-evidence/kolibri1-combined-neon-cache-20261008.md). Before naming a lever from that reading, the profile had to be re-attributed on the current tree — and it was wrong: the
moe_gluescope (src/vllm/model_executor/models/kolibri1_forward.cpp:369) wraps the routed experts'ExpertMlp→LinearBTcalls (kolibri1_forward.cpp:372-408), so its time is mostly the already-attributed nested linear_gemm and dequant_fp8_block work. Nothing was attributable until that was corrected.How is it implemented (file:line anchors)?
VT_KOLIBRI1_PROFILE=1, src/vllm/model_executor/models/kolibri1_forward.cpp:71-112). WHOLE-RUN final report at 60 forwards (1 prefill + 59 decode forwards, prefill included in every number), 16 GiB: linear_gemm 17.54 s, moe_glue 15.63 s (nested), dequant_fp8_block 9.31 s (= 18717 cold misses x ~0.50 ms whole-matrix pool decodes, kolibri1_dequant_cache.h:292-301), attn_core+rope 6.06 s, lm_head 1.57 s, norms 0.26 s. These are WHOLE-RUN shares of a 35.24 s denominator and are not applied to a single decode step.E4m3ToF32x4x block scale, one bf16 rounding viaF32x4ToBf16x4, kolibri1_fp8_dequant.h:60-91) insideBt16Neon's exact loop shape (src/vt/cpu/cpu_matmul_elem.cpp:155-189). It is BIT-EXACT (0 differing outputs vs the two-step reference on [512,2560] and [2560,512]) and 8x SLOWER (0.297 ms -> 2.475 ms; 0.275 ms -> 2.181 ms). The order-preserving contract (bit-exact across outputs, never along K) forbidsvfmaq/fmlaand K reassociation, so the kernel is ALU-bound at ~9 GB/s/thread, not bandwidth-bound: halving weight bytes buys nothing and the in-register decode pays ~6x the arithmetic on the ALU wall. This falsifies that implementation in that regime; it does not measure — and does not bound — other GEMM scheduling, layout, or threading variants, which remain open.What are the exact commands and observed numbers?
Baseline and budget legs (quiet windows, verified per leg: load < 12, no W3/qwen/bench process, > 140 GB available; every leg's chain byte-identical, md5
8de1463a…; counters byte-identical to the records per budget):Phase-split legs (2026-10-09 review repair; production config, profiling ON, verified quiet before each leg: load < 1, no test/bench/qwen process by exact
commname, > 200 GB available; both PASS, chain exact — anchor 109726, alternating 101807/109726; 60-forward counters byte-identical to the records; max RSS 92.3 GiB both legs):Decode phase (EXACT, forwards #11..#60 = R60-R10, 50 pure t=1 decode forwards; ms per decode step, leg #1 / #2): linear_gemm 178/182 (59%), moe_glue 79/80 (26%, nested), attn_core 82/82 (27%), attn_rope 10/10 (3%), lm_head 14/14 (5%), dequant_fp8_block 10/8 (3%, steady-state cold misses), norms 3/3 (1%); total_forward 302/304. Decode-window cache counters (identical in both legs): hits=61717 misses=783 evictions=783 decode_calls=62500 — 15.7 misses/step steady state (7.8/step in the last block #51..#60). Prefill phase (DERIVED, upper bounds; the bench prefill window 15.93/15.68 s is the lower anchor): linear_gemm ~7.0 s, dequant_fp8_block ~7.8-8.1 s (~half of the prefill forward), lm_head ~0.7 s, attention ~0.6 s. 95.8% of the run's 18717 misses occur within the first 10 forwards (17934); the pure-decode window #11..#60 accrues only 783.
The knee is not at 16 GiB: 16 -> 32 GiB still buys +5-6% by cutting evictions 2.6x, at +16 GiB resident RSS (W3 max RSS measured 92.7 GiB at 16 GiB on this tree). OFF costs 2.5x on every run that does not set the var. The choice is the developer's; the numbers above are the deliverable.
Thread scaling (16 GiB): T=4 2.108, T=8 3.17-3.21, T=16 3.854-3.908 tok/s (8 threads stays the operating point; 16 threads is a measured latency option).
Gates on this tree (bf7b654 base, no source change): test_kolibri1, test_kolibri1_dequant, test_kolibri1_dequant_cache, test_kolibri1_dequant_cache_default, test_kolibri1_moe_glue, test_kolibri1_w2 all SUCCESS;
test_kolibri1_w3at 16 GiB: 900/900 assertions, ARGMAX CHAIN 141/145 (4 flips: 4 near-tie, 0 hard), fingerprints 2.18646/26763 identical to the baseline, cache counters hits=348418 misses=519935 evictions=514482 decode_calls=868353 byte-identical to the 2026-10-07 reland record; decode bench anchor 109726 on every leg; budget invariance holds (identical CHAIN across all budgets). The Tenstorrent unit gates (test_kolibri1_tt*) are not runnable in this CPU-only build —VLLM_CPP_TENSTORRENT=ONrequires the TT-Metalium install tree that belongs to the active TT repair agent — and this change touches no TT-path file; that row's gates remain its own obligation.Measurement-discipline note, recorded because the brief's warning is load-bearing: two quiet-gate defects were hit and fixed mid-unit — a
pgrep -fgate matches its own shell's cmdline (fixed with the bracket form and exactcommnames), and an incompletely-killed monitor loop resumed later and raced live legs (one log truncated); every surviving leg was re-measured in a final clean pass. The 2026-10-09 phase-split legs re-verified quiet the same way (exactcommnames, load < 1).Review repair (2026-10-09) — answers the blocking review
Correction 1 (whole-run attribution presented as decode attribution): fixed. The stage table is labeled WHOLE-RUN (60 forwards, 35.24 s denominator, prefill included) and its percentages are no longer applied to a decode step; the old per-decode-step reading is replaced by a correction note in the evidence doc. Separate prefill/decode counters were captured from two fresh bench legs at the production config with profiling ON (numbers above and in the evidence doc §2b): the decode-phase table is exact (R60-R10 = the 50 pure decode forwards #11..#60, because the bench's forward structure is deterministic and the profiler reports every 10 forwards); the prefill-phase numbers are derived and labeled with their bias; the exact prefill/decode counter split is not measurable with this profiler and is stated as such. The measured decode step is ~59% linear_gemm, ~30% attention, ~5% lm_head, ~3% steady-state cold misses — the ~26% whole-run cold-miss share is prefill/warmup-concentrated (95.8% of misses in the first 10 forwards), so the whole-run shares were never decode evidence.
Correction 2 (the experiments do not establish that all bit-exact code levers are exhausted): fixed. The ceiling wording is narrowed everywhere it appeared as universal (the evidence doc, the spec's NEXT-LEVER UNIT paragraph, the issue Resolution, and this body). The negative measurements (fp8-direct GEMV bit-exact and 8x slower) and the budget table are preserved as-is; the text now states what was tested (two shapes, single thread, that implementation; the landed inner loop under the row's contract; the landed miss path against the existing uncached decoder) and explicitly leaves unmeasured alternatives open per AGENTS.md's no-ceiling rule: GEMM scheduling/layout/threading variants, decoder improvements, attention (~30% of the measured decode step, never adjudicated), and prefill-path levers (~44% of the whole-run wall). The recommendation still names the cache budget as the next DECISION (it is a policy knob, measured) — not as "the only remaining lever".
This repair is records-only: no product code, no new performance claim, and W3 was not re-run (the host is contended and no code changed; the 2026-10-08 W3 gate result stands).
Resolves ISSUE-LOCAL-01M4EEX40G7NH571CR5GQY4CA1.
Spec: .agents/specs/kolibri-1-cpu.md
## NowNEXT-LEVER UNIT paragraph (narrowed 2026-10-09).FOLLOWING_AGENTS_PROTOCOL
Following-Agents-Protocol: true
AI-Assisted: true
Assisted-by: AGENT:mistral/mistral-large-4 [maki]