Skip to content

record(kolibri1): the next decode lever is the cache budget; falsify the rest - #3423

Open
lu-zero wants to merge 4 commits into
localai-org:mainfrom
lu-zero:row/kolibri-perf-next
Open

lu-zero wants to merge 4 commits into
localai-org:mainfrom
lu-zero:row/kolibri-perf-next

Conversation

@lu-zero

@lu-zero lu-zero commented Oct 8, 2026 •

Copy link
Copy Markdown
Collaborator

record(kolibri1): narrow the next-lever evidence to tested candidates; split prefill/decode attribution

What does this change do, in one paragraph?

It resolves the Kolibri-1 CPU next-performance-lever unit as a measured proposal and lands no product code. The TESTED next-lever candidates are closed by direct measurement: the remaining wall is attributed per stage and split into prefill and decode phases, two candidate families are falsified by direct measurement, and the unmeasured alternatives are named and left open per AGENTS.md's no-ceiling rule. The next DECISION is the dequant-cache production budget (VT_KOLIBRI1_DEQUANT_CACHE_MB, currently default 0) — a measured policy knob that belongs to the developer, not the only remaining lever. The unit delivers the corrected attribution, the phase split, the falsifications with their numbers, the budget-sensitivity table (0/8/16/32 GiB), thread scaling (4/8/16), and the re-run correctness gates, so the budget default can be chosen from evidence.

Why is it needed now?

Production stood at ~3.1 decode tok/s (8 threads, 16 GiB) after two landed levers, and the recorded stage profile attributed a ~0.25 s/forward "moe_glue residual" that suggested recoverable per-expert work remained (docs/bench-evidence/kolibri1-combined-neon-cache-20261008.md). Before naming a lever from that reading, the profile had to be re-attributed on the current tree — and it was wrong: the moe_glue scope (src/vllm/model_executor/models/kolibri1_forward.cpp:369) wraps the routed experts' ExpertMlp → LinearBT calls (kolibri1_forward.cpp:372-408), so its time is mostly the already-attributed nested linear_gemm and dequant_fp8_block work. Nothing was attributable until that was corrected.

How is it implemented (file:line anchors)?

  • Evidence: docs/bench-evidence/kolibri1-perf-next-lever-20261008.md (baseline reproduction, whole-run attribution, the prefill/decode phase split, falsifications, budget table, thread scaling, gates).
  • Attribution uses the shipped stage profiler (VT_KOLIBRI1_PROFILE=1, src/vllm/model_executor/models/kolibri1_forward.cpp:71-112). WHOLE-RUN final report at 60 forwards (1 prefill + 59 decode forwards, prefill included in every number), 16 GiB: linear_gemm 17.54 s, moe_glue 15.63 s (nested), dequant_fp8_block 9.31 s (= 18717 cold misses x ~0.50 ms whole-matrix pool decodes, kolibri1_dequant_cache.h:292-301), attn_core+rope 6.06 s, lm_head 1.57 s, norms 0.26 s. These are WHOLE-RUN shares of a 35.24 s denominator and are not applied to a single decode step.
  • The phase split (2026-10-09 review repair, two fresh legs at the production config with profiling ON): the bench's forward structure is deterministic (forward perf(gdn): SANCTIONED Triton AOT fast-path for delta_h (VLLM_CPP_TRITON, default OFF) #1 is the prefill t=128, forwards perf(gdn): SANCTIONED Triton AOT fast-path for chunk_o + WU/WY (stacked on #1) #2..MiniMax-H3 W-FP4a GB10 leg: CUDA Marlin-W4A16 speed gate + fp4-resident e2e #64 are t=1 decode steps, the profiler reports every 10 forwards), so R60-R10 is exactly the 50 pure decode forwards fix(build): clear GCC 12 production warnings #11..perf(fa2): GB10 num_splits cap (gated-OFF) + MXFP4 parity TERMINAL residual #60. Measured DECODE phase (~302-304 ms/step profiled, 315 ms/step bench window): linear_gemm ~59%, attention ~30% (attn_core 27% + rope 3%), lm_head ~5%, steady-state cold misses ~3% (decode-window counters hits=61717 misses=783 evictions=783 decode_calls=62500, 15.7 misses/step steady state). 95.8% of the run's misses occur within the first 10 forwards — the ~26% whole-run cold-miss share is prefill/warmup-concentrated. The PREFILL phase (~15.7-15.9 s bench window) is about half cold-miss dequant; its stage split is derived (R10 minus the nine warming decode forwards) and labeled with its bias in the evidence doc; the exact prefill/decode counter split is not measurable with this profiler and is stated as such.
  • Falsification 1 — fp8-direct GEMV (TESTED candidate: one implementation, two real shapes, single thread): a scratch micro-benchmark (outside the tree, /tmp/kolibri-micro/) composes the landed NEON pieces (E4m3ToF32x4 x block scale, one bf16 rounding via F32x4ToBf16x4, kolibri1_fp8_dequant.h:60-91) inside Bt16Neon's exact loop shape (src/vt/cpu/cpu_matmul_elem.cpp:155-189). It is BIT-EXACT (0 differing outputs vs the two-step reference on [512,2560] and [2560,512]) and 8x SLOWER (0.297 ms -> 2.475 ms; 0.275 ms -> 2.181 ms). The order-preserving contract (bit-exact across outputs, never along K) forbids vfmaq/fmla and K reassociation, so the kernel is ALU-bound at ~9 GB/s/thread, not bandwidth-bound: halving weight bytes buys nothing and the in-register decode pays ~6x the arithmetic on the ALU wall. This falsifies that implementation in that regime; it does not measure — and does not bound — other GEMM scheduling, layout, or threading variants, which remain open.
  • Falsification 2 (tested candidates) — the landed GEMM inner loop: the same roof measurement closes every bit-exact change to the LANDED inner loop (the contract forbids fmla/dot); GEMM scheduling/layout/threading variants are unmeasured and remain open.
  • Falsification 3 (negative result recorded as such, tested path): the cold-miss path is already the uncached decode partitioned over the ONE pool (kolibri1_dequant_cache.h:292-301) under LRU with stable-identity keys — a miss cannot cost less than the decode it replaces. Equality with the existing uncached decoder does not prove that no decoder improvement exists; decoder variants are unmeasured and remain open.

What are the exact commands and observed numbers?

Baseline and budget legs (quiet windows, verified per leg: load < 12, no W3/qwen/bench process, > 140 GB available; every leg's chain byte-identical, md5 8de1463a…; counters byte-identical to the records per budget):

VLLM_CPP_CPU_THREADS=8 VT_KOLIBRI1_DEQUANT_CACHE_MB=<0|8192|16384|32768> \
  VT_KOLIBRI1_PROFILE=1 ./tests/test_kolibri1_decode_bench
budget decode tok/s hits/misses/evictions
0 (OFF) 1.230 / 1.258 —
8 GiB 2.992 / 3.134 69377/20497/18321
16 GiB 3.206 / 3.170 71157/18717/13264
32 GiB 3.369 / 3.363 72824/17050/5043

Phase-split legs (2026-10-09 review repair; production config, profiling ON, verified quiet before each leg: load < 1, no test/bench/qwen process by exact comm name, > 200 GB available; both PASS, chain exact — anchor 109726, alternating 101807/109726; 60-forward counters byte-identical to the records; max RSS 92.3 GiB both legs):

VLLM_CPP_CPU_THREADS=8 VT_KOLIBRI1_DEQUANT_CACHE_MB=16384 VT_KOLIBRI1_PROFILE=1 \
  ./tests/test_kolibri1_decode_bench
leg wall s prefill s decode s decode tok/s
#1 35.79 15.93 19.87 3.171
#2 35.56 15.68 19.89 3.168

Decode phase (EXACT, forwards #11..#60 = R60-R10, 50 pure t=1 decode forwards; ms per decode step, leg #1 / #2): linear_gemm 178/182 (59%), moe_glue 79/80 (26%, nested), attn_core 82/82 (27%), attn_rope 10/10 (3%), lm_head 14/14 (5%), dequant_fp8_block 10/8 (3%, steady-state cold misses), norms 3/3 (1%); total_forward 302/304. Decode-window cache counters (identical in both legs): hits=61717 misses=783 evictions=783 decode_calls=62500 — 15.7 misses/step steady state (7.8/step in the last block #51..#60). Prefill phase (DERIVED, upper bounds; the bench prefill window 15.93/15.68 s is the lower anchor): linear_gemm ~7.0 s, dequant_fp8_block ~7.8-8.1 s (~half of the prefill forward), lm_head ~0.7 s, attention ~0.6 s. 95.8% of the run's 18717 misses occur within the first 10 forwards (17934); the pure-decode window #11..#60 accrues only 783.

The knee is not at 16 GiB: 16 -> 32 GiB still buys +5-6% by cutting evictions 2.6x, at +16 GiB resident RSS (W3 max RSS measured 92.7 GiB at 16 GiB on this tree). OFF costs 2.5x on every run that does not set the var. The choice is the developer's; the numbers above are the deliverable.

Thread scaling (16 GiB): T=4 2.108, T=8 3.17-3.21, T=16 3.854-3.908 tok/s (8 threads stays the operating point; 16 threads is a measured latency option).

Gates on this tree (bf7b654 base, no source change): test_kolibri1, test_kolibri1_dequant, test_kolibri1_dequant_cache, test_kolibri1_dequant_cache_default, test_kolibri1_moe_glue, test_kolibri1_w2 all SUCCESS; test_kolibri1_w3 at 16 GiB: 900/900 assertions, ARGMAX CHAIN 141/145 (4 flips: 4 near-tie, 0 hard), fingerprints 2.18646/26763 identical to the baseline, cache counters hits=348418 misses=519935 evictions=514482 decode_calls=868353 byte-identical to the 2026-10-07 reland record; decode bench anchor 109726 on every leg; budget invariance holds (identical CHAIN across all budgets). The Tenstorrent unit gates (test_kolibri1_tt*) are not runnable in this CPU-only build — VLLM_CPP_TENSTORRENT=ON requires the TT-Metalium install tree that belongs to the active TT repair agent — and this change touches no TT-path file; that row's gates remain its own obligation.

Measurement-discipline note, recorded because the brief's warning is load-bearing: two quiet-gate defects were hit and fixed mid-unit — a pgrep -f gate matches its own shell's cmdline (fixed with the bracket form and exact comm names), and an incompletely-killed monitor loop resumed later and raced live legs (one log truncated); every surviving leg was re-measured in a final clean pass. The 2026-10-09 phase-split legs re-verified quiet the same way (exact comm names, load < 1).

Review repair (2026-10-09) — answers the blocking review

Correction 1 (whole-run attribution presented as decode attribution): fixed. The stage table is labeled WHOLE-RUN (60 forwards, 35.24 s denominator, prefill included) and its percentages are no longer applied to a decode step; the old per-decode-step reading is replaced by a correction note in the evidence doc. Separate prefill/decode counters were captured from two fresh bench legs at the production config with profiling ON (numbers above and in the evidence doc §2b): the decode-phase table is exact (R60-R10 = the 50 pure decode forwards #11..#60, because the bench's forward structure is deterministic and the profiler reports every 10 forwards); the prefill-phase numbers are derived and labeled with their bias; the exact prefill/decode counter split is not measurable with this profiler and is stated as such. The measured decode step is ~59% linear_gemm, ~30% attention, ~5% lm_head, ~3% steady-state cold misses — the ~26% whole-run cold-miss share is prefill/warmup-concentrated (95.8% of misses in the first 10 forwards), so the whole-run shares were never decode evidence.

Correction 2 (the experiments do not establish that all bit-exact code levers are exhausted): fixed. The ceiling wording is narrowed everywhere it appeared as universal (the evidence doc, the spec's NEXT-LEVER UNIT paragraph, the issue Resolution, and this body). The negative measurements (fp8-direct GEMV bit-exact and 8x slower) and the budget table are preserved as-is; the text now states what was tested (two shapes, single thread, that implementation; the landed inner loop under the row's contract; the landed miss path against the existing uncached decoder) and explicitly leaves unmeasured alternatives open per AGENTS.md's no-ceiling rule: GEMM scheduling/layout/threading variants, decoder improvements, attention (~30% of the measured decode step, never adjudicated), and prefill-path levers (~44% of the whole-run wall). The recommendation still names the cache budget as the next DECISION (it is a policy knob, measured) — not as "the only remaining lever".

This repair is records-only: no product code, no new performance claim, and W3 was not re-run (the host is contended and no code changed; the 2026-10-08 W3 gate result stands).

Resolves ISSUE-LOCAL-01M4EEX40G7NH571CR5GQY4CA1.

Spec: .agents/specs/kolibri-1-cpu.md ## Now NEXT-LEVER UNIT paragraph (narrowed 2026-10-09).

FOLLOWING_AGENTS_PROTOCOL

Following-Agents-Protocol: true
AI-Assisted: true
Assisted-by: AGENT:mistral/mistral-large-4 [maki]

@localai-org-maint-bot localai-org-maint-bot left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

@mudler @richiejp Two evidence corrections needed at 6f45aa7d9a4d4a5ea02aa274cbf329211b49a455 before treating this as the decode-lever decision:

  • P2: whole-run attribution is presented as decode attribution. Section 2 explicitly says the 60-forward profile includes prefill and uses 35.24 s as its denominator, but then applies its 50%/25%/17% shares to a ~315 ms decode step. Section 1 itself reports 16.84 s prefill versus 19.65 s decode, so prefill is a substantial part of that profile. Different GEMM shapes and cold-cache behavior can give those phases different bottlenecks. Keep the table labeled whole-run, and capture separate prefill/decode counters before using these percentages to choose a decode lever.
  • P2: the experiments do not establish that all bit-exact code levers are exhausted. The two-shape, single-thread fp8-direct microbenchmark falsifies that implementation in that regime. It does not prove that every GEMM scheduling/layout/threading change is impossible, nor does equality with the existing uncached decoder prove the decoder cannot improve. The spec and outcome currently turn these limited observations into a universal ceiling and declare cache budget the only remaining lever. Narrow those conclusions to the tested candidates and leave unmeasured alternatives open, consistent with AGENTS.md's no-ceiling rule. Preserve the negative measurements and budget table.

I reviewed all three changed files and checked the reported phase denominators. I did not rerun the native benchmark on the author's hardware; this is an evidence review, not acceptance of its performance claims. No default cache-budget change is proposed by this review.

lu-zero added a commit to lu-zero/vllm.cpp that referenced this pull request Oct 9, 2026
…; split prefill/decode attribution

Records-only repair of the blocking review on PR localai-org#3423 (two corrections,
no product code, no new performance claim, W3 not re-run: the host is
contended and no code changed).

Correction 1, attribution: the whole-run stage table in
docs/bench-evidence/kolibri1-perf-next-lever-20261008.md is now labeled
WHOLE-RUN (60 forwards, 35.24 s denominator, prefill included) and its
shares are no longer applied to a decode step; the old "per decode step
~315 ms: ~50% GEMV, ~25% decodes, ~17% attention" reading is replaced by
a correction note. New §2b captures the phase split from two fresh bench
legs at the production config (VLLM_CPP_CPU_THREADS=8
VT_KOLIBRI1_DEQUANT_CACHE_MB=16384, VT_KOLIBRI1_PROFILE=1, verified quiet):
wall 35.79/35.56 s, prefill 15.93/15.68 s, decode 19.87/19.89 s,
3.171/3.168 tok/s, chains exact, counters byte-identical. The profiler
cannot split by phase itself, but the bench's forward structure is
deterministic (forward localai-org#1 prefill t=128, forwards localai-org#2..localai-org#64 decode t=1,
reports every 10 forwards), so R60-R10 is exactly the 50 pure decode
forwards localai-org#11..localai-org#60: the decode-phase table is exact (linear_gemm 59%,
attention 27%+3%, lm_head 5%, steady-state cold misses 3%, ~302-304 ms per
step profiled, 315 ms per step bench window; counters hits=61717
misses=783 evictions=783 decode_calls=62500, 15.7 misses/step steady
state). The prefill-phase numbers are derived (R10 minus the nine warming
decode forwards) and labeled as upper bounds; the exact prefill/decode
counter split is not measurable with this profiler and is not claimed.
95.8% of the run's misses occur in the first 10 forwards: the ~26%
whole-run cold-miss share is prefill/warmup-concentrated and was never
decode evidence.

Correction 2, ceiling wording: the universal claims are narrowed to the
tested candidates everywhere they appeared (the evidence doc, the spec's
NEXT-LEVER UNIT paragraph, the issue Resolution, and the PR body). The
negative measurements (fp8-direct GEMV bit-exact 8x slower, two shapes
single thread; no bit-exact change to the landed GEMM inner loop) and the
budget table are preserved as-is; the text now says what was tested and
leaves unmeasured alternatives open per AGENTS.md's no-ceiling rule:
attention (~30% of the measured decode step, never adjudicated), GEMM
scheduling/layout/threading variants, decoder improvements, and
prefill-path levers. The cache budget remains the next DECISION (a
measured policy knob), not the only remaining lever.

Files: docs/bench-evidence/kolibri1-perf-next-lever-20261008.md,
.agents/specs/kolibri-1-cpu.md,
.agents/issues/MODEL-TEXT-kolibri-1-kolibri1-for-causal-lm/ISSUE-LOCAL-01M4EEX40G7NH571CR5GQY4CA1.md.

FOLLOWING_AGENTS_PROTOCOL

Following-Agents-Protocol: true
AI-Assisted: true
Assisted-by: AGENT:mistral/mistral-large-4 [maki]
The combined NEON-plus-cache run leaves moe_glue and linear_gemm at
roughly 0.25 s each per forward and nothing has attributed the
remaining decode wall by measurement on the current tree. File the
profile-first issue that names the lever from a fresh attribution
before any code lands.

FOLLOWING_AGENTS_PROTOCOL

Following-Agents-Protocol: true
AI-Assisted: true
Assisted-by: AGENT:zai/glm-5.3-flash [maki]
…the rest

The baseline reproduces on bf7b654 at 3.17-3.21 decode tok/s (8 threads,
16 GiB), with the anchor, chain, and cache counters byte-identical to the
records. The profile attribution corrects the recorded moe_glue reading: the
scope nests the routed experts' GEMM and dequant, so its separable residual
is small, and the wall is linear_gemm ~50 percent (at the order-preserving
ALU roof), cold-miss whole-matrix decodes ~26 percent, attention ~17 percent.

Two candidate levers are falsified by measurement. An fp8-direct GEMV
(dequantize the packed weights in-register, halve the bytes) is bit-exact in
a scratch micro-benchmark over the real expert shapes and 8x SLOWER, because
the order-preserving contract forces one mul and one add per element per lane
and the kernel is ALU-bound, not bandwidth-bound. The same roof measurement
closes the bit-exact GEMM inner-loop family: fmla and bf16 dot break the
row's contract.

The one open lever is the dequant-cache production budget, which is the
developer's policy decision, so this unit delivers numbers, not a default:
OFF 1.23-1.26 tok/s, 8 GiB 2.99-3.13, 16 GiB 3.17-3.21, 32 GiB 3.36-3.37,
with evictions 18321/13264/5043 and identical chains and counters per
budget; thread scaling 4/8/16 = 2.11/3.2/3.9 tok/s. The W3 gate is green on
this tree at 16 GiB: 900/900, ARGMAX 141/145 (4 near-tie, 0 hard),
fingerprints 2.18646/26763, counters byte-identical, 92.7 GiB RSS.

Resolves ISSUE-LOCAL-01M4EEX40G7NH571CR5GQY4CA1.

FOLLOWING_AGENTS_PROTOCOL

Following-Agents-Protocol: true
AI-Assisted: true
Assisted-by: AGENT:zai/glm-5.3-flash [maki]
…; split prefill/decode attribution

Records-only repair of the blocking review on PR localai-org#3423 (two corrections,
no product code, no new performance claim, W3 not re-run: the host is
contended and no code changed).

Correction 1, attribution: the whole-run stage table in
docs/bench-evidence/kolibri1-perf-next-lever-20261008.md is now labeled
WHOLE-RUN (60 forwards, 35.24 s denominator, prefill included) and its
shares are no longer applied to a decode step; the old "per decode step
~315 ms: ~50% GEMV, ~25% decodes, ~17% attention" reading is replaced by
a correction note. New §2b captures the phase split from two fresh bench
legs at the production config (VLLM_CPP_CPU_THREADS=8
VT_KOLIBRI1_DEQUANT_CACHE_MB=16384, VT_KOLIBRI1_PROFILE=1, verified quiet):
wall 35.79/35.56 s, prefill 15.93/15.68 s, decode 19.87/19.89 s,
3.171/3.168 tok/s, chains exact, counters byte-identical. The profiler
cannot split by phase itself, but the bench's forward structure is
deterministic (forward localai-org#1 prefill t=128, forwards localai-org#2..localai-org#64 decode t=1,
reports every 10 forwards), so R60-R10 is exactly the 50 pure decode
forwards localai-org#11..localai-org#60: the decode-phase table is exact (linear_gemm 59%,
attention 27%+3%, lm_head 5%, steady-state cold misses 3%, ~302-304 ms per
step profiled, 315 ms per step bench window; counters hits=61717
misses=783 evictions=783 decode_calls=62500, 15.7 misses/step steady
state). The prefill-phase numbers are derived (R10 minus the nine warming
decode forwards) and labeled as upper bounds; the exact prefill/decode
counter split is not measurable with this profiler and is not claimed.
95.8% of the run's misses occur in the first 10 forwards: the ~26%
whole-run cold-miss share is prefill/warmup-concentrated and was never
decode evidence.

Correction 2, ceiling wording: the universal claims are narrowed to the
tested candidates everywhere they appeared (the evidence doc, the spec's
NEXT-LEVER UNIT paragraph, the issue Resolution, and the PR body). The
negative measurements (fp8-direct GEMV bit-exact 8x slower, two shapes
single thread; no bit-exact change to the landed GEMM inner loop) and the
budget table are preserved as-is; the text now says what was tested and
leaves unmeasured alternatives open per AGENTS.md's no-ceiling rule:
attention (~30% of the measured decode step, never adjudicated), GEMM
scheduling/layout/threading variants, decoder improvements, and
prefill-path levers. The cache budget remains the next DECISION (a
measured policy knob), not the only remaining lever.

Files: docs/bench-evidence/kolibri1-perf-next-lever-20261008.md,
.agents/specs/kolibri-1-cpu.md,
.agents/issues/MODEL-TEXT-kolibri-1-kolibri1-for-causal-lm/ISSUE-LOCAL-01M4EEX40G7NH571CR5GQY4CA1.md.

FOLLOWING_AGENTS_PROTOCOL

Following-Agents-Protocol: true
AI-Assisted: true
Assisted-by: AGENT:mistral/mistral-large-4 [maki]
The PR body was edited after creation (the 2026-10-09 review-repair
section), which does not retrigger CI by itself; this empty commit does.
No tree change.

FOLLOWING_AGENTS_PROTOCOL

Following-Agents-Protocol: true
AI-Assisted: true
Assisted-by: AGENT:mistral/mistral-large-4 [maki]
@lu-zero
lu-zero force-pushed the row/kolibri-perf-next branch from f19d433 to 4c285f2 Compare October 9, 2026 10:58

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants