Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
Original file line number Diff line number Diff line change
@@ -0,0 +1,53 @@
ID: ISSUE-LOCAL-01M4E22DM790W0E69M07D5XA9D
Title: Kolibri-1 TT B2b-i completion: dense-resident device forward — one greedy decode on the P150
Row: MODEL-TEXT-kolibri-1-tenstorrent
State: OPEN
Kind: enhancement
GitHub: -
Mirror: PENDING
Availability: FULL
Created: 2026-10-08
Updated: 2026-10-08
Closed: -

## Problem

The B2b-i bring-up slice landed 2026-10-08 (merged as cf6258c76 + b52c0baeb): the TT build links against the pin source via the fresh /tmp/pin-build lib64, the resident non-expert slice (3,311,163,520 B = 3.084 GiB, 753 tensors) stages on the P150 with byte-exact readback of all 1103 operands, and one device op (embedding bit-exact + layer-0 q_proj GEMM within its stated envelope) is verified against the CPU row. What remains is the B2b-i completion condition named in .agents/specs/kolibri-tt.md ### B2 scope — B2b addendum: the dense-resident device forward running entirely on device with the routed-expert tier absent, completing one greedy decode of a golden prompt on the card. Concretely (addendum lines 360-382): (1) attention — the hybrid geometry as staged, 40 sliding-window layers at window 513 and 10 full-attention layers with full RNoPE (no rope tables on the full group), two-group KV per the wave-A design, per-head q/k RMS norms inside the attention path; (2) the four sandwich norms (input_layernorm, post_attn_norm, post_attention_layernorm, post_ffn_norm) per the landed CPU row (.agents/specs/kolibri-1-cpu.md); (3) the router — bf16 [384,2560] gate, f32 e_score_correction_bias, sigmoid-logit-add scoring, top-6-of-384, used in this slice only for the shared expert plus a nameable unimplemented-routed-expert refusal; (4) embed + untied lm_head staged bf16; (5) on-device sampling through the landed decode seam (ModelRegistry::Forward, dense_attn::AttnBlock) where the TT backend provides it; no streaming, no slot pool. Gates: device-free host-side tests first (red-first, manifest-driven style in tests/vllm/models/), then the token gate vs the CPU golden chains (W3 methodology: 141/145 argmax positions, the 4 known flips adjudicated inside the 2.5-nat band, 0 hard flips allowed — any flip outside the band is a gate failure), then the production bench anchor (only after the token gate passes). No device measurement is B2b evidence until the greedy decode completes. Stop conditions per the addendum lines 444-456: pin tree cannot compile the B2b device TU -> stop, record, row stays ACTIVE; stale _ttnncpp.so link blocker (fresh /tmp/pin-build lib64 is the verified workaround, already in place); no card window -> row stays ACTIVE with implementation owed.

## Resolution

- 2026-10-08 (branch row/kolibri-tt-b2bi-fwd, commit cbdd7cce5): the
dense-resident device forward landed and the COMPLETION CONDITION is met.
One greedy decode of the first golden prompt completed on the P150 card
through the production seam (ModelRegistry::Prepare builds the 350-projection
/ 3.540 GiB bf16 resident context; ModelRegistry::Forward runs prefill + 8
decode steps; 78/78 assertions). The routed-expert refusal fired BY NAME 450
times (50 MoE blocks x 9 steps) and the shared expert carried the step; the
full-model 141/145 token gate stays owed to B2b-ii (the goldens cannot be
replayed without the routed experts). Host gates green: test_kolibri1
234/234, w2 1608/1608, w3 900/900 (141/145, 4 near-tie, 0 hard),
test_kolibri1_tt 235/235, b2i 204/204, the new test_kolibri1_tt_b2bi host
cases green after red-first capture. Evidence:
docs/bench-evidence/kolibri1-tt-b2bi-fwd-20261008.md. One finding recorded
against the earlier Owed note: the non-pin tree's stale lib64 does not link
this row's TT binaries at all (missing use_mcast symbol); the pin fresh
/tmp/pin-build libs are the working link. Issue stays OPEN until the slice
merges; the token gate and bench anchor are B2b-ii's.
-
- 2026-10-08 (same branch, review-repair commits): the fresh mutation
review's three findings — every guarantee had NO host-side coverage, so
all three re-applied mutations left every card-less gate green — are
REPAIRED red-first WITHOUT a card. New host cases in
test_kolibri1_tt_b2bi.cpp: a HOST op census over the PRODUCTION forward
(host-memory backend + platform in the kTENSTORRENT slot, recording op
providers over the existing vt::OpProvider seam, disabled on scope exit;
per-step op order + the four sandwich norms consuming their OWN sentinel
weights + per-head q/k norms + router/shared/lm_head matmuls) carrying the
refusal counter assertion host-side (refusals >= layers x steps), and a
registry dispatch-identity case (Resolve -> Load -> Prepare ->
ModelRegistry::Forward on a TT queue completes and fires the refusal).
Mutations re-applied and RED: silenced counter (:851/:963), wrong
post_attn_norm weight (:907, records 3 and 9), deleted kTENSTORRENT arm
(:918 throw); each restored byte-for-byte and GREEN (10 cases, 182
assertions; full battery at landed counts, W3 900/900 + 141/145).
Evidence: docs/bench-evidence/kolibri1-tt-b2bi-fwd-20261008.md §9.
31 changes: 26 additions & 5 deletions .agents/specs/kolibri-tt.md
Original file line number Diff line number Diff line change
Expand Up @@ -469,8 +469,23 @@ leg is green on BOTH builds: the pin build against the fresh `/tmp/pin-build`
libs (smoke `SMOKE_RC=0`, 36,880/36,880 assertions) and the non-pin build +
gdn stub. The earlier pin device-op hangs were the pin tree's STALE in-tree
`build_Release/lib64`, not the pin source or the KMD/fw pair — no pin bump is
owed (§ Owed). Slice completion (one greedy decode on the card), B2b-ii, and
the token gate / bench anchor remain owed.
owed (§ Owed). B2b-i completion landed (2026-10-08, same day): the
dense-resident device forward (`kolibri1_tt_forward.cpp` + the production
registry dispatch) ran ONE GREEDY DECODE of a golden prompt ON THE CARD —
the slice's completion condition — with the routed-expert refusal firing by
name 450 times (50 MoE blocks x 9 steps) and the shared expert carrying the
step; evidence
`docs/bench-evidence/kolibri1-tt-b2bi-fwd-20261008.md`. The token gate
(141/145 argmax vs the full-model goldens) stays OWED to B2b-ii: the
goldens are full-model decodes and cannot be replayed without the routed
experts; the refusal firing by name is the recorded proof. The fresh mutation
review's three host-coverage findings were REPAIRED the same day (2026-10-08,
PR #3421 follow-up commits): the forward's op sequence, the refusal firing,
and the registry dispatch identity each now have a RED-first HOST-side case
in test_kolibri1_tt_b2bi.cpp (the host op census over a recording kTENSTORRENT
stand-in through the EXISTING vt::OpProvider seam — no parallel forward path;
evidence `docs/bench-evidence/kolibri1-tt-b2bi-fwd-20261008.md` §9). B2b-ii
(the streaming MoE), then the token gate and the bench anchor, remain owed.

## Git integration

Expand All @@ -481,9 +496,15 @@ One pull request for wave A (spec + implementation together), branched from

- B2b implementation (the § B2 scope — B2b addendum): the B2b-i bring-up
slice landed 2026-10-08 (build/link, resident staging, one verified device
op — see `## Now`); the dense-resident device forward's completion
condition (one greedy decode on the card), then the streaming MoE
(B2b-ii), remain owed.
op — see `## Now`); the dense-resident device forward's COMPLETION
CONDITION landed the same day (one greedy decode on the card, refusal
firing by name — see `## Now`). Still owed: the streaming MoE (B2b-ii),
then the full-model token gate (141/145 argmax, 4 flips adjudicated in
the 2.5-nat band, 0 hard) and the production bench anchor. The review
repair (2026-10-08) closed the three host-coverage findings: the refusal
firing, the forward's op sequence (per-norm weight identity), and the
kTENSTORRENT dispatch arm are each RED-first testable WITHOUT a card
(test_kolibri1_tt_b2bi.cpp host census; evidence doc §9).
- The `_ttnncpp.so` pin rebuild (verified-fresh `lib64/_ttnncpp.so`, ninja
`ttnn tt_metal` + copy) — a named prerequisite for every TT test binary
before the first B2b device run (the stale-lib64 blocker; the issue
Expand Down
7 changes: 7 additions & 0 deletions CMakeLists.txt
Original file line number Diff line number Diff line change
Expand Up @@ -856,7 +856,14 @@ add_library(vllm STATIC
# hybrid forward (RNoPE, qk-norm, sandwich norms, sigmoid-logit-add MoE).
# kolibri1_tt.cpp is the MODEL-TEXT-kolibri-1-tenstorrent wave-A staging
# plan (device-free; spec .agents/specs/kolibri-tt.md).
# kolibri1_shared.h is the CPU row's routing/step-input helpers, shared
# VERBATIM with the TT row (a pure relocation; the CPU gates pin it).
# kolibri1_tt_forward.cpp is the B2b-i dense-resident device forward
# (spec .agents/specs/kolibri-tt.md ### B2 scope — B2b addendum, slice i);
# it is backend-agnostic (vt ops + the shared residency seam) and refuses
# non-Tenstorrent queues by name.
src/vllm/model_executor/models/kolibri1_tt.cpp
src/vllm/model_executor/models/kolibri1_tt_forward.cpp
src/vllm/model_executor/models/kolibri1_registry.cpp
src/vllm/model_executor/models/kolibri1_weights.cpp
src/vllm/model_executor/models/kolibri1_forward.cpp
Expand Down
Loading
Loading