Repository navigation
feat(kolibri1-tt): the B2b-i dense-resident device forward — one greedy decode on the P150 - #3421
Merged
Merged
Conversation
…orward) The B2b-i bring-up landed (cf6258c + b52c0ba): build/link, resident staging, and one verified device op. The completion condition — one greedy decode of a golden prompt on the P150 with the resident non-expert set running entirely on device and the routed tier absent — is not yet met, so the work is tracked in a canonical local issue before the implementation starts, per the every-change-starts-from-an-issue rule. FOLLOWING_AGENTS_PROTOCOL Following-Agents-Protocol: true AI-Assisted: true Assisted-by: AGENT:mistral/mistral-large-4 [maki]
The Kolibri-1 Tenstorrent slice B2b-i (issue ISSUE-LOCAL-01M4E22DM790W0E69M07D5XA9D): the resident non-expert set runs entirely on a Tenstorrent queue with the routed-expert tier absent. kolibri1_tt_forward.cpp runs the CPU row's op sequence op-for-op over the resident slice — hybrid attention (40 sliding-window layers at window 513, 10 full-attention RNoPE layers, two-group KV, per-head q/k RMS norms), the four sandwich norms, the router (bf16 [384,2560] gate, f32 e_score_correction_bias, sigmoid-logit-add top-6-of-384, the CPU row's f32 compute path), the shared expert, embed + untied lm_head — with the ONE deliberate compute difference documented in the header: the resident fp8-block projections dequant ONCE into device-resident bf16 buffers (memoized; the CPU row's documented R1 disposition) instead of per call. The router's top-6 always requests routed experts; slice i fires the routed-expert refusal BY NAME (counted, once-per-process message naming the missing part, the owning slice B2b-ii, the row, and the issue) and the step proceeds with the shared expert only. The registry dispatches CPU to the CPU row (unchanged) and Tenstorrent to the new forward; every other device is refused by name at the registry boundary. Prepare on a Tenstorrent queue builds the device context eagerly. kolibri1_shared.h relocates the sigmoid-logit-add routing and the per-step metadata upload out of kolibri1_forward.cpp VERBATIM so both arms consume the same math; test_kolibri1 (234/234), test_kolibri1_w2 (1608/1608) and test_kolibri1_w3 (900/900, ARGMAX chain 141/145, 4 near-tie flips in the 2.5-nat band, 0 hard) pin that relocation byte-identical. test_kolibri1_tt_b2bi.cpp is the gate TU. Host-side (no card): the refusal contract, the relocated routing contract (tie-break to the lower index, weights on the unbiased logits), the device context over the tiny fixture (7 projections per layer, BIT-EXACT agreement of every memoized dequant against kolibri1_fp8::DequantRowsBf16, byte accounting, the non-fp8 refusal, the non-TT forward refusal) and the real-manifest byte math (350 resident projections, the bf16 dequant cost = 2x the resident fp8 bytes, within 8 MiB of the spec's 1.587 + 0.188 GiB plan). Device leg (env-gated VT_KOLIBRI1_TT_B2BI_MODEL, operator-run): resident context through the production ModelRegistry::Prepare seam, staged-byte checks, and ONE GREEDY DECODE of a golden prompt on the card — the slice's completion condition — with the refusal firing by name throughout. The goldens' expected tokens are NOT asserted there: the goldens are full-model decodes that cannot be replayed without the routed experts, so the 141/145 token gate stays owed to B2b-ii, exactly as the addendum's gate ordering records. FOLLOWING_AGENTS_PROTOCOL Following-Agents-Protocol: true AI-Assisted: true Assisted-by: AGENT:zai/glm-5.3-flash [maki]
…the card The slice's completion condition is met (2026-10-08, P150 under the host GPU mutex): the dense-resident forward ran one greedy decode of the first golden prompt on the card through the production seam — 350 projections / 3.540 GiB bf16 resident context, prefill + 8 decode steps, 78/78 assertions — with the routed-expert refusal firing by name 450 times (50 MoE blocks x 9 steps) and the shared expert carrying every step. The spec's ## Now records the completion and what stays owed (B2b-ii, then the full-model 141/145 token gate and the bench anchor); the Owed item moves the completion condition from owed to landed. The issue's Resolution carries the dated gate evidence. The evidence file records the build recipe (the pin fresh /tmp/pin-build libs — with the finding that the non-pin stale lib64 does not link this row's TT binaries at all: its chunk_gated_delta_rule export predates the use_mcast parameter, so the "no-op stub" note in ## Owed is not reachable for a vLLM TT link), the lease window, the staged byte totals vs the plan, the host gate table with exact counts, the decode log, the per-op agreement envelope, and the red-first capture (the inherited draft had never been built; two compile errors fixed forward; the new gate TU ran red on three wrong test expectations before green). The env-tt-common.sh LD_LIBRARY_PATH preemption of the binary's pin RUNPATH is recorded in the evidence file's runtime recipe — a device run that sources that script loads the stale non-pin lib64 and aborts. FOLLOWING_AGENTS_PROTOCOL Following-Agents-Protocol: true AI-Assisted: true Assisted-by: AGENT:zai/glm-5.3-flash [maki]
…rage The fresh mutation review returned FAIL with three findings of one form: every guarantee had no host-side coverage, so re-applying each mutation left all card-less gates green. The repair adds a host op census over the PRODUCTION ForwardKolibri1TTResidentForward and a registry dispatch- identity case, both running without a card. The seam is the existing vt::OpProvider provider table plus a registered backend, not a parallel path: a host-memory backend and platform stand in the kTENSTORRENT slot for the scope of each case (the test_resident_weight_host_addressable.cpp pattern), and recording providers at priority 100 capture which ops the forward emits, in what order, consuming which norm weight. The providers register under one test-only name and are DISABLED on scope exit, the previous backend and platform are restored, and TenstorrentPresent() excludes the stand-in, so a card box's device leg selects its native kernels exactly as landed. The tiny fixture's sandwich and per-head norm weights now carry distinct bf16 sentinels so the census identifies WHICH weight a recorded norm consumed. Mutation -> red -> restore -> green, captured (evidence doc section 9): silencing the refusal counter goes red at the census and dispatch cases; swapping post_attn_norm to the input_layernorm weight goes red on norm records 3 and 9 in both layers; deleting the kTENSTORRENT registry arm goes red on the dispatch case. Full host battery green at the landed counts, W3 900/900 and 141/145 unchanged. FOLLOWING_AGENTS_PROTOCOL Following-Agents-Protocol: true AI-Assisted: true Assisted-by: AGENT:zai/glm-5.3-flash [maki]
…nce recorded The three mutation-review findings are closed with the host op census and the registry dispatch-identity case. The evidence doc gains section 9: each of the reviewer's exact mutations is RED host-side (silenced refusal counter at the census and dispatch cases, the wrong post_attn_norm weight on norm records 3 and 9, the deleted kTENSTORRENT arm on the dispatch case), each restored byte-for-byte and green, with the full host battery at the landed counts and W3 900/900 / 141/145 unchanged. The spec's ## Now and ## Owed and the issue's Resolution record the repaired coverage; the device leg's semantics are exactly as landed. FOLLOWING_AGENTS_PROTOCOL Following-Agents-Protocol: true AI-Assisted: true Assisted-by: AGENT:zai/glm-5.3-flash [maki]
gh pr edit does not retrigger CI. Empty commit so the checks read the updated body (review-repair section appended; trailer block still last; agent-pr-body.py --pr 3421 exits 0). FOLLOWING_AGENTS_PROTOCOL Following-Agents-Protocol: true AI-Assisted: true Assisted-by: AGENT:zai/glm-5.3-flash [maki]
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
feat(kolibri1-tt): the B2b-i dense-resident device forward — one greedy decode on the P150
What changed
Kolibri-1 Tenstorrent slice B2b-i (row MODEL-TEXT-kolibri-1-tenstorrent):
the resident non-expert set runs entirely on a Tenstorrent queue with the
routed-expert tier absent, and ONE GREEDY DECODE of a golden prompt
completed on the P150 — the slice's completion condition.
kolibri1_tt_forward.cpp/.h(new): the CPU row's op sequence op-for-opover the resident slice — hybrid attention (40 sliding-window layers at
window 513, 10 full-attention RNoPE layers, two-group KV, per-head q/k
RMS norms), the four sandwich norms, the router (bf16 [384,2560] gate,
f32 e_score_correction_bias, sigmoid-logit-add top-6-of-384, the CPU
row's f32 compute path), the shared expert, embed + untied lm_head.
The one deliberate compute difference, documented in the header: the
resident fp8-block projections dequant ONCE into device-resident bf16
buffers (memoized; the CPU row's documented R1 disposition) instead of
per call. The router's top-6 always requests routed experts; slice i
fires the routed-expert refusal BY NAME (counted, message names the
missing part, slice B2b-ii, the row, the issue) and proceeds with the
shared expert only.
kolibri1_registry.cpp: the dispatch — CPU to the CPU row (unchanged),Tenstorrent to the new forward; every other device refused by name.
Prepare on a Tenstorrent queue builds the device context eagerly.
kolibri1_shared.h(new): the sigmoid-logit-add routing and theper-step metadata upload relocated out of kolibri1_forward.cpp
VERBATIM so both arms consume the same math.
tests/vllm/models/test_kolibri1_tt_b2bi.cpp(new) + CMakeregistration: host-side cases (no card) — the refusal contract, the
routing contract (tie to the lower index, weights on the unbiased
logits), the device context over the tiny fixture (bit-exact dequant
agreement, byte accounting, the non-fp8 and non-TT refusals), the
real-manifest byte math (350 resident projections; bf16 cost = 2x the
resident fp8 bytes) — plus the env-gated device leg
(VT_KOLIBRI1_TT_B2BI_MODEL, allowlisted progress env
VT_KOLIBRI1_TT_B2BI_PROGRESS).
## Now/## Owed, issue resolution, and the evidence filedocs/bench-evidence/kolibri1-tt-b2bi-fwd-20261008.md.Why
The addendum's gate ordering makes the greedy decode the first admissible
B2b device evidence; everything before it is build or smoke output. The
refusal-instead-of-throw polarity keeps the decode completing (the
completion condition) while making the absent tier visible and counted.
How a reviewer verifies
Host (no card), fresh build dir:
Observed 2026-10-08: test_kolibri1 234/234, w2 1608/1608, w3 900/900
(ARGMAX chain 141/145, 4 near-tie flips in the 2.5-nat band, 0 hard —
byte-identical to the landed CPU-row baseline, pinning the shared-header
relocation), test_kolibri1_tt 235/235 (B2a contracts unchanged), b2i
204/204, and the new TU's host cases green. test_kolibri1_dequant_cache
is 48/49 on this host — the fork() probe failure is PRE-EXISTING at HEAD
129e997 (verified by building that test at HEAD in a scratch worktree).
Device (P150 under the host GPU mutex; card reset + cache clear first;
runtime env TT_METAL_HOME=TT_METAL_RUNTIME_ROOT=~/Sources/tt/tt-metal-pin
and LD_LIBRARY_PATH=/tmp/pin-build/lib64:/tmp/pin-build/libexec/tt-metalium
— sourcing env-tt-common.sh instead loads the stale non-pin lib64 and
aborts):
Observed: load 18.9 s; context 350 projections, 3.540 GiB bf16 in 1.0 s;
prefill + 8 decode steps; "GREEDY DECODE COMPLETE on card: 15 tokens";
the routed-expert refusal fired 450 times by name; 78/78 assertions,
Status SUCCESS. The goldens' expected tokens are deliberately NOT
asserted: the goldens are full-model decodes that cannot be replayed
without the routed experts, so the 141/145 token gate stays owed to
B2b-ii — the refusal firing by name is the recorded proof.
Red-first: the inherited uncommitted draft had never been built (first
build failed on two compile errors, fixed forward); the new gate TU ran
red on three wrong test expectations before green; no product assertion
was weakened. Mutation-style reachability: the device leg enters through
ModelRegistry::Prepare/Forward — the production seam — not a hand-built
context.
What remains out of scope
router readback dtype pivot, the runtime stream-bound assert).
the production bench anchor — only after B2b-ii.
StreamingPlan contracts stay green unchanged); no tt-metal pin bump.
lib64 does not link this row's TT binaries at all (its
chunk_gated_delta_rule export predates the use_mcast parameter), so the
pin fresh /tmp/pin-build libs are the working link on this host.
Fixes ISSUE-LOCAL-01M4E22DM790W0E69M07D5XA9D (dense-resident slice; the
issue stays open for the merge record per the local close-on-land rule —
close when this lands).
Review repair (2026-10-08, commits 8db5d58 + 7615d04)
What changed: the fresh mutation review returned FAIL with three findings,
all of one form — every guarantee had no host-side coverage, so re-applying
each mutation left all card-less gates green. The repair adds two host
cases to tests/vllm/models/test_kolibri1_tt_b2bi.cpp (no product code
changed; the device leg's semantics are exactly as landed):
host-memory backend + platform stand in the kTENSTORRENT slot for the
scope of each case, and recording op providers over the EXISTING
vt::OpProvider seam capture which ops fire, in what order, consuming
which norm weight (the tiny fixture's norm weights now carry distinct
bf16 sentinels). This carries the refusal counter assertion
(refusals >= layers x steps) HOST-SIDE — previously only inside the
device leg — and pins the four sandwich norms to their OWN weights in
the CPU row's order, plus per-head q/k norms, router, shared expert,
and the lm_head. No device kernel executes; no parallel forward path
exists; the recorders are disabled on scope exit and the previous
backend/platform restored, so a card box's device leg is untouched.
ModelRegistry::Forward on a TT queue completes through the
kTENSTORRENT arm and fires the refusal, so deleting the dispatch arm
goes red host-side.
Why: closes the three findings without weakening any assertion and
without touching the CPU row or the device leg.
How to verify (each mutation is the reviewer's exact mutation, then the
product code restored byte-for-byte):
:963 (8/10 cases, 180/182 assertions) -> restore -> 10/10, 182/182.
(norm records 3 and 9 carry the 1.0 sentinel instead of 2.0, in BOTH
layers) -> restore -> 10/10, 182/182.
case threw the arm-not-implemented error) -> restore -> 10/10, 182/182.
Full host battery at the landed counts: test_kolibri1 234/234,
test_kolibri1_tt 235/235, test_kolibri1_tt_b2i 204/204,
test_kolibri1_tt_b2bi 10 cases / 182 assertions (was 8/62),
test_kolibri1_dequant 10/10, test_kolibri1_moe_glue 21/21,
test_kolibri1_w2 1608/1608, test_kolibri1_decode_bench 2/2, W3 900/900
with ARGMAX CHAIN 141/145 (4 near-tie, 0 hard). test_kolibri1_dequant_cache
48/49 — the documented pre-existing fork() failure in the default-off
probe, reproduced identically and untouched by this repair. Evidence:
docs/bench-evidence/kolibri1-tt-b2bi-fwd-20261008.md section 9.
FOLLOWING_AGENTS_PROTOCOL
Following-Agents-Protocol: true
AI-Assisted: true
Assisted-by: AGENT:zai/glm-5.3-flash [maki]