Skip to content

feat(kolibri1-tt): the B2b-i dense-resident device forward — one greedy decode on the P150 - #3421

Merged
lu-zero merged 6 commits into
localai-org:mainfrom
lu-zero:row/kolibri-tt-b2bi-fwd
Oct 8, 2026
Merged

lu-zero merged 6 commits into
localai-org:mainfrom
lu-zero:row/kolibri-tt-b2bi-fwd

Conversation

@lu-zero

@lu-zero lu-zero commented Oct 8, 2026 •

Copy link
Copy Markdown
Collaborator

feat(kolibri1-tt): the B2b-i dense-resident device forward — one greedy decode on the P150

What changed

Kolibri-1 Tenstorrent slice B2b-i (row MODEL-TEXT-kolibri-1-tenstorrent):
the resident non-expert set runs entirely on a Tenstorrent queue with the
routed-expert tier absent, and ONE GREEDY DECODE of a golden prompt
completed on the P150 — the slice's completion condition.

  • kolibri1_tt_forward.cpp/.h (new): the CPU row's op sequence op-for-op
    over the resident slice — hybrid attention (40 sliding-window layers at
    window 513, 10 full-attention RNoPE layers, two-group KV, per-head q/k
    RMS norms), the four sandwich norms, the router (bf16 [384,2560] gate,
    f32 e_score_correction_bias, sigmoid-logit-add top-6-of-384, the CPU
    row's f32 compute path), the shared expert, embed + untied lm_head.
    The one deliberate compute difference, documented in the header: the
    resident fp8-block projections dequant ONCE into device-resident bf16
    buffers (memoized; the CPU row's documented R1 disposition) instead of
    per call. The router's top-6 always requests routed experts; slice i
    fires the routed-expert refusal BY NAME (counted, message names the
    missing part, slice B2b-ii, the row, the issue) and proceeds with the
    shared expert only.
  • kolibri1_registry.cpp: the dispatch — CPU to the CPU row (unchanged),
    Tenstorrent to the new forward; every other device refused by name.
    Prepare on a Tenstorrent queue builds the device context eagerly.
  • kolibri1_shared.h (new): the sigmoid-logit-add routing and the
    per-step metadata upload relocated out of kolibri1_forward.cpp
    VERBATIM so both arms consume the same math.
  • tests/vllm/models/test_kolibri1_tt_b2bi.cpp (new) + CMake
    registration: host-side cases (no card) — the refusal contract, the
    routing contract (tie to the lower index, weights on the unbiased
    logits), the device context over the tiny fixture (bit-exact dequant
    agreement, byte accounting, the non-fp8 and non-TT refusals), the
    real-manifest byte math (350 resident projections; bf16 cost = 2x the
    resident fp8 bytes) — plus the env-gated device leg
    (VT_KOLIBRI1_TT_B2BI_MODEL, allowlisted progress env
    VT_KOLIBRI1_TT_B2BI_PROGRESS).
  • Spec ## Now/## Owed, issue resolution, and the evidence file
    docs/bench-evidence/kolibri1-tt-b2bi-fwd-20261008.md.

Why

The addendum's gate ordering makes the greedy decode the first admissible
B2b device evidence; everything before it is build or smoke output. The
refusal-instead-of-throw polarity keeps the decode completing (the
completion condition) while making the absent tier visible and counted.

How a reviewer verifies

Host (no card), fresh build dir:

cmake -S <worktree> -B /tmp/b -G Ninja -DCMAKE_BUILD_TYPE=Release \
  -DVLLM_CPP_TENSTORRENT=ON \
  -DCMAKE_PREFIX_PATH="/tmp/pin-build/lib64/cmake;/tmp/pin-build/share/cmake"
ninja -C /tmp/b test_kolibri1 test_kolibri1_w2 test_kolibri1_w3 \
  test_kolibri1_tt test_kolibri1_tt_b2i test_kolibri1_tt_b2bi

Observed 2026-10-08: test_kolibri1 234/234, w2 1608/1608, w3 900/900
(ARGMAX chain 141/145, 4 near-tie flips in the 2.5-nat band, 0 hard —
byte-identical to the landed CPU-row baseline, pinning the shared-header
relocation), test_kolibri1_tt 235/235 (B2a contracts unchanged), b2i
204/204, and the new TU's host cases green. test_kolibri1_dequant_cache
is 48/49 on this host — the fork() probe failure is PRE-EXISTING at HEAD
129e997 (verified by building that test at HEAD in a scratch worktree).

Device (P150 under the host GPU mutex; card reset + cache clear first;
runtime env TT_METAL_HOME=TT_METAL_RUNTIME_ROOT=~/Sources/tt/tt-metal-pin
and LD_LIBRARY_PATH=/tmp/pin-build/lib64:/tmp/pin-build/libexec/tt-metalium
— sourcing env-tt-common.sh instead loads the stale non-pin lib64 and
aborts):

VT_KOLIBRI1_TT_B2BI_MODEL=/mnt/models/Aleph-Alpha/Kolibri-1 \
  ./tests/test_kolibri1_tt_b2bi

Observed: load 18.9 s; context 350 projections, 3.540 GiB bf16 in 1.0 s;
prefill + 8 decode steps; "GREEDY DECODE COMPLETE on card: 15 tokens";
the routed-expert refusal fired 450 times by name; 78/78 assertions,
Status SUCCESS. The goldens' expected tokens are deliberately NOT
asserted: the goldens are full-model decodes that cannot be replayed
without the routed experts, so the 141/145 token gate stays owed to
B2b-ii — the refusal firing by name is the recorded proof.

Red-first: the inherited uncommitted draft had never been built (first
build failed on two compile errors, fixed forward); the new gate TU ran
red on three wrong test expectations before green; no product assertion
was weakened. Mutation-style reachability: the device leg enters through
ModelRegistry::Prepare/Forward — the production seam — not a hand-built
context.

What remains out of scope

  • B2b-ii: the routed-expert streaming tier (slot pool, fetch executor,
    router readback dtype pivot, the runtime stream-bound assert).
  • The full-model token gate (141/145, 4 flips in the band, 0 hard) and
    the production bench anchor — only after B2b-ii.
  • No B2a policy change (Kolibri1TTExpertSlotPolicy / DispatchPlan /
    StreamingPlan contracts stay green unchanged); no tt-metal pin bump.
  • One finding recorded for the record: the non-pin tt-metal tree's stale
    lib64 does not link this row's TT binaries at all (its
    chunk_gated_delta_rule export predates the use_mcast parameter), so the
    pin fresh /tmp/pin-build libs are the working link on this host.

Fixes ISSUE-LOCAL-01M4E22DM790W0E69M07D5XA9D (dense-resident slice; the
issue stays open for the merge record per the local close-on-land rule —
close when this lands).

Review repair (2026-10-08, commits 8db5d58 + 7615d04)

What changed: the fresh mutation review returned FAIL with three findings,
all of one form — every guarantee had no host-side coverage, so re-applying
each mutation left all card-less gates green. The repair adds two host
cases to tests/vllm/models/test_kolibri1_tt_b2bi.cpp (no product code
changed; the device leg's semantics are exactly as landed):

  • HOST op census over the PRODUCTION ForwardKolibri1TTResidentForward: a
    host-memory backend + platform stand in the kTENSTORRENT slot for the
    scope of each case, and recording op providers over the EXISTING
    vt::OpProvider seam capture which ops fire, in what order, consuming
    which norm weight (the tiny fixture's norm weights now carry distinct
    bf16 sentinels). This carries the refusal counter assertion
    (refusals >= layers x steps) HOST-SIDE — previously only inside the
    device leg — and pins the four sandwich norms to their OWN weights in
    the CPU row's order, plus per-head q/k norms, router, shared expert,
    and the lm_head. No device kernel executes; no parallel forward path
    exists; the recorders are disabled on scope exit and the previous
    backend/platform restored, so a card box's device leg is untouched.
  • Registry dispatch identity: Resolve -> Load -> Prepare ->
    ModelRegistry::Forward on a TT queue completes through the
    kTENSTORRENT arm and fires the refusal, so deleting the dispatch arm
    goes red host-side.

Why: closes the three findings without weakening any assertion and
without touching the CPU row or the device leg.

How to verify (each mutation is the reviewer's exact mutation, then the
product code restored byte-for-byte):

  • Silenced refusal counter -> RED at test_kolibri1_tt_b2bi.cpp:851 and
    :963 (8/10 cases, 180/182 assertions) -> restore -> 10/10, 182/182.
  • post_attn_norm swapped to the input_layernorm weight -> RED at :907
    (norm records 3 and 9 carry the 1.0 sentinel instead of 2.0, in BOTH
    layers) -> restore -> 10/10, 182/182.
  • kTENSTORRENT registry arm replaced with a throw -> RED at :918 (test
    case threw the arm-not-implemented error) -> restore -> 10/10, 182/182.

Full host battery at the landed counts: test_kolibri1 234/234,
test_kolibri1_tt 235/235, test_kolibri1_tt_b2i 204/204,
test_kolibri1_tt_b2bi 10 cases / 182 assertions (was 8/62),
test_kolibri1_dequant 10/10, test_kolibri1_moe_glue 21/21,
test_kolibri1_w2 1608/1608, test_kolibri1_decode_bench 2/2, W3 900/900
with ARGMAX CHAIN 141/145 (4 near-tie, 0 hard). test_kolibri1_dequant_cache
48/49 — the documented pre-existing fork() failure in the default-off
probe, reproduced identically and untouched by this repair. Evidence:
docs/bench-evidence/kolibri1-tt-b2bi-fwd-20261008.md section 9.

FOLLOWING_AGENTS_PROTOCOL

Following-Agents-Protocol: true
AI-Assisted: true
Assisted-by: AGENT:zai/glm-5.3-flash [maki]

…orward)

The B2b-i bring-up landed (cf6258c + b52c0ba): build/link, resident
staging, and one verified device op. The completion condition — one greedy
decode of a golden prompt on the P150 with the resident non-expert set
running entirely on device and the routed tier absent — is not yet met, so
the work is tracked in a canonical local issue before the implementation
starts, per the every-change-starts-from-an-issue rule.

FOLLOWING_AGENTS_PROTOCOL

Following-Agents-Protocol: true
AI-Assisted: true
Assisted-by: AGENT:mistral/mistral-large-4 [maki]
The Kolibri-1 Tenstorrent slice B2b-i (issue
ISSUE-LOCAL-01M4E22DM790W0E69M07D5XA9D): the resident non-expert set runs
entirely on a Tenstorrent queue with the routed-expert tier absent.

kolibri1_tt_forward.cpp runs the CPU row's op sequence op-for-op over the
resident slice — hybrid attention (40 sliding-window layers at window 513,
10 full-attention RNoPE layers, two-group KV, per-head q/k RMS norms), the
four sandwich norms, the router (bf16 [384,2560] gate, f32
e_score_correction_bias, sigmoid-logit-add top-6-of-384, the CPU row's f32
compute path), the shared expert, embed + untied lm_head — with the ONE
deliberate compute difference documented in the header: the resident
fp8-block projections dequant ONCE into device-resident bf16 buffers
(memoized; the CPU row's documented R1 disposition) instead of per call.
The router's top-6 always requests routed experts; slice i fires the
routed-expert refusal BY NAME (counted, once-per-process message naming
the missing part, the owning slice B2b-ii, the row, and the issue) and the
step proceeds with the shared expert only.

The registry dispatches CPU to the CPU row (unchanged) and Tenstorrent to
the new forward; every other device is refused by name at the registry
boundary. Prepare on a Tenstorrent queue builds the device context
eagerly. kolibri1_shared.h relocates the sigmoid-logit-add routing and
the per-step metadata upload out of kolibri1_forward.cpp VERBATIM so both
arms consume the same math; test_kolibri1 (234/234), test_kolibri1_w2
(1608/1608) and test_kolibri1_w3 (900/900, ARGMAX chain 141/145, 4
near-tie flips in the 2.5-nat band, 0 hard) pin that relocation
byte-identical.

test_kolibri1_tt_b2bi.cpp is the gate TU. Host-side (no card): the
refusal contract, the relocated routing contract (tie-break to the lower
index, weights on the unbiased logits), the device context over the tiny
fixture (7 projections per layer, BIT-EXACT agreement of every memoized
dequant against kolibri1_fp8::DequantRowsBf16, byte accounting, the
non-fp8 refusal, the non-TT forward refusal) and the real-manifest byte
math (350 resident projections, the bf16 dequant cost = 2x the resident
fp8 bytes, within 8 MiB of the spec's 1.587 + 0.188 GiB plan). Device leg
(env-gated VT_KOLIBRI1_TT_B2BI_MODEL, operator-run): resident context
through the production ModelRegistry::Prepare seam, staged-byte checks,
and ONE GREEDY DECODE of a golden prompt on the card — the slice's
completion condition — with the refusal firing by name throughout. The
goldens' expected tokens are NOT asserted there: the goldens are
full-model decodes that cannot be replayed without the routed experts, so
the 141/145 token gate stays owed to B2b-ii, exactly as the addendum's
gate ordering records.

FOLLOWING_AGENTS_PROTOCOL

Following-Agents-Protocol: true
AI-Assisted: true
Assisted-by: AGENT:zai/glm-5.3-flash [maki]
…the card

The slice's completion condition is met (2026-10-08, P150 under the host
GPU mutex): the dense-resident forward ran one greedy decode of the first
golden prompt on the card through the production seam — 350 projections /
3.540 GiB bf16 resident context, prefill + 8 decode steps, 78/78
assertions — with the routed-expert refusal firing by name 450 times (50
MoE blocks x 9 steps) and the shared expert carrying every step.

The spec's ## Now records the completion and what stays owed (B2b-ii,
then the full-model 141/145 token gate and the bench anchor); the Owed
item moves the completion condition from owed to landed. The issue's
Resolution carries the dated gate evidence. The evidence file records the
build recipe (the pin fresh /tmp/pin-build libs — with the finding that
the non-pin stale lib64 does not link this row's TT binaries at all: its
chunk_gated_delta_rule export predates the use_mcast parameter, so the
"no-op stub" note in ## Owed is not reachable for a vLLM TT link), the
lease window, the staged byte totals vs the plan, the host gate table
with exact counts, the decode log, the per-op agreement envelope, and the
red-first capture (the inherited draft had never been built; two compile
errors fixed forward; the new gate TU ran red on three wrong test
expectations before green).

The env-tt-common.sh LD_LIBRARY_PATH preemption of the binary's pin
RUNPATH is recorded in the evidence file's runtime recipe — a device run
that sources that script loads the stale non-pin lib64 and aborts.

FOLLOWING_AGENTS_PROTOCOL

Following-Agents-Protocol: true
AI-Assisted: true
Assisted-by: AGENT:zai/glm-5.3-flash [maki]
…rage

The fresh mutation review returned FAIL with three findings of one form:
every guarantee had no host-side coverage, so re-applying each mutation
left all card-less gates green. The repair adds a host op census over the
PRODUCTION ForwardKolibri1TTResidentForward and a registry dispatch-
identity case, both running without a card.

The seam is the existing vt::OpProvider provider table plus a registered
backend, not a parallel path: a host-memory backend and platform stand in
the kTENSTORRENT slot for the scope of each case (the
test_resident_weight_host_addressable.cpp pattern), and recording
providers at priority 100 capture which ops the forward emits, in what
order, consuming which norm weight. The providers register under one
test-only name and are DISABLED on scope exit, the previous backend and
platform are restored, and TenstorrentPresent() excludes the stand-in, so
a card box's device leg selects its native kernels exactly as landed. The
tiny fixture's sandwich and per-head norm weights now carry distinct bf16
sentinels so the census identifies WHICH weight a recorded norm consumed.

Mutation -> red -> restore -> green, captured (evidence doc section 9):
silencing the refusal counter goes red at the census and dispatch cases;
swapping post_attn_norm to the input_layernorm weight goes red on norm
records 3 and 9 in both layers; deleting the kTENSTORRENT registry arm
goes red on the dispatch case. Full host battery green at the landed
counts, W3 900/900 and 141/145 unchanged.

FOLLOWING_AGENTS_PROTOCOL

Following-Agents-Protocol: true
AI-Assisted: true
Assisted-by: AGENT:zai/glm-5.3-flash [maki]
…nce recorded

The three mutation-review findings are closed with the host op census and
the registry dispatch-identity case. The evidence doc gains section 9:
each of the reviewer's exact mutations is RED host-side (silenced refusal
counter at the census and dispatch cases, the wrong post_attn_norm weight
on norm records 3 and 9, the deleted kTENSTORRENT arm on the dispatch
case), each restored byte-for-byte and green, with the full host battery
at the landed counts and W3 900/900 / 141/145 unchanged. The spec's
## Now and ## Owed and the issue's Resolution record the repaired
coverage; the device leg's semantics are exactly as landed.

FOLLOWING_AGENTS_PROTOCOL

Following-Agents-Protocol: true
AI-Assisted: true
Assisted-by: AGENT:zai/glm-5.3-flash [maki]
gh pr edit does not retrigger CI. Empty commit so the checks read the
updated body (review-repair section appended; trailer block still last;
agent-pr-body.py --pr 3421 exits 0).

FOLLOWING_AGENTS_PROTOCOL

Following-Agents-Protocol: true
AI-Assisted: true
Assisted-by: AGENT:zai/glm-5.3-flash [maki]
@lu-zero
lu-zero merged commit 6b63e62 into localai-org:main Oct 8, 2026
25 of 31 checks passed
@lu-zero
lu-zero deleted the row/kolibri-tt-b2bi-fwd branch October 8, 2026 21:43
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant