Skip to content

A spin whose loop calls spins is one shared call on the devices when one function runs it twice, so raytrace runs 1.8x faster and editdist 1.1x - #1282

Closed
wakamex wants to merge 1 commit into
bendlang:mainfrom
wakamex:share-rule-2
Closed

wakamex wants to merge 1 commit into
bendlang:mainfrom
wakamex:share-rule-2

Conversation

@wakamex

@wakamex wakamex commented Oct 2, 2026 •

Copy link
Copy Markdown
Contributor

Follow-up to #1102. Raytrace on the GPU got 38% slower between 2.0.3 and 2.0.26, when its bounce loop shrank under SPIN_FAR and got inlined into a shader that is itself inlined four times. Copies times size, as suggested on #1102, fixed raytrace (1.6x) but made gameoflife 7x slower, mandelbrot 1.85x and kmeans 1.45x, since it turned small helpers with many copies into calls inside hot loops. ncu on my RTX 3090 shows what actually costs raytrace (higher active threads and lower instruction counts are better):

Active threads per warp Issued instructions
raytrace, shader inlined 4 times (now) 2.1 of 32 101 billion
raytrace, one shared shader 3.7 of 32 60 billion
gameoflife, spin_8 inlined 16 times (now) 32 of 32 213 billion
gameoflife, spin_8 as one call 32 of 32 1,192 billion

Bend's GPU lanes run unrelated tasks, so a warp only runs threads together at the same instruction. Raytrace's pixel loop runs the shader four times in a row, and the bounce loop in it ends at a different depth per ray, so lanes drift into different copies and the warp runs the copies one after another. One shared copy lets them meet again. Gameoflife has no loops, so its lanes never drift and a call only adds work.

That only happens when all three hold, and each one has a case against it:

  • One function runs the copies in turn. Mandelbrot's escape wrapper is called from two different functions; its lanes sit at 32 of 32 either way, and sharing it cost 1.6%.
  • The spin holds a loop. Loop-free code rejoins after every branch (gameoflife).
  • The loop calls other spins. Hashmap's bucket walk calls none, does a few loads per step, and sharing it cost 2.2% at the same 7.5 of 32 active threads.

So a spin becomes SHARE when it holds a loop that calls spins and some function runs two or more copies of it. Each segment records the spins it calls as emit_fuse emits them; copies count through inlined spins and stop at shared or FAR ones, deciding callers first. SHARE is FAR on the devices and INLINE on the host, where these calls cost raytrace 22%. SPIN_FAR, FAR, term_drop and the meaning of a segment's spin flag don't change. comp.ts goes from 63980 to 64305 ttok (+29 -4 lines), which puts it 305 over the 64k cap.

bend-bench paired (fast-gpu-v2 sizes, RTX 3090, CUDA 13.1), main at a3f1782 against this PR in 10 blocks of A B B A / B A A B, each block scored (A1 + A2) / (B1 + B2); above 1x means this PR is faster:

Median block 95% interval All 10 blocks
raytrace 1.844x 1.827x to 1.875x 1.795x to 1.898x
editdist 1.145x 1.125x to 1.151x 1.107x to 1.161x

The other 20 workloads compile to the same program as main: the same GPU program and the same host code, since the SHARE definitions they don't use only change the source copy the binary embeds. So they weren't timed. The only other programs with SHARE spins are three demos and two tests: app_slash_boss_3d's headless probe gives the same output with frames at 60.9 -> 60.2 ms, pure_hvm5_mini is too short to time, and app_ray_tracer_3d needs a display. With bend-bench's compiler-check, every test keeps its result. Metal is unmeasured since I don't have Apple hardware; SHARE is FAR there too, so timings of raytrace and editdist on an M-series Mac would be welcome.

@wakamex

wakamex commented Oct 2, 2026

Copy link
Copy Markdown
Contributor Author

I put up #1230 to save 200 ttoks which would then have allowed this to come in under the cap. Looks like we've blown past it since then, but it's no longer binding? I see main at bend2/comp.ts: 64251 > 64000 ttok right now.

@wakamex

wakamex commented Oct 2, 2026 •

Copy link
Copy Markdown
Contributor Author

These five merges on 10-01 each left comp.ts under the 64,000 ttok cap:

Merged (UTC) PR Author / merged by comp.ts ttok added main after
10-01 15:10 #1196 fix(compiler): avoid expanding shared display layouts vicmcorrea / Lorenzobattistela +13 63,120
10-01 15:22 #1239 An operation's argument nested past 32 brackets goes to a local nicolas-abril / nicolas-abril +93 63,213
10-01 16:36 #1156 A file's names are keyed ns:name inside the compiler nicolas-abril / nicolas-abril +96 63,309
10-01 18:07 #1245 A deadline or a timer fires while computations keep the loop busy nicolas-abril / nicolas-abril +193 63,502
10-01 19:13 #1074 The emitted C builds at -O0 aldeni / Lorenzobattistela +34 63,536

Together they added 429 tokens and left main 464 under the cap.

The one that crossed it was #1269 on 10-02 (@nicolas-abril, self-merged): +714 tokens, taking main to 64,250. Its "179 under the cap" was measured on a branch based before the five merges above, and the gate wasn't rerun after merging. @Lorenzobattistela's repo gate 49/49 on main a few hours later couldn't have been counting: with a ttok that prints nothing, gates/repo.ts reads Number("") as 0 and passes every file (fix here: #1283).

Token counts are ttok over bend2/comp.ts at each merge commit. That's the gate's own metric: it gives 64,172 at d7af6bf, the figure Taelin recorded there.

@wakamex

wakamex commented Oct 2, 2026 •

Copy link
Copy Markdown
Contributor Author

I think this shows a thorny issue of trying to land +ttok improvements. It forces you to go code-golfing for -ttok savings. I tried to keep that separate. #1230 was -200 ttok at the time I submitted it. #1235 came later with larger savings and was merged before mine, leaving me with only -90 ttok. How many bend-bux of credit should I apply toward my +325 ttok improvement (this PR)? At the time of #1230, this PR would stay under 64k ttok limit. But since then, comp.ts has grown several times, and is now actually over the cap.

I could have submitted this PR as #/1231. But I chose not too, especially since you can't make stacked PRs on a fork. Making me a contributor solves this 😅

@Lorenzobattistela

Copy link
Copy Markdown
Collaborator

Hey @wakamex thanks for the comments. Yes I noted the bug on my ttok returning 0 (thanks for the fix btw) and that's how some merges happened that shouldn't. I was reviewing them mostly in order of arrival, that's why yours ended up merging on a "wrong state". I'll look into it and get back.

@wakamex

wakamex commented Oct 4, 2026

Copy link
Copy Markdown
Contributor Author

Updated the timing in the description. I added a paired mode to bend-bench that runs base and PR as A B B A / B A A B blocks, so background load cancels within each block, and only times workloads whose compiled program actually changes. Here that's raytrace (1.867x, 95% interval 1.857x to 1.889x) and editdist (1.139x, 1.135x to 1.147x); the other 20 compile to the same program as main. The earlier 1.37x editdist round was a slow base half.

Small correction to my comment above: #1230 kept -110 ttok after #1235, not -90 (90 is what it lost).

…one function runs it twice

Bend's GPU lanes run unrelated tasks, so a warp's threads only execute
together at the same instruction. When one function runs two copies of a
spin holding a loop in turn, a lane that finishes the first copy early
moves on to the second while its neighbours are still in the first, and
the warp runs the copies one after another. raytrace's shader, inlined
four times in its pixel loop around the bounce loop, ran 2.1 of 32
threads per warp under ncu; one shared copy lifts that to 3.7 and cuts
issued instructions 1.7x.

So a spin is SHARE when it holds a loop that calls spins and some
function runs two or more copies of it. Each segment records the spins
it calls as emit_fuse emits them, and copies count through inlined
spins and stop at shared or FAR ones, deciding callers first. SHARE is
FAR on the devices and INLINE on the host, where clang decides (as host
calls, raytrace's shared loops cost 22% on 16 threads). Each condition
rests on a measured case:
  - copies in different functions run in different tasks, so sharing
    them buys no reconvergence: mandelbrot's escape wrapper, called from
    two functions, ran 32 of 32 threads either way and lost 1.6% to the
    call;
  - a loop that calls no spin is a few statements per step: sharing
    hashmap's bucket walk issued 10.7% more instructions at the same 7.5
    of 32 active threads and lost 2.2%;
  - loop-free code rejoins after each branch: gameoflife runs 32 of 32
    and a call there costs 5.6x the instructions.
SPIN_FAR, FAR, term_drop and the meaning of a segment's spin flag are
unchanged. comp.ts 63980 -> 64305 ttok, +29 -4 lines.

Measured with bend-bench paired (fast-gpu-v2 sizes, RTX 3090, CUDA 13.1,
4GB heap, shared GPU) against a3f1782: after a check and a warmup per
side, 10 blocks of A B B A / B A A B, each scored (A1 + A2) / (B1 + B2),
median and 95% bootstrap interval of the median:
  raytrace 1.844x (1.827x to 1.875x);
  editdist 1.145x (1.125x to 1.151x).
The other 20 workloads compile to the same program as a3f1782 (the same
GPU program and host code; the SHARE definitions they don't use only
change the source copy the binary embeds), so they were not timed.
bend-bench compiler-check against a3f1782: every test keeps its status
(691 pass, 11 wrong, 841 build-fail), and only raytrace, editdist, three
demos and two tests get SHARE spins.
Of the demos it touches, app_slash_boss_3d's 120-frame headless probe
gave the same output and frames of 60.9 -> 60.2 ms on a loaded host
(measured on 1adb0a6); pure_hvm5_mini runs in 12 ms, too short to time;
app_ray_tracer_3d needs a display. Metal is unmeasured: no Apple
hardware here.
@Lorenzobattistela

Copy link
Copy Markdown
Collaborator

Thanks for the careful measurements and for #1283.

I'm closing this one, for three reasons.

  1. The gain on Metal is small for what it costs. I merged the PR onto main 1159f0b and ran the perf gate on the M4 minis 7 times, alternating main and this PR, and the gains were ~noise (+0.10x and +0.05x)

The other 15 benches and both CPU modes are unchanged, as you said. For a rule that is still a heuristic, +325 ttok is a lot.

Also since this is an heuristic change, it can cause regressions in other kinds of programs (no guarantee).

  1. It isn't the rule we asked for. On A native is a call on the device when it loops and is long #1102 the direction was to count inline copies times size. This counts copies and doesn't look at size. It adds two more conditions instead ("a loop", "calls spins"), each tuned to one bench. Your write-up does explain why copies times size hurt gameoflife, mandelbrot and kmeans, which is a good argument for rethinking the rule, maybe not for adding another heuristic next to SPIN_FAR. The choice of what to inline has no correct answer, and we'd rather not grow the special cases.

  2. Compile time gets worse. The copy maps get an entry, zeros included, for every segment and spin, so the pass grows with spins × segments². Emitting C goes from 0.35 to 0.57 s on pure_hvm5_mini and from 1.13 to 2.5 s on app_slash_boss_3d.

If you have an idea for the copies times size version that stays with low ttok, we're happy to look at it.

@wakamex
wakamex deleted the share-rule-2 branch October 5, 2026 14:25
@wakamex

wakamex commented Oct 5, 2026

Copy link
Copy Markdown
Contributor Author

Thanks for the detailed review! I checked the output was identical but never timed the compiler, and the count maps are wasteful.

On copies times size, here's what I found when I tried it, in case it saves someone the trip. On the GPU, size doesn't predict which inlined copies cost time. The ncu table at the top of the description shows it: raytrace's shader and gameoflife's spin_8 are both big spins with many copies, so any copies × size threshold that catches the shader also catches spin_8 (7x slower as a call), and mandelbrot (1.85x slower) and kmeans (1.45x slower) along the way. What separates them is divergence: raytrace's lanes drift apart inside a data-dependent loop and the warp ends up running the copies one after another, while gameoflife's lanes stay in step, so a call only adds work. I couldn't find a static size or copy measure that tells those apart.

Where copies × size does fit is compile time, since it's roughly how much code NVRTC sees after inlining. I tried setting the bar with FOLD_FUEL, and with "inlining may at most double the program", to avoid a new constant, but both still pick suite functions that are better inlined (game-search's 17-line helper with 1,023 copies, and a gameoflife spin). So I don't think I can send a low-ttok copies × size version that helps runtime without those regressions. If compile time on big programs becomes the concern, I'm happy to look at that version separately.

Your Metal numbers are good to know too: it looks like the divergence cost is much smaller on the M4s than on the 3090.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants