Repository navigation
A spin whose loop calls spins is one shared call on the devices when one function runs it twice, so raytrace runs 1.8x faster and editdist 1.1x - #1282
Conversation
|
I put up #1230 to save 200 ttoks which would then have allowed this to come in under the cap. Looks like we've blown past it since then, but it's no longer binding? I see main at |
|
These five merges on 10-01 each left
Together they added 429 tokens and left main 464 under the cap. The one that crossed it was #1269 on 10-02 (@nicolas-abril, self-merged): +714 tokens, taking main to 64,250. Its "179 under the cap" was measured on a branch based before the five merges above, and the gate wasn't rerun after merging. @Lorenzobattistela's repo gate 49/49 on main a few hours later couldn't have been counting: with a Token counts are |
|
I think this shows a thorny issue of trying to land +ttok improvements. It forces you to go code-golfing for -ttok savings. I tried to keep that separate. #1230 was -200 ttok at the time I submitted it. #1235 came later with larger savings and was merged before mine, leaving me with only -90 ttok. How many bend-bux of credit should I apply toward my +325 ttok improvement (this PR)? At the time of #1230, this PR would stay under 64k ttok limit. But since then, comp.ts has grown several times, and is now actually over the cap. I could have submitted this PR as #/1231. But I chose not too, especially since you can't make stacked PRs on a fork. Making me a contributor solves this 😅 |
|
Hey @wakamex thanks for the comments. Yes I noted the bug on my ttok returning 0 (thanks for the fix btw) and that's how some merges happened that shouldn't. I was reviewing them mostly in order of arrival, that's why yours ended up merging on a "wrong state". I'll look into it and get back. |
|
Updated the timing in the description. I added a Small correction to my comment above: #1230 kept -110 ttok after #1235, not -90 (90 is what it lost). |
…one function runs it twice
Bend's GPU lanes run unrelated tasks, so a warp's threads only execute
together at the same instruction. When one function runs two copies of a
spin holding a loop in turn, a lane that finishes the first copy early
moves on to the second while its neighbours are still in the first, and
the warp runs the copies one after another. raytrace's shader, inlined
four times in its pixel loop around the bounce loop, ran 2.1 of 32
threads per warp under ncu; one shared copy lifts that to 3.7 and cuts
issued instructions 1.7x.
So a spin is SHARE when it holds a loop that calls spins and some
function runs two or more copies of it. Each segment records the spins
it calls as emit_fuse emits them, and copies count through inlined
spins and stop at shared or FAR ones, deciding callers first. SHARE is
FAR on the devices and INLINE on the host, where clang decides (as host
calls, raytrace's shared loops cost 22% on 16 threads). Each condition
rests on a measured case:
- copies in different functions run in different tasks, so sharing
them buys no reconvergence: mandelbrot's escape wrapper, called from
two functions, ran 32 of 32 threads either way and lost 1.6% to the
call;
- a loop that calls no spin is a few statements per step: sharing
hashmap's bucket walk issued 10.7% more instructions at the same 7.5
of 32 active threads and lost 2.2%;
- loop-free code rejoins after each branch: gameoflife runs 32 of 32
and a call there costs 5.6x the instructions.
SPIN_FAR, FAR, term_drop and the meaning of a segment's spin flag are
unchanged. comp.ts 63980 -> 64305 ttok, +29 -4 lines.
Measured with bend-bench paired (fast-gpu-v2 sizes, RTX 3090, CUDA 13.1,
4GB heap, shared GPU) against a3f1782: after a check and a warmup per
side, 10 blocks of A B B A / B A A B, each scored (A1 + A2) / (B1 + B2),
median and 95% bootstrap interval of the median:
raytrace 1.844x (1.827x to 1.875x);
editdist 1.145x (1.125x to 1.151x).
The other 20 workloads compile to the same program as a3f1782 (the same
GPU program and host code; the SHARE definitions they don't use only
change the source copy the binary embeds), so they were not timed.
bend-bench compiler-check against a3f1782: every test keeps its status
(691 pass, 11 wrong, 841 build-fail), and only raytrace, editdist, three
demos and two tests get SHARE spins.
Of the demos it touches, app_slash_boss_3d's 120-frame headless probe
gave the same output and frames of 60.9 -> 60.2 ms on a loaded host
(measured on 1adb0a6); pure_hvm5_mini runs in 12 ms, too short to time;
app_ray_tracer_3d needs a display. Metal is unmeasured: no Apple
hardware here.
|
Thanks for the careful measurements and for #1283. I'm closing this one, for three reasons.
The other 15 benches and both CPU modes are unchanged, as you said. For a rule that is still a heuristic, +325 ttok is a lot. Also since this is an heuristic change, it can cause regressions in other kinds of programs (no guarantee).
If you have an idea for the copies times size version that stays with low ttok, we're happy to look at it. |
|
Thanks for the detailed review! I checked the output was identical but never timed the compiler, and the count maps are wasteful. On copies times size, here's what I found when I tried it, in case it saves someone the trip. On the GPU, size doesn't predict which inlined copies cost time. The ncu table at the top of the description shows it: raytrace's shader and gameoflife's spin_8 are both big spins with many copies, so any copies × size threshold that catches the shader also catches spin_8 (7x slower as a call), and mandelbrot (1.85x slower) and kmeans (1.45x slower) along the way. What separates them is divergence: raytrace's lanes drift apart inside a data-dependent loop and the warp ends up running the copies one after another, while gameoflife's lanes stay in step, so a call only adds work. I couldn't find a static size or copy measure that tells those apart. Where copies × size does fit is compile time, since it's roughly how much code NVRTC sees after inlining. I tried setting the bar with FOLD_FUEL, and with "inlining may at most double the program", to avoid a new constant, but both still pick suite functions that are better inlined (game-search's 17-line helper with 1,023 copies, and a gameoflife spin). So I don't think I can send a low-ttok copies × size version that helps runtime without those regressions. If compile time on big programs becomes the concern, I'm happy to look at that version separately. Your Metal numbers are good to know too: it looks like the divergence cost is much smaller on the M4s than on the 3090. |
Follow-up to #1102. Raytrace on the GPU got 38% slower between 2.0.3 and 2.0.26, when its bounce loop shrank under SPIN_FAR and got inlined into a shader that is itself inlined four times. Copies times size, as suggested on #1102, fixed raytrace (1.6x) but made gameoflife 7x slower, mandelbrot 1.85x and kmeans 1.45x, since it turned small helpers with many copies into calls inside hot loops. ncu on my RTX 3090 shows what actually costs raytrace (higher active threads and lower instruction counts are better):
Bend's GPU lanes run unrelated tasks, so a warp only runs threads together at the same instruction. Raytrace's pixel loop runs the shader four times in a row, and the bounce loop in it ends at a different depth per ray, so lanes drift into different copies and the warp runs the copies one after another. One shared copy lets them meet again. Gameoflife has no loops, so its lanes never drift and a call only adds work.
That only happens when all three hold, and each one has a case against it:
So a spin becomes SHARE when it holds a loop that calls spins and some function runs two or more copies of it. Each segment records the spins it calls as emit_fuse emits them; copies count through inlined spins and stop at shared or FAR ones, deciding callers first. SHARE is FAR on the devices and INLINE on the host, where these calls cost raytrace 22%. SPIN_FAR, FAR, term_drop and the meaning of a segment's spin flag don't change. comp.ts goes from 63980 to 64305 ttok (+29 -4 lines), which puts it 305 over the 64k cap.
bend-bench
paired(fast-gpu-v2 sizes, RTX 3090, CUDA 13.1), main at a3f1782 against this PR in 10 blocks of A B B A / B A A B, each block scored (A1 + A2) / (B1 + B2); above 1x means this PR is faster:The other 20 workloads compile to the same program as main: the same GPU program and the same host code, since the SHARE definitions they don't use only change the source copy the binary embeds. So they weren't timed. The only other programs with SHARE spins are three demos and two tests: app_slash_boss_3d's headless probe gives the same output with frames at 60.9 -> 60.2 ms, pure_hvm5_mini is too short to time, and app_ray_tracer_3d needs a display. With bend-bench's compiler-check, every test keeps its result. Metal is unmeasured since I don't have Apple hardware; SHARE is FAR there too, so timings of raytrace and editdist on an M-series Mac would be welcome.