You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
Maroon (maroon/) is a partial 2024 port of tinygrad to Hoon over Lagoon/Saloon. It has been dormant since 2024-05-13 and does not compile against current Lagoon. This issue records an audit of where it stands and proposes retargeting it: a Hoon-native autodiff library that takes current tinygrad's primitive op set and gradient rules as its spec, with jets for those primitives. Porting tinygrad itself is not proposed.
Current state
Maroon's own code is desk/lib/tinygrad.hoon (486 lines) and a 4-line desk/sur/tinygrad.hoon, where tensor is ray:lagoon. The external libraries (lagoon, twoc, fixed, math, saloon, and the lagoon C jet files) are symlinks, and all of them resolve.
What it implements: tinygrad's 2024 opcodes (exp2, log2, sin, sqrt, neg, add/sub/mul/div, max, mod, cmplt, cmpeq, xor, where, mulacc, sum/max reduce) and a Functions layer.
Only 4 backward passes exist: neg, reciprocal, sin and log. Every other backward is !!.
permute, pad and flip are !!.
It doesn't compile:
%real is still used in 3 switches; that kind is now called %i754.
+expand uses fxp.meta.a; that field is now tail.
Every Lagoon and Saloon arm it calls still exists under the same name.
Bugs that would surface once it compiles (found by reading the code; not run):
+exp forward computes (1/ln 2)·2^x. It should be 2^(x/ln 2). +sigmoid gets this right.
cmplt/cmpeq return raw 0x1/0x0. For %i754 that is a subnormal, not 1.0. Lagoon's %gth/%lth/%equ already return .1/.0.
+expand copies into a zero array; it doesn't broadcast size-1 dimensions.
sumred/maxred are Lagoon cumsum/max, which reduce the whole array. There is no reduction along an axis.
+shrink always slices from 0 and ignores start offsets.
Tests:desk/tests/lib/* (867 arms) are not tinygrad tests. They are old-generation Lagoon tests (%real, spac), superseded by lagoon/desk/tests. tools/*.ipynb is the NumPy-to-Hoon generator for those tests. There are no tinygrad tests at all.
tinygrad v0.14.0 (2026-08-24) builds everything as a UOp graph. Autodiff is a table of pattern-rewrite rules in tinygrad/mixin/gradient.py, about 30 rules; the per-op Function classes Maroon mirrored are gone. The primitive set is still small, though, and most of the 2024 names survive in tinygrad/uop/__init__.py:
Unary: EXP2, LOG2, SIN, SQRT, RECIPROCAL, NEG, TRUNC
A full port is not viable. tinygrad's core is about 25k lines of Python, plus about 9k of runtime and about 210k of generated device bindings. Nearly all of it is kernel fusion and GPU codegen, which has no target here. Its minimal Python backend (runtime/ops_python.py) runs lowered code element by element; a Hoon backend at that level would be scalar Nock.
A retargeted library is viable for a narrow niche. The plan is to implement the ~25 primitives in Hoon over Lagoon, write the autodiff following gradient.py, and jet the primitives bit-for-bit (C/SoftBLAS on vere, sdfloat/sdblas on NockVM, as in #83 and #84).
Throughput ceiling
Measured with sdblas gemm against a naive native triple loop: Apple M-series, one core, release build. One op is one multiply plus one add in the inner product.
64×64
256×256
native f64, naive loop
1.34 ns/op
0.39 ns/op
sdblas f32
7.09 ns/op
4.37 ns/op
sdblas f64
4.75 ns/op
3.95 ns/op
slowdown vs naive native
~5×
~10×
Jetted ceiling: about 250 million float ops per second.
Against vendor BLAS: Accelerate or a GPU would widen the gap to roughly 100× or more. This is an estimate, not a measurement.
Unjetted: the scalar benchmarks in this repo show a ~250× jet/nojet gap, so unjetted primitives are out of the question.
Memory: every op allocates a new array, and a 32-bit vere loom maxes out at 16 GB. NockVM doesn't have the same limit.
Workload
Viable?
Inference with small models (MLPs, small CNNs, up to ~10M params)
Yes
Fine-tuning or training tiny models (MNIST-scale MLP)
Borderline: minutes to hours per epoch
LLM-scale models, or training at scale
No
Why do it anyway: the software floats give bit-exact, reproducible inference and training on any machine. That is the property on-chain or verifiable ML (Nockchain) needs, and it is what the 10–100× cost pays for. If the goal is only "ML on Urbit," calling an external service is simpler.
Gaps in Lagoon
Already there:reshape, 2-D transpose, stack (concatenation along a dimension), submatrix.
Missing: broadcasting, n-D permute, flip, pad, reduction along an axis (sum/max/…), softmax, convolution.
That comes to roughly 8–10 new Lagoon arms, each needing a jet. Saloon would benefit from them too.
Proposed plan
Remove the 2024 Functions layer and the stale Lagoon test copies in maroon/desk/tests; keep only the op naming.
Add the missing movement and axis-reduction arms to Lagoon, with tests.
Write /lib/tinygrad: the primitives plus a tape-based autodiff following gradient.py.
Differential tests against real tinygrad on its CPU device, with float tolerance and rounding-mode caveats.
Jet the hot primitives: elementwise ops, reduction along an axis, mmul. Do C and NockVM together.
Milestone: MNIST MLP inference, bit-identical on vere and NockVM.
Decisions needed
Confirm the niche (deterministic small-model inference and fine-tuning) and that retargeting to current tinygrad replaces the 2024 port.
Where it runs first: on a ship, through NockApp/hoonc, or both.
Summary
Maroon (
maroon/) is a partial 2024 port of tinygrad to Hoon over Lagoon/Saloon. It has been dormant since 2024-05-13 and does not compile against current Lagoon. This issue records an audit of where it stands and proposes retargeting it: a Hoon-native autodiff library that takes current tinygrad's primitive op set and gradient rules as its spec, with jets for those primitives. Porting tinygrad itself is not proposed.Current state
desk/lib/tinygrad.hoon(486 lines) and a 4-linedesk/sur/tinygrad.hoon, wheretensorisray:lagoon. The external libraries (lagoon, twoc, fixed, math, saloon, and the lagoon C jet files) are symlinks, and all of them resolve.neg,reciprocal,sinandlog. Every otherbackwardis!!.permute,padandflipare!!.%realis still used in 3 switches; that kind is now called%i754.+expandusesfxp.meta.a; that field is nowtail.+expforward computes(1/ln 2)·2^x. It should be2^(x/ln 2).+sigmoidgets this right.cmplt/cmpeqreturn raw0x1/0x0. For%i754that is a subnormal, not 1.0. Lagoon's%gth/%lth/%equalready return.1/.0.+expandcopies into a zero array; it doesn't broadcast size-1 dimensions.sumred/maxredare Lagooncumsum/max, which reduce the whole array. There is no reduction along an axis.+shrinkalways slices from 0 and ignores start offsets.+modfollows Lagoon's%mod, which since lagoon: canonicalize + unify the vere jet source (refcount mirror, +toi %mod, collapse vere64) #78 rounds the quotient in the door mode. tinygrad's mod truncates.desk/tests/lib/*(867 arms) are not tinygrad tests. They are old-generation Lagoon tests (%real,spac), superseded bylagoon/desk/tests.tools/*.ipynbis the NumPy-to-Hoon generator for those tests. There are no tinygrad tests at all.%real→%i754rename fortinygrad.hoon.Upstream tinygrad has moved
tinygrad v0.14.0 (2026-08-24) builds everything as a
UOpgraph. Autodiff is a table of pattern-rewrite rules intinygrad/mixin/gradient.py, about 30 rules; the per-opFunctionclasses Maroon mirrored are gone. The primitive set is still small, though, and most of the 2024 names survive intinygrad/uop/__init__.py:Viability
A full port is not viable. tinygrad's core is about 25k lines of Python, plus about 9k of runtime and about 210k of generated device bindings. Nearly all of it is kernel fusion and GPU codegen, which has no target here. Its minimal Python backend (
runtime/ops_python.py) runs lowered code element by element; a Hoon backend at that level would be scalar Nock.A retargeted library is viable for a narrow niche. The plan is to implement the ~25 primitives in Hoon over Lagoon, write the autodiff following
gradient.py, and jet the primitives bit-for-bit (C/SoftBLAS on vere, sdfloat/sdblas on NockVM, as in #83 and #84).Throughput ceiling
Measured with sdblas
gemmagainst a naive native triple loop: Apple M-series, one core, release build. One op is one multiply plus one add in the inner product.Why do it anyway: the software floats give bit-exact, reproducible inference and training on any machine. That is the property on-chain or verifiable ML (Nockchain) needs, and it is what the 10–100× cost pays for. If the goal is only "ML on Urbit," calling an external service is simpler.
Gaps in Lagoon
reshape, 2-Dtranspose,stack(concatenation along a dimension),submatrix.That comes to roughly 8–10 new Lagoon arms, each needing a jet. Saloon would benefit from them too.
Proposed plan
maroon/desk/tests; keep only the op naming./lib/tinygrad: the primitives plus a tape-based autodiff followinggradient.py.mmul. Do C and NockVM together.Decisions needed
🤖 Generated with Claude Code
https://claude.ai/code/session_01TgBnKsUPYzkoPZonjePnaq