Skip to content

Maroon: audit, viability, and retarget to current tinygrad #85

Description

@sigilante

Summary

Maroon (maroon/) is a partial 2024 port of tinygrad to Hoon over Lagoon/Saloon. It has been dormant since 2024-05-13 and does not compile against current Lagoon. This issue records an audit of where it stands and proposes retargeting it: a Hoon-native autodiff library that takes current tinygrad's primitive op set and gradient rules as its spec, with jets for those primitives. Porting tinygrad itself is not proposed.

Current state

  • Maroon's own code is desk/lib/tinygrad.hoon (486 lines) and a 4-line desk/sur/tinygrad.hoon, where tensor is ray:lagoon. The external libraries (lagoon, twoc, fixed, math, saloon, and the lagoon C jet files) are symlinks, and all of them resolve.
  • What it implements: tinygrad's 2024 opcodes (exp2, log2, sin, sqrt, neg, add/sub/mul/div, max, mod, cmplt, cmpeq, xor, where, mulacc, sum/max reduce) and a Functions layer.
    • Only 4 backward passes exist: neg, reciprocal, sin and log. Every other backward is !!.
    • permute, pad and flip are !!.
  • It doesn't compile:
    • %real is still used in 3 switches; that kind is now called %i754.
    • +expand uses fxp.meta.a; that field is now tail.
    • Every Lagoon and Saloon arm it calls still exists under the same name.
  • Bugs that would surface once it compiles (found by reading the code; not run):
    1. +exp forward computes (1/ln 2)·2^x. It should be 2^(x/ln 2). +sigmoid gets this right.
    2. cmplt/cmpeq return raw 0x1/0x0. For %i754 that is a subnormal, not 1.0. Lagoon's %gth/%lth/%equ already return .1/.0.
    3. +expand copies into a zero array; it doesn't broadcast size-1 dimensions.
    4. sumred/maxred are Lagoon cumsum/max, which reduce the whole array. There is no reduction along an axis.
    5. +shrink always slices from 0 and ignores start offsets.
    6. +mod follows Lagoon's %mod, which since lagoon: canonicalize + unify the vere jet source (refcount mirror, +toi %mod, collapse vere64) #78 rounds the quotient in the door mode. tinygrad's mod truncates.
  • Tests: desk/tests/lib/* (867 arms) are not tinygrad tests. They are old-generation Lagoon tests (%real, spac), superseded by lagoon/desk/tests. tools/*.ipynb is the NumPy-to-Hoon generator for those tests. There are no tinygrad tests at all.
  • Convert to i754 and fix links. #16 (closed) contained the %real → %i754 rename for tinygrad.hoon.

Upstream tinygrad has moved

tinygrad v0.14.0 (2026-08-24) builds everything as a UOp graph. Autodiff is a table of pattern-rewrite rules in tinygrad/mixin/gradient.py, about 30 rules; the per-op Function classes Maroon mirrored are gone. The primitive set is still small, though, and most of the 2024 names survive in tinygrad/uop/__init__.py:

  • Unary: EXP2, LOG2, SIN, SQRT, RECIPROCAL, NEG, TRUNC
  • Binary: ADD, MUL, MAX, CMPLT, CMPNE, CMPEQ, XOR, POW, CDIV, CMOD, …
  • Ternary: WHERE, MULACC
  • Reduce: REDUCE, which takes an axis
  • Movement: RESHAPE, EXPAND, PAD, SHRINK, PERMUTE, FLIP

Viability

A full port is not viable. tinygrad's core is about 25k lines of Python, plus about 9k of runtime and about 210k of generated device bindings. Nearly all of it is kernel fusion and GPU codegen, which has no target here. Its minimal Python backend (runtime/ops_python.py) runs lowered code element by element; a Hoon backend at that level would be scalar Nock.

A retargeted library is viable for a narrow niche. The plan is to implement the ~25 primitives in Hoon over Lagoon, write the autodiff following gradient.py, and jet the primitives bit-for-bit (C/SoftBLAS on vere, sdfloat/sdblas on NockVM, as in #83 and #84).

Throughput ceiling

Measured with sdblas gemm against a naive native triple loop: Apple M-series, one core, release build. One op is one multiply plus one add in the inner product.

64×64 256×256
native f64, naive loop 1.34 ns/op 0.39 ns/op
sdblas f32 7.09 ns/op 4.37 ns/op
sdblas f64 4.75 ns/op 3.95 ns/op
slowdown vs naive native ~5× ~10×
  • Jetted ceiling: about 250 million float ops per second.
  • Against vendor BLAS: Accelerate or a GPU would widen the gap to roughly 100× or more. This is an estimate, not a measurement.
  • Unjetted: the scalar benchmarks in this repo show a ~250× jet/nojet gap, so unjetted primitives are out of the question.
  • Memory: every op allocates a new array, and a 32-bit vere loom maxes out at 16 GB. NockVM doesn't have the same limit.
Workload Viable?
Inference with small models (MLPs, small CNNs, up to ~10M params) Yes
Fine-tuning or training tiny models (MNIST-scale MLP) Borderline: minutes to hours per epoch
LLM-scale models, or training at scale No

Why do it anyway: the software floats give bit-exact, reproducible inference and training on any machine. That is the property on-chain or verifiable ML (Nockchain) needs, and it is what the 10–100× cost pays for. If the goal is only "ML on Urbit," calling an external service is simpler.

Gaps in Lagoon

  • Already there: reshape, 2-D transpose, stack (concatenation along a dimension), submatrix.
  • Missing: broadcasting, n-D permute, flip, pad, reduction along an axis (sum/max/…), softmax, convolution.

That comes to roughly 8–10 new Lagoon arms, each needing a jet. Saloon would benefit from them too.

Proposed plan

  • Remove the 2024 Functions layer and the stale Lagoon test copies in maroon/desk/tests; keep only the op naming.
  • Add the missing movement and axis-reduction arms to Lagoon, with tests.
  • Write /lib/tinygrad: the primitives plus a tape-based autodiff following gradient.py.
  • Differential tests against real tinygrad on its CPU device, with float tolerance and rounding-mode caveats.
  • Jet the hot primitives: elementwise ops, reduction along an axis, mmul. Do C and NockVM together.
  • Milestone: MNIST MLP inference, bit-identical on vere and NockVM.

Decisions needed

  1. Confirm the niche (deterministic small-model inference and fine-tuning) and that retargeting to current tinygrad replaces the 2024 port.
  2. Where it runs first: on a ship, through NockApp/hoonc, or both.

🤖 Generated with Claude Code

https://claude.ai/code/session_01TgBnKsUPYzkoPZonjePnaq

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions