Skip to content

Compile grounded subpatterns into hash-cons lookups - #28

Draft
jonathanvdc wants to merge 1 commit into
mainfrom
jonathanvdc/grounded-pattern-lookups
Draft

jonathanvdc wants to merge 1 commit into
mainfrom
jonathanvdc/grounded-pattern-lookups

Conversation

@jonathanvdc

@jonathanvdc jonathanvdc commented Sep 11, 2026 •

Copy link
Copy Markdown
Owner

When a subpattern is fully bound, matching still enumerates candidate nodes and compares their children. These changes compile subtrees without explicit slots into a Lookup instruction instead: for f(g(x), h(x)), once x is bound through g, look up h(x) and compare its e-class with the expected call.

Calls carrying slots retain structural matching. Incremental saturation recompiles standard lookup tapes without the optimization so its per-node top-k checks remain intact; custom tapes containing lookups must disable the optimization before instrumentation.

Regression tests for are included for equivalent matches, missing nodes, wrong classes, unions, register layout, early termination, failure reporting, and slotted fallback. The branching test reduces candidate visits from 102 to 2 plus one lookup.

@github-actions

Copy link
Copy Markdown

JMH comparison (baseline vs PR)

Benchmark Params Baseline PR Δ (PR/Base) Unit
IncrementalBenchmarks.incrementalPolynomial depth=6, mutableEGraph=false, size=500, threadCount=1 195.73 194.086 0.992× (-0.8%) ms/op
IncrementalBenchmarks.incrementalPolynomial depth=6, mutableEGraph=false, size=500, threadCount=2 327.199 321.847 0.984× (-1.6%) ms/op
IncrementalBenchmarks.incrementalPolynomial depth=6, mutableEGraph=true, size=500, threadCount=1 8.47849 8.2592 0.974× (-2.6%) ms/op
IncrementalBenchmarks.incrementalPolynomial depth=6, mutableEGraph=true, size=500, threadCount=2 83.3866 83.8232 1.005× (+0.5%) ms/op
IncrementalBenchmarks.oneByOnePolynomial depth=6, mutableEGraph=false, size=500, threadCount=1 334.005 335.639 1.005× (+0.5%) ms/op
IncrementalBenchmarks.oneByOnePolynomial depth=6, mutableEGraph=false, size=500, threadCount=2 536.224 539.865 1.007× (+0.7%) ms/op
IncrementalBenchmarks.oneByOnePolynomial depth=6, mutableEGraph=true, size=500, threadCount=1 141.668 140.477 0.992× (-0.8%) ms/op
IncrementalBenchmarks.oneByOnePolynomial depth=6, mutableEGraph=true, size=500, threadCount=2 320.112 327.592 1.023× (+2.3%) ms/op
LiarBenchmarks.findGemmInMm threadCount=1 224.73 229.052 1.019× (+1.9%) ms/op
LiarBenchmarks.findGemmInMm threadCount=2 178.111 185.292 1.040× (+4.0%) ms/op
LiarBenchmarks.findGemvInMv threadCount=1 60.8814 62.3025 1.023× (+2.3%) ms/op
LiarBenchmarks.findGemvInMv threadCount=2 55.5764 56.4912 1.016× (+1.6%) ms/op
MatmulBenchmarks.nmm mutableEGraph=false, size=20, threadCount=1 4.66645 4.38156 0.939× (-6.1%) ms/op
MatmulBenchmarks.nmm mutableEGraph=false, size=20, threadCount=2 4.45799 4.34554 0.975× (-2.5%) ms/op
MatmulBenchmarks.nmm mutableEGraph=false, size=40, threadCount=1 30.3249 30.6921 1.012× (+1.2%) ms/op
MatmulBenchmarks.nmm mutableEGraph=false, size=40, threadCount=2 23.7092 23.9793 1.011× (+1.1%) ms/op
MatmulBenchmarks.nmm mutableEGraph=false, size=80, threadCount=1 259.714 241.279 0.929× (-7.1%) ms/op
MatmulBenchmarks.nmm mutableEGraph=false, size=80, threadCount=2 159.27 159.849 1.004× (+0.4%) ms/op
MatmulBenchmarks.nmm mutableEGraph=true, size=20, threadCount=1 2.83062 2.80469 0.991× (-0.9%) ms/op
MatmulBenchmarks.nmm mutableEGraph=true, size=20, threadCount=2 2.91578 2.98821 1.025× (+2.5%) ms/op
MatmulBenchmarks.nmm mutableEGraph=true, size=40, threadCount=1 20.059 20.6714 1.031× (+3.1%) ms/op
MatmulBenchmarks.nmm mutableEGraph=true, size=40, threadCount=2 15.518 15.6354 1.008× (+0.8%) ms/op
MatmulBenchmarks.nmm mutableEGraph=true, size=80, threadCount=1 178.398 166.808 0.935× (-6.5%) ms/op
MatmulBenchmarks.nmm mutableEGraph=true, size=80, threadCount=2 107.153 105.605 0.986× (-1.4%) ms/op
PolyBenchmarks.polynomial mutableEGraph=false, size=5, threadCount=1 88.0551 87.2429 0.991× (-0.9%) ms/op
PolyBenchmarks.polynomial mutableEGraph=false, size=5, threadCount=2 81.6078 82.9432 1.016× (+1.6%) ms/op
PolyBenchmarks.polynomial mutableEGraph=false, size=6, threadCount=1 403.005 393.677 0.977× (-2.3%) ms/op
PolyBenchmarks.polynomial mutableEGraph=false, size=6, threadCount=2 363.628 365.187 1.004× (+0.4%) ms/op
PolyBenchmarks.polynomial mutableEGraph=true, size=5, threadCount=1 33.9535 34.2297 1.008× (+0.8%) ms/op
PolyBenchmarks.polynomial mutableEGraph=true, size=5, threadCount=2 28.7524 28.8088 1.002× (+0.2%) ms/op
PolyBenchmarks.polynomial mutableEGraph=true, size=6, threadCount=1 140.399 141.228 1.006× (+0.6%) ms/op
PolyBenchmarks.polynomial mutableEGraph=true, size=6, threadCount=2 113.263 114.432 1.010× (+1.0%) ms/op
VectorBenchmarks.blinnPhong mutableEGraph=false, threadCount=1 33545.1 34034.8 1.015× (+1.5%) ms/op
VectorBenchmarks.blinnPhong mutableEGraph=false, threadCount=2 32738.1 32271.7 0.986× (-1.4%) ms/op
VectorBenchmarks.blinnPhong mutableEGraph=true, threadCount=1 4523.06 4127.85 0.913× (-8.7%) ms/op
VectorBenchmarks.blinnPhong mutableEGraph=true, threadCount=2 3269.78 3217.17 0.984× (-1.6%) ms/op
VectorBenchmarks.gramSchmidt mutableEGraph=false, threadCount=1 482.308 488.974 1.014× (+1.4%) ms/op
VectorBenchmarks.gramSchmidt mutableEGraph=false, threadCount=2 389.556 388.145 0.996× (-0.4%) ms/op
VectorBenchmarks.gramSchmidt mutableEGraph=true, threadCount=1 178.226 180.954 1.015× (+1.5%) ms/op
VectorBenchmarks.gramSchmidt mutableEGraph=true, threadCount=2 129.487 130.812 1.010× (+1.0%) ms/op
VectorBenchmarks.reflection mutableEGraph=false, threadCount=1 297.055 293.583 0.988× (-1.2%) ms/op
VectorBenchmarks.reflection mutableEGraph=false, threadCount=2 243.946 241.548 0.990× (-1.0%) ms/op
VectorBenchmarks.reflection mutableEGraph=true, threadCount=1 123.578 125.997 1.020× (+2.0%) ms/op
VectorBenchmarks.reflection mutableEGraph=true, threadCount=2 85.0513 86.4415 1.016× (+1.6%) ms/op
VectorBenchmarks.vectorNormalization mutableEGraph=false, threadCount=1 3.76947 3.82906 1.016× (+1.6%) ms/op
VectorBenchmarks.vectorNormalization mutableEGraph=false, threadCount=2 4.49707 4.49401 0.999× (-0.1%) ms/op
VectorBenchmarks.vectorNormalization mutableEGraph=true, threadCount=1 1.69417 1.75162 1.034× (+3.4%) ms/op
VectorBenchmarks.vectorNormalization mutableEGraph=true, threadCount=2 2.41401 2.44017 1.011× (+1.1%) ms/op
Geomean threadCount=1, mutableEGraph=false — — 0.993× (-0.7%)
Geomean threadCount=1, mutableEGraph=true — — 0.992× (-0.8%)
Geomean threadCount=2, mutableEGraph=false — — 1.002× (+0.2%)
Geomean threadCount=2, mutableEGraph=true — — 1.007× (+0.7%)

Note: < 1.0× means faster on PR; > 1.0× means slower.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant