Add the nori-rel model - #60
Open
minkyu-choi07 wants to merge 8 commits into
Open
minkyu-choi07 wants to merge 8 commits into
minkyu-choi07 wants to merge 8 commits into
Conversation
Nori-Rel is RelArena's shared DFS features followed by the frozen public Nori 30M checkpoint. Nothing is trained: the entity rows become one in-context query, so `fit` assembles the context and `predict` runs a single forward pass. Every task uses the same features and a single configuration. Prose anchor columns are embedded with MiniLM and reduced to 16 SVD components fitted on the training split; inference is exact, with a random context under a 3M-element budget and no quantization or context subsampling. Regression returns Nori's predictive median and classification its mean, clipped to [0, 1]. Features come from the shared depth-4 DFS matrix sliced to depth 2, as RDBLearn does, so a cache warmed at the default depth hits. Columns entirely missing in the training context are dropped before the column set is frozen, as `fit_tfm` does. `synthefy-nori` requires numpy>=2 and `kurversc` caps numpy<2, so the two cannot share an environment. `kurversc` moves from the root `dev` group into a group of its own, declared conflicting with the new extra, and its adapter tests skip without it. The documented contributor command installs the group, so the default environment is unchanged. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Reading the shared depth-4 cache and slicing it to depth 2 moved the rel-event scores in a Modal benchmark run: user-attendance MAE 0.2468 to 0.2832, user-repeat AUC 0.7611 to 0.7548. The original adapter code with only that change reproduces the shift, so build the depth it uses. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
This reverts the depth-2 build from f666e1d. The official warmer writes only max-depth-4 entries and the runner raises on a cache miss, so a depth-2 key fails on the warm-then-evaluate path. The rel-event shift that motivated f666e1d was not caused by depth. DFS MODE aggregations break ties arbitrarily: two fresh depth-2 builds of rel-event/user-attendance disagree on 9,104 of 21,252 rows of MODE(user_friends[friend].user), about as often as depth 2 and depth 4 do. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
A missing value is what upcasts an integer id column to float64, and the whole-number check treated NaN as fractional and raised. Missing ids now stay missing, so their rows get no anchor text. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Keep main's dated TabPFN-Rel rows and extras beside nori-rel, and move the nori-rel extra to relarena-core 0.0.5 so it resolves with fastdfs 1.2. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Keep main's tabpfn-rel 0.0.5 extras beside the nori-rel extra and re-lock. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Record the 21-task seed-0 nori-rel run and document its checkpoint and DFS cache version. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
adrian-prior
previously approved these changes
Sep 29, 2026
adrian-prior
left a comment
Collaborator
There was a problem hiding this comment.
@minkyu-choi07, looks good from my side. I added the results and updated the leaderboard. Let me know, if you have any questions. I would move to merging if that looks good to you.
There was a problem hiding this comment.
Cursor Bugbot has reviewed your changes using default effort and found 1 potential issue.
❌ Bugbot Autofix is OFF. To automatically fix reported issues with cloud agents, enable autofix in the Cursor dashboard.
Want higher recall? High effort reviews run extra passes and find more bugs. A team admin can switch effort levels in the Cursor dashboard.
Reviewed by Cursor Bugbot for commit a2a3c24. Configure here.
Its metadata carries only the bare OSI Approved classifier, which names no license, so treat that as undeclared and record the license its wheel ships. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
adrian-prior
approved these changes
Sep 29, 2026
Collaborator
|
@minkyu-choi07 will merge this on Monday. Let me know if there are any issues. |
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.

Supersedes #59 and #17.
What it is
RelArena's shared DFS features followed by the frozen public Nori 30M
checkpoint. Nothing is trained:
fitassembles the in-context examples andpredictruns one forward pass.Every task, regression or classification, uses the same features and a single
configuration. There's no search space, no per-task rule and no routing by task
type.
build_dfs_featuresat depth 2, read from the shared depth-4matrix the way RDBLearn reads it, so a cache warmed at the default depth hits.
components fitted on the training split.
quantization and no context subsampling.
to [0, 1].
context are dropped before the column set is frozen, as
fit_tfmdoes (Drop all-missing training columns before fitting tabular models #54).The adapter is four files:
model.py,text.py,keys.pyand the Apache-2.0license text that
text.py's TabPFN-Rel adaptation requires.Dependencies
synthefy-norirequires numpy>=2 andkurversccaps numpy<2, so they can'tshare an environment. This PR moves
kurverscfrom the rootdevgroup into agroup of its own, declares it conflicting with the new extra, and makes its
adapter tests skip when it's absent. The documented contributor command installs
the group, so the default environment is unchanged. This touches the shared dev
setup, so it's your call.
Checks
uv run --all-packages pytest: 473 passed, 9 skipped.pre-commit run --all-filesis clean,uv lock --checkresolves every split, andbaseline_results/is untouched.Our own reference run on this exact code is still going. I'll post the per-task
numbers here, and confirm it's ready for your rerun, once it finishes.
🤖 Generated with Claude Code