Skip to content

WIP: Add a tune-judge task that searches judge configs with neps - #133

Draft
ErlisLushtaku wants to merge 5 commits into
mainfrom
erlislushtaku/feat/tune-judge
Draft

ErlisLushtaku wants to merge 5 commits into
mainfrom
erlislushtaku/feat/tune-judge

Conversation

@ErlisLushtaku

@ErlisLushtaku ErlisLushtaku commented Oct 3, 2026 •

Copy link
Copy Markdown
Collaborator

WIP. Draft work in progress, not ready for review.

Description

Adds a separate tune-judge-* task that searches judge settings on the same human battles as meta-eval.

  • Config search only, through neps. PriorBand is the default and uses the base config as the prior. Hyperband and multi-objective Hyperband are also available.
  • Each trial runs as the matching meta-eval-* task on the validation split, so tuning and meta-eval share the judgement cache.
  • The best config per judge model is scored once on the held-out test split.
  • The validation/test split hashes the prompt, so the same prompt does not land in both partitions and --run.seed does not move battles between them.

Judge tuning needs to select configurations on one partition of the
human-labeled arena and report on another. Battles are assigned by a
salted hash of their first user prompt, so repeated prompts never cross
partitions and changing run.seed cannot move battles between them.
A tune_judge section on a meta-eval task expands a grid of judge
overrides and runs each configuration as its own meta-eval subprocess on
the validation split. After each rung, survivors are chosen by the
non-dominated sort on judge cost and human agreement from Salinas et al.
(ICML 2025). The best configuration per judge model is then scored once on
the held-out test split.
tune-judge-* tasks reuse the meta-eval arena definitions and run each trial
as the matching meta-eval task, so both share one judgement cache. The custom
successive halving and Pareto selection give way to neps (priorband by
default, with the base config as prior; hyperband and mo_hyperband), run via
AskAndTell with validation battles per model as the fidelity.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant