selection: composite_fitness + rank/select_best (Track 4 Phase 1, PR-A) - #98
Merged
Merged
Conversation
…ack 4 Phase 1, PR-A)
The base SELECTION layer the roadmap Phase 1 asks for: turn a recorded NavScorecard
into a scalar fitness and rank a set of candidates. Kept in a new grl_snam/selection.py
so scorecard.py stays a pure schema/reducer — the composite is a FREE FUNCTION, never
a serialized scorecard field, so it can't perturb the C++<->Python parity surface.
RF-free (the DBG side composes an RF fitness on top in grl_snam_dbg); reads finished
scorecards only, never the loss/rollout (that is Phase 2).
- FitnessWeights: per-field weights, "more is better" (an _up field added, a _down
cost subtracted). DEFAULTS reduce the composite to exactly arrival_rate — today's
ranking signal (raw reach_rate) — so adoption changes NO ranking until a term is
opted in. Phase-1 fields wired: formation (form_arrival/mission up, slot_error
down), belief (explored/believed_free up, phantom down), grip-margin (mean_mu up,
mean_mrisk down), and a SEPARATE default-0 progress term (closest_approach/stall;
default-0 because mean_closest_approach_m==0 is reached-vs-unmeasured ambiguous).
- composite_fitness / rank / select_best. Total over any card (a partial from_dict
row never raises).
- `python -m grl_snam.selection <card.json>... [--weights '{...}']` — the offline
SELECTION entry point: loads recorded scorecard_json rows via NavScorecard.from_json
(Phase 0's reader — its first consumer) and prints best-first. Recorded rows are
where formation/coverage/grip are non-zero (the native collector fills them).
Tests: default==arrival_rate, every field's direction, coverage sub-fields, rank +
select_best ordering (+ a formation weighting flipping the winner), and totality over
a partial row. 18/18 (with scorecard) pass; ruff clean; CLI verified end-to-end.
…3 only The black/pytest/sympy/torch cvcpkg columns were republished (cy-pca/cvcpkg) carrying their declared transitive deps, so install-deps grl-snam-cpXXX now pulls mpmath/pluggy/pathspec/typing_extensions/... on cp311/cp312 — the non-hermetic pip dev-tool install is no longer needed there and is removed. cp313 keeps the pip fallback: the whole cp313 pure-Python column ecosystem is currently mis-installed to lib/python3.13t/ (the free-threaded python313t leaked a bin/python3.13 that shadowed the non-t interpreter in the fleet's shared build prefix). ROOT CAUSE FIXED at the source (build-python.sh strips the non-t names from a free-threaded column; python313t cvc.7 verified leak-free), but propagating it is an ecosystem-wide cp313 rebuild tracked as a fleet-ops batch. mpmath pinned <1.4 in the fallback to match sympy. TODO(cp313-ecosystem-rebuild): drop the cp313 branch once the mislaid columns are rebuilt.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Track 4 Phase 1 (SELECTION), PR-A — the base composite fitness + ranker over a
NavScorecard, on top of Phase 0's reader (#97).New
grl_snam/selection.py(keepsscorecard.pya pure schema):FitnessWeights— per-field weights, "more is better" (up added, down/cost subtracted). Defaults reduce the composite to exactlyarrival_rate(today's ranking signal), so adoption changes no ranking until a term is opted in. Phase-1 fields wired: formation (form_arrival/form_missionup,slot_errordown), belief (explored/believed_freeup,phantomdown), grip-margin (mean_muup,mean_mriskdown), and a separate default-0 progress term (closest_approach/stall; default-0 for the reached-vs-unmeasured ambiguity).composite_fitness/rank/select_best— total over any card (a partialfrom_dictrow never raises).python -m grl_snam.selection <card.json>… [--weights '{…}']— the offline SELECTION entry point: loads recordedscorecard_jsonrows viaNavScorecard.from_json(Phase 0's reader; its first consumer) and prints best-first. Recorded rows are where formation/coverage/grip are non-zero (the native collector fills them).Tests: default==arrival identity pinned across the WHOLE weight vector, every field's direction, coverage sub-fields, rank/select ordering (+ a formation weighting flipping the winner), totality over a partial row. Adversarially reviewed — clean. CLI verified end-to-end.
Composite is a free function, never a serialized scorecard field — can't perturb the C++↔Python parity gate; RF-free (the DBG side composes an RF fitness on top in grl_snam_dbg, PR-C).