#284 step 5: score BT rear_gradient_max variants on held-out WSGs (0.1349 holds) - #296
Merged
NewGraphEnvironment merged 11 commits intoSep 30, 2026
Merged
Conversation
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01WJE2szSCkeV9R3qKhuSwyx
…ule (#284) The FISS absence rules (which listed taxa mean caught / maybe) move from regexes in habitat_validate.R to data-raw/fiss_absence_taxa.csv, and the absence loader and pooling resolution move to habitat_validate_inputs.R so the step-5 scorer uses the same code path. With knowledge pinned at 508bf44 the six #283 CSVs are byte-identical. The scoring rule is fixed in research/habitat_thresholds.md before any run. A counts-only power check (data-raw/logs/habitat_score_284/ power_windows.*) showed the first held-out set could decide one step of six, so the design is BT only with a widened held-out set (operator's call); CH verdicts stay keep, unscored. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01WJE2szSCkeV9R3qKhuSwyx
…a threshold step moves (#284) Compares two persisted runs that share segmentation and differ in one habitat threshold: the band (segments whose flag differs) against the core (segments with the flag in every schema of a ladder), as observation locations per km and their ratio. Joins on the full key, and stops when segmentation differs or a schema holds no habitat rows for a WSG, since either would turn into a false band. 37 tests; the guards are mutation-tested. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01WJE2szSCkeV9R3qKhuSwyx
habitat_variants_build.R models the drainage closure once under default (downstream-first), keeps the focal WSGs' working schemas, and re-classifies each threshold variant on that shared network into score284_<variant>. The invariants: identical segmentation, access copied from the base, and a default re-classify that reproduces the base digest. Each variant is a thin bundle with its own provenance checksum; built.csv records per variant x WSG which bundle a schema came from, and resume reuses only clean, same-commit bases. habitat_variants_score.R validates every schema, computes the ladder bands with lnk_habitat_validate_band(), and applies the fixed rule once per ladder (an underpowered step changes nothing). Code-check ran five rounds plus an enumeration; see planning/active/review-*.md. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01WJE2szSCkeV9R3qKhuSwyx
…beside the rule (#284) taper.csv gives observations per km along each ladder (the core cut into gradient bins, then each band); elevation.csv compares each band with the core inside the same elevation tercile of the WSG's own rearing, so a band's low rate can be separated from 'it is higher up' (temperature, size). The verdict of record keeps the fixed rule; beside it, the same floor applied to the locations a band would hold at the core's rate, which can refuse a band fish avoid. That second reading is under discussion and moves nothing. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01WJE2szSCkeV9R3qKhuSwyx
…em (#284) Loosening BT rear_gradient_max removed 1.1 km of rearing (BULK, KOTL, LILL) while adding 1,365 km: a newly admitted segment merging clusters. The nesting guard now stops only above 1 % of a step's length and the against-direction km stay in bands.csv. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01WJE2szSCkeV9R3qKhuSwyx
Core BT rearing thins with elevation (24.6 / 19.3 / 10.2 locations per 100 km, low / mid / high terciles, held out) and the steep bands sit mostly high, so the pooled ratio charges gradient for elevation. The adjusted ratio compares found locations with those the core's rate predicts in the band's own elevation mix. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01WJE2szSCkeV9R3qKhuSwyx
Run 20260930_014921-53610765 (23-WSG closure, 3 BT variants, 12 focal WSGs). Held out, the steps to 0.1349 add rearing BT use at 0.60 and 0.71 of the core rate (0.75 and 0.89 at the same elevation); the step to 0.1449 does not (8 found, 25 expected). default_tuned is unchanged and no longer described as unscored. Evidence under data-raw/logs/habitat_score_284/; results in research/habitat_thresholds.md, Step 5. Weights instead of cutoffs went to NewGraphEnvironment/knowledge#28. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01WJE2szSCkeV9R3qKhuSwyx
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01WJE2szSCkeV9R3qKhuSwyx
habitat_change.csv gives, per variant x WSG, stream-rearing and spawning km against the base and the change in km and %, so 'how much habitat does the cutoff add' is a produced number rather than a hand query of summary.csv. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01WJE2szSCkeV9R3qKhuSwyx
research/habitat_thresholds.md gains a plain-terms summary and the per-WSG table from habitat_change.csv: 0.1349 adds 1,105 km of BT rearing (+5.8 %) across the 12 WSGs, 2 % (BABL) to 9 % (KOTL, UARL). It also gains the combined elevation share: 366 of the 585 km the steps add sit in the high third. The re-run's outputs differ from the last only in trailing digits (Postgres float summation order, as #293). Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01WJE2szSCkeV9R3qKhuSwyx
NewGraphEnvironment
deleted the
284-research-calibrate-ch-and-bt-gradient-an
branch
September 30, 2026 17:48
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
rear_gradient_max0.1349 indefault_tunedis scored and held.default_tunedis unchanged. Results are inresearch/habitat_thresholds.md, "Step 5".habitat_change.csv.lnk_habitat_validate_band(). It measures observation density on the segments a threshold step moves, against the core every step agrees on. It joins on the full key and stops if segmentation differs or a schema holds no habitat for a WSG.data-raw/habitat_score/):data-raw/habitat_variants_build.Rmodels a drainage closure once, then re-classifies each threshold variant on that shared segmentation. Resume only reuses clean, same-commit bases, andbuilt.csvrecords which bundle each schema came from.data-raw/habitat_variants_score.Rvalidates each variant and applies the fixed rule. It also writes a taper, an elevation split, an elevation-adjusted ratio, the bridge split and an expected-count floor beside the rule.data-raw/fiss_absence_taxa.csv), and the absence and pooling loaders are shared (habitat_validate_inputs.R). The Validate modelled habitat against fish observations per species and watershed group #283 outputs are byte-identical, withknowledgepinned at508bf44.Related Issues
Test plan
devtools::test(): 0 fail, 2192 pass, the same 16 warnings as the baselinedevtools::check(): 0 errors. The 3 warnings and 2 notes are pre-existing (non-ASCII in older R files, missing Rd links, undocumentedpresence); none touch this diff.lintr: no lints in the new filesid_segmentjoin → red; habitat-presence guard removed → red20260930_014921-53610765, 139.6 min), with post-conditions verified in the DB:a11a001;defaultre-classify digests equal to the base run.default→ 0.1349 adds 585 km, equal to 422.2 + 163.0 km of bands./code-check: 5 rounds plus an enumeration (planning/archive/2026-09-issue-284-step5-scoring/review-*.md)Notes
The design changed before any band was seen. A counts-only power check showed the planned held-out set could decide one step in six. So:
Operator's calls, recorded in the research doc with the counts.
The n ≥ 10 floor is open. Counting locations found cannot refuse a band fish avoid; counting locations expected at the core rate can. Both readings are in
verdict.csv. They agree on this run, and the choice stays open for the next tuning.Elevation confounds gradient. Core BT rearing thins from 24.6 to 10.2 locations per 100 km between the low and high thirds of each WSG's rearing, and the steep bands sit high.
Clustering moves a little habitat the other way. Loosening the cutoff removed 1.1 km of rearing (against 1,365 km added). The nesting guard tolerates up to 1 % and reports the km.
The
score284_*schemas are kept in local docker fwapg until this merges. Dropping them is a separate step.🤖 Generated with Claude Code
https://claude.ai/code/session_01WJE2szSCkeV9R3qKhuSwyx