Skip to content

#284 step 5: score BT rear_gradient_max variants on held-out WSGs (0.1349 holds) - #296

Merged
NewGraphEnvironment merged 11 commits into
mainfrom
284-research-calibrate-ch-and-bt-gradient-an
Sep 30, 2026
Merged

NewGraphEnvironment merged 11 commits into
mainfrom
284-research-calibrate-ch-and-bt-gradient-an

Conversation

@NewGraphEnvironment

@NewGraphEnvironment NewGraphEnvironment commented Sep 30, 2026 •

Copy link
Copy Markdown
Owner

Summary

  • Research: calibrate CH and BT gradient and channel-width thresholds from observations #284 step 5: BT rear_gradient_max 0.1349 in default_tuned is scored and held.
    • It was scored on eight watershed groups held out from its calibration.
    • The steps to 0.1349 add rearing that BT use at 0.60 and 0.71 of the core rate (0.75 and 0.89 at the same elevation).
    • The step to 0.1449 does not hold up: 8 found where 25 were expected.
    • default_tuned is unchanged. Results are in research/habitat_thresholds.md, "Step 5".
    • What that means for habitat: 0.1349 adds 1,105 km of BT rearing (+5.8 %) across the 12 watershed groups, from 2 % (BABL) to 9 % (KOTL, UARL). The per-group table is habitat_change.csv.
  • New export lnk_habitat_validate_band(). It measures observation density on the segments a threshold step moves, against the core every step agrees on. It joins on the full key and stops if segmentation differs or a schema holds no habitat for a WSG.
  • Harness, all inputs as data (data-raw/habitat_score/):
    • data-raw/habitat_variants_build.R models a drainage closure once, then re-classifies each threshold variant on that shared segmentation. Resume only reuses clean, same-commit bases, and built.csv records which bundle each schema came from.
    • data-raw/habitat_variants_score.R validates each variant and applies the fixed rule. It also writes a taper, an elevation split, an elevation-adjusted ratio, the bridge split and an expected-count floor beside the rule.
  • FISS absence taxa are now data (data-raw/fiss_absence_taxa.csv), and the absence and pooling loaders are shared (habitat_validate_inputs.R). The Validate modelled habitat against fish observations per species and watershed group #283 outputs are byte-identical, with knowledge pinned at 508bf44.

Related Issues

Test plan

  • devtools::test(): 0 fail, 2192 pass, the same 16 warnings as the baseline
  • devtools::check(): 0 errors. The 3 warnings and 2 notes are pre-existing (non-ASCII in older R files, missing Rd links, undocumented presence); none touch this diff.
  • lintr: no lints in the new files
  • Band guards mutation-tested: segmentation guard removed → red; bare id_segment join → red; habitat-presence guard removed → red
  • BULL pre-flight, then the full build (run 20260930_014921-53610765, 139.6 min), with post-conditions verified in the DB:
    • 23 base WSGs, all clean at a11a001;
    • 12 focal WSGs per variant;
    • default re-classify digests equal to the base run.
  • Bands reconcile with the rollup: default → 0.1349 adds 585 km, equal to 422.2 + 163.0 km of bands.
  • /code-check: 5 rounds plus an enumeration (planning/archive/2026-09-issue-284-step5-scoring/review-*.md)

Notes

  • The design changed before any band was seen. A counts-only power check showed the planned held-out set could decide one step in six. So:

    • the BT held-out set was widened;
    • CH was dropped, and its verdicts stay "keep", unscored;
    • an underpowered step changes nothing.

    Operator's calls, recorded in the research doc with the counts.

  • The n ≥ 10 floor is open. Counting locations found cannot refuse a band fish avoid; counting locations expected at the core rate can. Both readings are in verdict.csv. They agree on this run, and the choice stays open for the next tuning.

  • Elevation confounds gradient. Core BT rearing thins from 24.6 to 10.2 locations per 100 km between the low and high thirds of each WSG's rearing, and the steep bands sit high.

  • Clustering moves a little habitat the other way. Loosening the cutoff removed 1.1 km of rearing (against 1,365 km added). The nesting guard tolerates up to 1 % and reports the km.

  • The score284_* schemas are kept in local docker fwapg until this merges. Dropping them is a separate step.

🤖 Generated with Claude Code

https://claude.ai/code/session_01WJE2szSCkeV9R3qKhuSwyx

NewGraphEnvironment and others added 11 commits September 29, 2026 17:26
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01WJE2szSCkeV9R3qKhuSwyx
…ule (#284)

The FISS absence rules (which listed taxa mean caught / maybe) move from
regexes in habitat_validate.R to data-raw/fiss_absence_taxa.csv, and the
absence loader and pooling resolution move to habitat_validate_inputs.R so
the step-5 scorer uses the same code path. With knowledge pinned at
508bf44 the six #283 CSVs are byte-identical.

The scoring rule is fixed in research/habitat_thresholds.md before any
run. A counts-only power check (data-raw/logs/habitat_score_284/
power_windows.*) showed the first held-out set could decide one step of
six, so the design is BT only with a widened held-out set (operator's
call); CH verdicts stay keep, unscored.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01WJE2szSCkeV9R3qKhuSwyx
…a threshold step moves (#284)

Compares two persisted runs that share segmentation and differ in one
habitat threshold: the band (segments whose flag differs) against the core
(segments with the flag in every schema of a ladder), as observation
locations per km and their ratio. Joins on the full key, and stops when
segmentation differs or a schema holds no habitat rows for a WSG, since
either would turn into a false band. 37 tests; the guards are
mutation-tested.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01WJE2szSCkeV9R3qKhuSwyx
habitat_variants_build.R models the drainage closure once under default
(downstream-first), keeps the focal WSGs' working schemas, and
re-classifies each threshold variant on that shared network into
score284_<variant>. The invariants: identical segmentation, access copied
from the base, and a default re-classify that reproduces the base digest.
Each variant is a thin bundle with its own provenance checksum; built.csv
records per variant x WSG which bundle a schema came from, and resume
reuses only clean, same-commit bases.

habitat_variants_score.R validates every schema, computes the ladder
bands with lnk_habitat_validate_band(), and applies the fixed rule once
per ladder (an underpowered step changes nothing). Code-check ran five
rounds plus an enumeration; see planning/active/review-*.md.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01WJE2szSCkeV9R3qKhuSwyx
…beside the rule (#284)

taper.csv gives observations per km along each ladder (the core cut into
gradient bins, then each band); elevation.csv compares each band with the
core inside the same elevation tercile of the WSG's own rearing, so a band's
low rate can be separated from 'it is higher up' (temperature, size). The
verdict of record keeps the fixed rule; beside it, the same floor applied to
the locations a band would hold at the core's rate, which can refuse a band
fish avoid. That second reading is under discussion and moves nothing.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01WJE2szSCkeV9R3qKhuSwyx
…em (#284)

Loosening BT rear_gradient_max removed 1.1 km of rearing (BULK, KOTL,
LILL) while adding 1,365 km: a newly admitted segment merging clusters.
The nesting guard now stops only above 1 % of a step's length and the
against-direction km stay in bands.csv.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01WJE2szSCkeV9R3qKhuSwyx
Core BT rearing thins with elevation (24.6 / 19.3 / 10.2 locations per
100 km, low / mid / high terciles, held out) and the steep bands sit mostly
high, so the pooled ratio charges gradient for elevation. The adjusted
ratio compares found locations with those the core's rate predicts in the
band's own elevation mix.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01WJE2szSCkeV9R3qKhuSwyx
Run 20260930_014921-53610765 (23-WSG closure, 3 BT variants, 12 focal
WSGs). Held out, the steps to 0.1349 add rearing BT use at 0.60 and 0.71
of the core rate (0.75 and 0.89 at the same elevation); the step to 0.1449
does not (8 found, 25 expected). default_tuned is unchanged and no longer
described as unscored. Evidence under data-raw/logs/habitat_score_284/;
results in research/habitat_thresholds.md, Step 5. Weights instead of
cutoffs went to NewGraphEnvironment/knowledge#28.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01WJE2szSCkeV9R3qKhuSwyx
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01WJE2szSCkeV9R3qKhuSwyx
habitat_change.csv gives, per variant x WSG, stream-rearing and spawning km
against the base and the change in km and %, so 'how much habitat does the
cutoff add' is a produced number rather than a hand query of summary.csv.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01WJE2szSCkeV9R3qKhuSwyx
research/habitat_thresholds.md gains a plain-terms summary and the
per-WSG table from habitat_change.csv: 0.1349 adds 1,105 km of BT
rearing (+5.8 %) across the 12 WSGs, 2 % (BABL) to 9 % (KOTL, UARL).
It also gains the combined elevation share: 366 of the 585 km the steps
add sit in the high third. The re-run's outputs differ from the last
only in trailing digits (Postgres float summation order, as #293).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01WJE2szSCkeV9R3qKhuSwyx
@NewGraphEnvironment
NewGraphEnvironment merged commit b4dcc11 into main Sep 30, 2026
1 check passed
@NewGraphEnvironment
NewGraphEnvironment deleted the 284-research-calibrate-ch-and-bt-gradient-an branch September 30, 2026 17:48
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Research: calibrate CH and BT gradient and channel-width thresholds from observations

1 participant