Region-scoped observation pooling as bundle data: lnk_species_pooling() (#290) - #291
Merged
NewGraphEnvironment merged 11 commits intoSep 28, 2026
Merged
Conversation
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01WJE2szSCkeV9R3qKhuSwyx
inst/extdata/wsg_regions.csv assigns each of the 246 groups the region its outlet drains to (first wscode segment) and, where curated, a sub-region (longest whole-segment prefix). Kootenay (300.625474) is the first. Coastal and cross-border region names are drafts, marked in the defs. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01WJE2szSCkeV9R3qKhuSwyx
Which observation species count as evidence for which model species is now a tracker in the bundle (species_pooling.csv), scoped to region, sub-region or WSG through inst/extdata/wsg_regions.csv, with named groups (species_groups.csv) usable on either side. The resolver knows no species: DV into BT, or all salmon together, is a row. Most specific scope wins, then the row naming fewer groups; unresolved ties error; anything unlisted is not pooled. default declares both (seeded: DV -> BT in Fraser, Mackenzie, Skeena, Columbia, Kootenay); default_tuned inherits them. config_hash now also covers wsg_regions.csv for bundles with a tracker. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01WJE2szSCkeV9R3qKhuSwyx
#290) lnk_habitat_validate() accepts lnk_species_pooling()'s per-WSG table as species_obs (checked as a data frame first; the list form and its default are unchanged). habitat_validate.R resolves pooling once, from the first bundle that declares a tracker (--pooling=), applies it to both bundles and stamps it. query_habitat_thresholds_obs.R joins the resolved table in place of the hard-coded DV CASE. Both refuse a pooling bundle whose presence differs, WSG by WSG, from the scored one's. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01WJE2szSCkeV9R3qKhuSwyx
research/species_pooling.md holds the tracker's state per region. The #284 evidence is refreshed: only the province-wide ledger step 3 moves (DV 8,888 -> 6,138); every other CSV re-runs byte-identical, as does the #283 validation baseline. Method lines in the research docs now point at the tracker. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01WJE2szSCkeV9R3qKhuSwyx
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01WJE2szSCkeV9R3qKhuSwyx
Interior DV is a legacy name (the DV share of char records falls from ~90% before 1990 to under 15% after 2000). Skeena DV is not: 86% or more in every decade, and 85% of Skeena DV streams were re-sampled after 1995 and still recorded only DV. default_tuned's BT rear_gradient_max of 0.1349 depends on the Skeena DV records: without them the rule gives 0.1249, the BT-only value. The above-Hazelton cut keeps 0.1349, by 0.0009. data-raw/species_pooling_evidence.R runs four scenarios (S3, pre-1995 interior pooling, computed from the S0 evidence and checked against the producer on S0-S2) and scores S1/S2 through the validator. Figures and maps are in data-raw/logs/species_pooling_290/; write-up in research/species_pooling.md. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01WJE2szSCkeV9R3qKhuSwyx
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01WJE2szSCkeV9R3qKhuSwyx
…290) Skeena DV is a current name, not a legacy one, so the Skeena region row now keeps DV apart and a new sub-region, 'Skeena above Hazelton' (BULK, MORR, KISP, BABL, BABR, SUST, MSKE, USKE), pools it. wsg_regions_defs.csv gains explicit-membership sub-regions, since MSKE and USKE outlets sit on the Skeena mainstem like LSKE and no wscode prefix can split them. species_pooling.csv gains an optional obs_year_max: a row then pools only records dated in or before that year (undated records do not qualify). lnk_species_pooling() carries the winning row's limit; the validator and the #284 obs producer apply it per record, and the validator refuses a year limit on a source with no observation_date. Empty in the seed. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01WJE2szSCkeV9R3qKhuSwyx
The reimplementation applied a year limit to already-deduplicated S0 evidence, while the producer applies it per record before deduplicating, so a location with DV records on both sides of 1995 was judged by whichever record survived (review round, 2026-09-27). Every scenario now reports its own producer run; the reimplementation cross-checks the scenarios without a year limit and supplies BT_only. The S3 figures published earlier came from the reimplementation and are superseded. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01WJE2szSCkeV9R3qKhuSwyx
…n split (#290) Under the adopted tracker 852 Skeena DV records below Hazelton drop out of the BT evidence. #284: pooled BT rearing gradient P95 0.1348 -> 0.1309, n 4,764 -> 3,988; the rule still gives 0.1349, now by 0.0009, and no verdict moves. #283: the shared BT set is 3,771 locations (was 4,600); CH is byte-identical. The scenario table now reports producer runs for all five scenarios (S3 corrected: 543 DV records, P95 0.1202) and adds S4, the split with the interior limited to pre-1995 records (P95 0.1310). Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01WJE2szSCkeV9R3qKhuSwyx
7 tasks done
NewGraphEnvironment
deleted the
290-region-scoped-dv-bt-observation-pooling
branch
September 28, 2026 00:48
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
configs/default/species_pooling.csv: one dated, sourced row per decision, scoped to a region, sub-region or WSG. Either side can be a group fromspecies_groups.csv.lnk_species_pooling()knows no species: DV into BT, or all salmon together, is a row.inst/extdata/wsg_regions.csv, which uses each group's outlet wscode and is built bydata-raw/wsg_regions.R. It is also hashed intoconfig_hashfor any bundle that declares a tracker.lnk_habitat_validate()accepts the per-WSG table asspecies_obs; the list form and its default are unchanged.data-raw/habitat_validate.Rresolves pooling once and applies it to both bundles. The Research: calibrate CH and BT gradient and channel-width thresholds from observations #284 obs producer joins the table in place of its hard-coded DVCASE. Both drivers stop if the pooling bundle's presence table disagrees, WSG by WSG, with the scored bundle's.Skeena above Hazelton, defined by an explicit group list. MSKE and USKE have their outlets on the Skeena mainstem, like LSKE, so no code prefix can separate them.obs_year_maxcolumn pools only records dated in or before a year. The validator and the Research: calibrate CH and BT gradient and channel-width thresholds from observations #284 script apply it per record. It is empty in the seed.data-raw/species_pooling_evidence.R, with maps and figures. The write-up isresearch/species_pooling.md, and the walkthrough page is When a Dolly Is a Bull Trout (private).Related Issues
Test plan
devtools::test()(at4e9893d): 2151 pass, 3 fail. The 3 are all the:63333bcfishpass tunnel being down (test-lnk_db_conn.R:10,test-lnk_wsg_resolve.R:143,:154), in files this branch does not touch.test-lnk_species_pooling.R, on synthetic codes (AAA/BBB/GRP), so nothing passes on BT/DV by accident;species_obsform, end to end against the list form;config_hash;consumed_byref lands on a line naming its column.pool = no, the presence gate, the df branch, the hash block, and wrong dictionary refs each turn their tests red.911bed3).defaultvsbcfishpass, 55 + 59 WSGs, knowledge @ 508bf44) re-ran byte-identical under the region rule. Under the split, its shared BT set falls from 4,600 to 3,771 locations, CH is unchanged, and the committed CSVs are updated.pool = no: exactly 1,929 DV rows drop, in the 9 Skeena WSGs; nothing else moves.devtools::document()andpkgdown::check_pkgdown()clean.lintr: not clean.test-lnk_log.Ron main).object_usage_linterfalse positives for helpers that are absent from the stale installed package.Notes
config_hashchanges fordefaultanddefault_tuned, because they declare two new files and now hash the region lookup. No model output changed, so a hash mismatch against pre-Region-scoped DV→BT observation pooling, tracked in bundle CSVs #290 log rows is expected, not drift./code-checkrounds. That is 7 agents, past the usual bound; each round found real defects.planning/archive/2026-09-issue-290-species-pooling/review-*.md.lnk_presence()'s hard-coded groups.candidates.csv, from a double-precisionsum().🤖 Generated with Claude Code
https://claude.ai/code/session_01WJE2szSCkeV9R3qKhuSwyx