Skip to content

Region-scoped DV→BT observation pooling, tracked in bundle CSVs #290

Description

@NewGraphEnvironment

If we do it: whether a DV record counts as BT evidence is decided per region, sub-region or WSG, in bundle CSVs that record why, when and from what source. If we never do: DV is pooled with BT wherever BT is present. That rule is hard-coded in three places, and it is least defensible in exactly the places where DV and BT differ.

Split out of #236 (question 1, narrowed to observation evidence). #284 step 5 is parked until this lands, because its BT evidence is defined by this rule.

Problem

DV→BT pooling is a biological position, but it lives in code, with one global scope:

  • R/lnk_habitat_validate.R:202: species_obs = list(BT = c("BT", "DV")), the default.
  • data-raw/query_habitat_thresholds_obs.R:79-87: CASE WHEN species_code = 'DV' THEN 'BT', gated only on BT presence in the WSG.
  • inst/extdata/configs/default/parameters_fresh.csv: observation_species = BT;DV for the BT access override. This one is out of scope here and stays with Research: DV as BT access evidence — regional scope + multi-rule observation overrides #236, which needs multi-rule overrides. It should consume the same table when that lands.

The rule matters for #284. BT rear_gradient_max 0.1249 vs 0.1349 in default_tuned is exactly BT-only vs BT+DV pooled, and in BULK and MORR the "BT" evidence is ~90 % DV records (442 DV / 40 BT, 344 DV / 49 BT, raw A/B locations).

The operator's position (2026-09-26): pool in named regions only, and build up the state of knowledge over time in CSVs.

Proposed Solution

Revised at plan approval (2026-09-26). The region lookup is geography, so it ships once at package level. Only the biological files live in the bundle, declared under files:. And pooling is abstract: either side of a rule can be a species or a group, so the same tool pools DV into BT or all salmon together.

1. inst/extdata/wsg_regions.csv (package-level), one row per WSG (246): watershed_group_code, region, subregion, wscode_outlet. Built by data-raw/wsg_regions.R from data-raw/wsg_regions_defs.csv.

region is generated from the first segment of each group's outlet wscode_ltree (fresh@v0.33.0 inst/extdata/wsg_outlet.csv). Reading it at the outlet matters because some groups straddle codes: LFRA touches both 100 and 900. Measured on local fwapg:

code region e.g.
100 Fraser LFRA, UNTH, LNTH
200 Mackenzie PARS, FINA, LIAR
300 Columbia ELKR, KOTL, BULL
400 Skeena BULK, MORR

The 500–800 and 9xx codes are northern rivers and the coast; names get confirmed when the generator runs. subregion is hand-curated, starting with Kootenay under Columbia. The generator is a data-raw/ script.

2. configs/default/overrides/species_groups.csv: group_code, species_code, seeded with SALMON (CH, CM, CO, PK, SK) and CHAR (BT, DV).

3. configs/default/overrides/species_pooling.csv, the tracker. species_code (the target) and species_obs (the evidence) each take a species or a group code. Columns: species_code, species_obs, scope_level, scope, pool, confidence, rationale, source, verified, issue.

  • scope_level is region, subregion or wsg.
  • Resolution: the most specific scope wins; at equal scope a species row beats a group row; a tie left with a different pool errors; anything unlisted is not pooled. A WSG row overrides its sub-region, which overrides its region; a WSG with no applicable row keeps DV separate.
  • Knowledge accrues as dated, sourced rows. Git history of the file is the record.
  • Seed rows: BT, DV, region, {Fraser, Mackenzie, Skeena, Columbia}, pool = yes and BT, DV, subregion, Kootenay, pool = yes. Source: operator call 2026-09-26 (inland DV are bull trout recorded under the other name). The coast and the north are unlisted. Revised 2026-09-27: the Skeena region row is now pool = no, and a sub-region row Skeena above Hazelton is pool = yes (see Status).

default declares the two bundle files (default_tuned inherits them). bcfishpass declares neither.

"No tracker" has one meaning per layer (settled in review):

  • lnk_species_pooling() without a tracker pools nothing: a species counts only as itself.
  • lnk_habitat_validate() keeps its list default (list(BT = c("BT","DV")), global) when called without a table, for back-compat.
  • data-raw/habitat_validate.R resolves pooling once, from the first of --bundles that declares a tracker (or --pooling=), and applies it to every bundle it scores, so a default vs bcfishpass diff is on the same observations. Only when no bundle declares one does the list default apply. The stamp records which.

Status (2026-09-27)

Built in PR #291, including three things added after the naming history was measured
(research/species_pooling.md):

  • The Skeena is split at Hazelton (operator call).
    • A new sub-region, Skeena above Hazelton, pools DV → BT in BULK, MORR, KISP, BABL, BABR, SUST, MSKE and USKE.
    • The Skeena region row is no, because Skeena DV is a current name, not a legacy one.
    • wsg_regions_defs.csv gains explicit-membership sub-regions for this.
  • The tracker gains an optional obs_year_max, which pools only records dated in or before a year. It is empty in the seed.
  • Research: calibrate CH and BT gradient and channel-width thresholds from observations #284's BT evidence moves. 852 DV records below Hazelton drop out. The rule still gives 0.1349, now by 0.0009.

Tasks

  • Generator for inst/extdata/wsg_regions.csv: region from the outlet code, sub-region from a small curated list; commit the output
  • Seed species_groups.csv and species_pooling.csv; add both to default's files: and to the config dictionaries
  • One resolver, lnk_species_pooling(loaded, aoi, species, regions = NULL) → a per-WSG table (watershed_group_code, species_code, obs_species, scope_level, scope, rule). Species-agnostic: it knows nothing about BT or DV
  • lnk_habitat_validate() and data-raw/habitat_validate.R take species_obs from the resolver when the bundle declares the tracker
  • data-raw/query_habitat_thresholds_obs.R pools through the resolver instead of the hard-coded CASE
  • Re-run Research: calibrate CH and BT gradient and channel-width thresholds from observations #284 step 1 under the regional rule; record how far the BT P95 and the verdicts move in research/habitat_thresholds.md, and in default_tuned if a value changes
  • Tests (synthetic codes, so nothing passes on BT/DV by accident): resolution precedence (WSG > sub-region > region > unlisted), group expansion on either side, unlisted means not pooled, and no tracker means nothing pooled

Out of scope: the access override's observation_species (#236), and DV as a modelled species with its own thresholds (#236, fresh).

Relates to #236, #284, #283, #189

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions