Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
22 changes: 21 additions & 1 deletion CLAUDE.md
Original file line number Diff line number Diff line change
Expand Up @@ -11,6 +11,25 @@ Experimental package — breaking all the time and loving the learning curve. St
**Prefix:** `lnk_`
**Branch:** `main` (current version: `DESCRIPTION` / [`NEWS.md`](NEWS.md))

## Status (2026-10-02) — the validator scores each group on its own model (#299)

**`lnk_habitat_validate()` re-tests a `mad` group on discharge, not width.**
- **One rule for the model:** fresh's `.frs_habitat_models()`, through `.lnk_wsg_model()`,
for classify and validator both. The method table is the one in the `cfg` you pass, so
swap it on that object; `lnk_pipeline_classify(method_csv =)` is invisible to it.
- **Labels are shared** (`fails_width` means the group's size); `model` / `mad_m3s` split
them. **`no_mad_threshold`** is a `mad` miss a MAD range alone would admit.
- Evidence: `data-raw/logs/habitat_validate_299/`; archive
`planning/archive/2026-10-issue-299-validate-mad/`.

**Facts not worth re-deriving:**
- **"Persisted TRUE, predicate FALSE" cannot catch a wrong model.** A cw predicate on a
`mad` run is looser, so the error shows as `post_predicate` (27 BT rearing misses on
ADMS), never as a disagreement.
- **`no_mad_threshold` is a floor, not a count, and cannot score a fix**: same-cause
misses also read `width_null` / `fails_gradient_and_width`, and once a range exists the
label cannot fire. Score MAD thresholds with capture, cost and the band score.

## Status (2026-10-01) — per-WSG `cw`/`mad` habitat model threaded (#286)

**Each bundle's `parameters_habitat_method.csv` reaches fresh as `params_method`.**
Expand All @@ -32,7 +51,8 @@ Experimental package — breaking all the time and loving the learning curve. St
- **fresh 0.36.0 reversed `frs_db_conn()`'s precedence** (`PG*` first). `lnk_db_conn()`
still reads `PG_*_SHARE` first; on a machine with both groups set they connect to
different databases.
- **`lnk_habitat_validate()` is cw-only** (#299).
- **`lnk_habitat_validate()` scores each group on its own model** (#299), from the
`cfg` it is given; `no_mad_threshold` names a species with no MAD range.

## Status (2026-09-29) — #284 step 5: BT `rear_gradient_max` 0.1349 scored and held

Expand Down
17 changes: 17 additions & 0 deletions R/lnk_config.R
Original file line number Diff line number Diff line change
Expand Up @@ -356,6 +356,23 @@ print.lnk_config <- function(x, ...) {
fresh_path
}

# A habitat method table read as classify hands it to fresh: every column
# character, and no `na.strings`, so a group coded "NA" stays a code rather
# than becoming a missing value that fresh would then resolve to cw.
.lnk_habitat_method_read <- function(path) {
utils::read.csv(path, stringsAsFactors = FALSE, colClasses = "character",
na.strings = character(0))
}

# The model (`cw` or `mad`) each watershed group classifies on, resolved by
# fresh's own rule so classify and the validator cannot disagree with it:
# an unlisted group is cw, and a model other than cw/mad or a duplicated
# group code is an error. Unnamed, one per `wsg`.
.lnk_wsg_model <- function(params_method, wsg) {
resolve <- utils::getFromNamespace(".frs_habitat_models", "fresh")
unname(resolve(wsg, params_method))
}

# Absolute path of a provenance entry. Keys are relative to the bundle that
# declared them: the leaf for its own entries, `.dir` for inherited ones.
.lnk_provenance_path <- function(cfg, rel) {
Expand Down
288 changes: 224 additions & 64 deletions R/lnk_habitat_validate.R

Large diffs are not rendered by default.

6 changes: 2 additions & 4 deletions R/lnk_pipeline_classify.R
Original file line number Diff line number Diff line change
Expand Up @@ -91,9 +91,7 @@ lnk_pipeline_classify <- function(conn, aoi, cfg, loaded, schema,
if (!nzchar(method_csv) || !file.exists(method_csv)) {
stop("method_csv not found: ", method_csv, call. = FALSE)
}
params_method <- utils::read.csv(method_csv, stringsAsFactors = FALSE,
colClasses = "character",
na.strings = character(0))
params_method <- .lnk_habitat_method_read(method_csv)

species <- species %||% lnk_pipeline_species(cfg, loaded, aoi)
if (length(species) == 0L) {
Expand Down Expand Up @@ -164,7 +162,7 @@ lnk_pipeline_classify <- function(conn, aoi, cfg, loaded, schema,
#
# Channel-width model only (#286): the bypass stands in for a width test,
# and bcfp applies it inside its cw branch alone. A `mad` group skips it.
aoi_model <- params_method$model[match(aoi, params_method$watershed_group_code)]
aoi_model <- .lnk_wsg_model(params_method, aoi)
species_bypass <- if (identical(aoi_model, "mad")) character(0) else species
for (sp in species_bypass) {
rear_rules <- params[[sp]][["rules"]][["rear"]]
Expand Down
14 changes: 9 additions & 5 deletions R/lnk_preflight_fresh.R
Original file line number Diff line number Diff line change
Expand Up @@ -120,9 +120,12 @@ lnk_preflight_fresh <- function(required = .lnk_fresh_required(),

# Arguments link passes that older fresh releases do not accept. The call
# would fail with "unused argument" only once a WSG reached that phase.
# R/lnk_pipeline_classify.R: params_method arrived in fresh 0.35.0 (#286).
# R/lnk_pipeline_classify.R: params_method arrived in fresh 0.35.0 (#286);
# R/lnk_habitat_validate.R: frs_habitat_predicates(model) in the same
# release (#299).
.lnk_fresh_required_formals <- function() {
list(frs_habitat_classify = "params_method")
list(frs_habitat_classify = "params_method",
frs_habitat_predicates = "model")
}

# "fn(arg)" for each required formal the namespace does not provide. A
Expand All @@ -149,10 +152,11 @@ lnk_preflight_fresh <- function(required = .lnk_fresh_required(),
out
}

# Non-exported fresh objects link reaches via getFromNamespace().
# R/lnk_pipeline_connect.R:101.
# Non-exported fresh objects link reaches via getFromNamespace():
# .frs_run_connectivity in lnk_pipeline_connect(), .frs_habitat_models in
# .lnk_wsg_model() (R/lnk_config.R, #299).
.lnk_fresh_required_internal <- function() {
".frs_run_connectivity"
c(".frs_run_connectivity", ".frs_habitat_models")
}

# Every `fresh::sym` / `fresh:::sym` reached from link's own namespace,
Expand Down
10 changes: 8 additions & 2 deletions RUNBOOK.md
Original file line number Diff line number Diff line change
Expand Up @@ -714,8 +714,14 @@ next dispatch.**
test, so a `mad` group without coverage loses all of its stream habitat with no
error. BULK has none in the local fwapg. Check `count(mad_m3s)` before moving a
group.
- **`lnk_habitat_validate()` is still cw-only**: its miss-reason relaxation
rewrites `s.channel_width`, so it scores a `mad` group as if it were `cw`.
- **`lnk_habitat_validate()` scores each group on its own model** (#299), from the
`cfg` it is handed, resolved by fresh's `.frs_habitat_models()` (one rule for classify
and validator). A `mad` group's predicates and size relaxation read `mad_m3s`, joined
from the discharge table on `linear_feature_id`; reason labels are shared with `cw`
and the `model` / `mad_m3s` columns split them, plus `no_mad_threshold` for a species
with no MAD range. To score a run made with a swapped method table, swap it on the
same `cfg` object (`cfg$files$parameters_habitat_method$path`) before validating:
`lnk_pipeline_classify(method_csv =)` is invisible to the validator.
- **Connectivity reads no size model.** `.frs_run_connectivity` takes no
`params_method`; its one width test (`.frs_connected_waterbody`,
`spawn_connected_cw_min`) is 0 for SK/KO in every bundle, so it is inert.
Expand Down
17 changes: 12 additions & 5 deletions data-raw/habitat_validate.R
Original file line number Diff line number Diff line change
Expand Up @@ -200,7 +200,8 @@ if (nrow(bundles) == 2L) {
# totals' n_habitat; its reason is then the rear reason.
stage_rows <- function(obs) {
base <- data.frame(
species_code = obs$species_code, gradient = obs$gradient,
species_code = obs$species_code, model = obs$model,
gradient = obs$gradient,
channel_width = obs$channel_width, edge_type = obs$edge_type,
is_spawn = obs$is_spawn, is_rear = obs$is_rear,
reason_any = ifelse(obs$spawning %in% TRUE, NA_character_,
Expand All @@ -209,7 +210,7 @@ stage_rows <- function(obs) {
reason_rear = obs$miss_reason_rear, stringsAsFactors = FALSE)
pick <- function(d, stage, col) {
cbind(data.frame(stage = rep(stage, nrow(d))), d[, c("species_code",
"gradient", "channel_width", "edge_type")], reason = d[[col]])
"model", "gradient", "channel_width", "edge_type")], reason = d[[col]])
}
rbind(pick(base, "any", "reason_any"),
pick(base[base$is_spawn, , drop = FALSE], "spawn", "reason_spawn"),
Expand All @@ -223,18 +224,24 @@ binned <- list()
for (r in runs[vapply(runs, function(r) r$summary$buffer_m[1] == 0, TRUE)]) {
d <- stage_rows(r$observations)
d$reason[is.na(d$reason)] <- "captured"
# model: a mad group's width reasons are about discharge (#299).
tab <- as.data.frame(table(stage = d$stage, species_code = d$species_code,
reason = d$reason), stringsAsFactors = FALSE)
model = d$model, reason = d$reason),
stringsAsFactors = FALSE)
tab <- tab[tab$Freq > 0, ]
names(tab)[names(tab) == "Freq"] <- "n"
misses[[length(misses) + 1L]] <- cbind(
data.frame(bundle = rep(r$bundle, nrow(tab))), tab)
m <- d[!d$reason %in% c("captured", "no_segment"), ]
m$gradient_bin <- as.character(cut(m$gradient, breaks_g, right = TRUE))
m$width_bin <- ifelse(is.na(m$channel_width), "NULL",
# Width bins describe the cw size dimension only; a mad row keeps its
# count under "mad" so the binned table still sums to misses.csv.
m$width_bin <- ifelse(m$model %in% "mad", "mad",
ifelse(is.na(m$channel_width), "NULL",
as.character(cut(m$channel_width, breaks_w,
right = FALSE)))
right = FALSE))))
b <- as.data.frame(table(stage = m$stage, species_code = m$species_code,
model = m$model,
reason = m$reason, gradient_bin = m$gradient_bin,
width_bin = m$width_bin, useNA = "ifany"),
stringsAsFactors = FALSE)
Expand Down
76 changes: 76 additions & 0 deletions data-raw/logs/habitat_validate_299/20261002_adms_compare.txt
Original file line number Diff line number Diff line change
@@ -0,0 +1,76 @@
## 1. cw no-change (fresh_default, ADMS)
observations identical apart from new columns: TRUE
summary identical apart from `model`: TRUE
branch model column: cw
observations: 94

## 2. persisted TRUE, re-built predicate FALSE (zz299_mad, ADMS on mad)
branch:
species_code n spawn_true_pred_false rear_true_pred_false
BT 37 0 0
CH 11 0 0
CO 46 0 0
main (cw predicates on a mad run):
species_code n spawn_true_pred_false rear_true_pred_false
BT 37 0 0
CH 11 0 0
CO 46 0 0

## 3. miss reasons on the mad run (locations)
spawn_all
species reason Freq_branch Freq_main
BT fails_gradient 0 6
BT fails_gradient_and_width 8 2
BT fails_width 0 2
BT no_mad_threshold 27 0
BT post_predicate 0 24
BT rule_excludes 2 2
BT width_null 0 1
CH captured 8 8
CH fails_gradient 1 1
CH fails_width 1 0
CH post_predicate 0 1
CH rule_excludes 1 1
CO captured 41 41
CO fails_gradient 1 1
CO fails_width 2 0
CO not_accessible 1 1
CO rule_excludes 1 1
CO width_null 0 2
spawn_staged
species reason Freq_branch Freq_main
BT no_mad_threshold 1 0
BT post_predicate 0 1
CH captured 5 5
CO captured 22 22
rear_all
species reason Freq_branch Freq_main
BT captured 2 2
BT fails_gradient 0 3
BT fails_gradient_and_width 4 1
BT fails_width 0 3
BT no_mad_threshold 31 0
BT post_predicate 0 27
BT width_null 0 1
CH captured 9 9
CH fails_gradient 1 1
CH post_predicate 1 1
CO captured 38 38
CO fails_gradient 3 3
CO fails_width 4 0
CO not_accessible 1 1
CO post_predicate 0 2
CO width_null 0 2
rear_staged
species reason Freq_branch Freq_main
BT fails_width 0 2
BT no_mad_threshold 6 0
BT post_predicate 0 4
CH captured 1 1
CH post_predicate 1 1
CO captured 12 12
CO fails_width 1 0
CO width_null 0 1

mad_m3s present on 94 of 94 mad-group locations
validate seconds: branch cw 12 main cw 15 branch mad 11 main mad 15
7 changes: 7 additions & 0 deletions data-raw/logs/habitat_validate_299/20261002_adms_mad_km.txt
Original file line number Diff line number Diff line change
@@ -0,0 +1,7 @@
species_code | spawning_km | rearing_km | lake_wetland_rearing_km
--------------+-------------+------------+-------------------------
BT | 0.0 | 0.0 | 318.2
CH | 256.4 | 317.5 | 98.7
CO | 292.4 | 317.8 | 113.9
(3 rows)

24 changes: 24 additions & 0 deletions data-raw/logs/habitat_validate_299/20261002_adms_model_mad.txt
Original file line number Diff line number Diff line change
@@ -0,0 +1,24 @@
Warning messages:
1: In st_read.DBIObject(conn, query = query) :
Could not find a simple features geometry column. Will return a `data.frame`.
2: In st_read.DBIObject(conn, query = query) :
Could not find a simple features geometry column. Will return a `data.frame`.
3: In st_read.DBIObject(conn, query = query) :
Could not find a simple features geometry column. Will return a `data.frame`.
4: In st_read.DBIObject(conn, query = query) :
Could not find a simple features geometry column. Will return a `data.frame`.
5: In st_read.DBIObject(conn, query = query) :
Could not find a simple features geometry column. Will return a `data.frame`.
6: In st_read.DBIObject(conn, query = query) :
Could not find a simple features geometry column. Will return a `data.frame`.
7: In st_read.DBIObject(conn, query = query) :
Could not find a simple features geometry column. Will return a `data.frame`.
8: In st_read.DBIObject(conn, query = query) :
Could not find a simple features geometry column. Will return a `data.frame`.
9: In st_read.DBIObject(conn, query = query) :
Could not find a simple features geometry column. Will return a `data.frame`.
10: In st_read.DBIObject(conn, query = query) :
Could not find a simple features geometry column. Will return a `data.frame`.
## model_mad ADMS -> zz299_mad method=method_adms_mad.csv link=0.55.0 fresh=0.36.2 2.8 min
n_segments
1 39422
Original file line number Diff line number Diff line change
@@ -0,0 +1 @@
## validate fresh_default CH,CO,BT method=- sha=be3be30 fresh=0.36.2 12 s, 94 locations
Original file line number Diff line number Diff line change
@@ -0,0 +1 @@
## validate zz299_mad CH,CO,BT method=method_adms_mad.csv sha=be3be30 fresh=0.36.2 11 s, 94 locations
Original file line number Diff line number Diff line change
@@ -0,0 +1 @@
## validate fresh_default CH,CO,BT method=- sha=4df1ffc fresh=0.36.2 15 s, 94 locations
Original file line number Diff line number Diff line change
@@ -0,0 +1 @@
## validate zz299_mad CH,CO,BT method=- sha=4df1ffc fresh=0.36.2 15 s, 94 locations
52 changes: 52 additions & 0 deletions data-raw/logs/habitat_validate_299/README.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,52 @@
# #299 live verification — validator scores each group on its own model

Local docker fwapg (:5432), 2026-10-02 UTC. Config `default`, ADMS, species CH, CO, BT
(DV pooled per `species_pooling.csv`), buffer 0. fresh 0.36.2. Branch runs are the
working tree on top of `be3be30`, before the Phase 2–3 commit (the `sha=` in their
headers is HEAD, not the tree that ran). Main runs are a detached worktree at `4df1ffc`.

- `model_mad.R` models ADMS with the `default` bundle into the scratch persist schema
`zz299_mad`, with its method table swapped for `method_adms_mad.csv` (ADMS on `mad`),
`mapping_code = TRUE`. 39,422 segments, 2.8 min. `km.sql` →
`20261002_adms_mad_km.txt`: stream spawning / rearing km BT 0 / 0, CH 256.4 / 317.5,
CO 292.4 / 317.8. BT keeps 318.2 km of lake and wetland rearing.
- `validate_run.R` runs `lnk_habitat_validate()` from a given repo, swapping the method
table on the same `cfg` object it validates with.
- `compare.R` compares the four runs: `20261002_adms_compare.txt`.

## Result

| check | result |
|---|---|
| cw no-change: `fresh_default` ADMS, branch vs main | `observations` identical apart from the new columns; `summary` identical apart from `model` (94 locations) |
| `mad_m3s` on mad-group locations | 94 of 94 |
| persisted TRUE, re-built predicate FALSE (outside UHC) | 0 on branch, **and 0 on main**: this metric does not discriminate (see below) |

**The planned metric was the wrong direction.** On a `mad` run, main re-builds the *cw*
predicate. For BT it is looser than the `mad` one that classified the run, so it passes
where classify failed. That is persisted FALSE with predicate TRUE, which the validator
reports as `post_predicate` ("removed by clustering or access gating"), not as a
disagreement. The discriminating numbers are the reasons, in locations
(`20261002_adms_compare.txt` §3). "All" tests every location against that stage's
predicate; "staged" keeps only locations whose records carry the stage, as the summary's
stage rows and the driver's `misses.csv` count them:

| species | reason column | main (cw predicates) | branch (mad predicates) |
|---|---|---|---|
| BT, 37 locations | spawn, all | 24 `post_predicate`, 6 `fails_gradient`, 2 `fails_width`, 1 `width_null`, 2 `fails_gradient_and_width`, 2 `rule_excludes` | 27 `no_mad_threshold`, 8 `fails_gradient_and_width`, 2 `rule_excludes` |
| BT | spawn, staged (1) | 1 `post_predicate` | 1 `no_mad_threshold` |
| BT | rear, all | 27 `post_predicate`, 3 `fails_gradient`, 3 `fails_width`, 1 `width_null`, 1 `fails_gradient_and_width`, 2 captured | 31 `no_mad_threshold`, 4 `fails_gradient_and_width`, 2 captured |
| BT | rear, staged (6) | 4 `post_predicate`, 2 `fails_width` | 6 `no_mad_threshold` |
| CO, 46 locations | spawn, all | 2 `width_null` | 2 `fails_width` |
| CO | rear, all | 2 `width_null`, 2 `post_predicate` | 4 `fails_width` |
| CO | rear, staged (13) | 1 `width_null` | 1 `fails_width` |
| CH, 11 locations | spawn, all | 1 `post_predicate` | 1 `fails_width` |

On the 27 BT locations main calls `post_predicate` against rearing, it tells a reader
the predicate passed and clustering or gating removed them. On a `mad` group they were
never in the predicate, because BT has no MAD range. Main's CO `width_null` are NULL
channel widths the `mad` run never tested; the branch tests their discharge (0.0127 and
0.018 m³/s on the two spawn misses), which is below CO's minimum.

Scratch schemas `zz299_mad` and `zz299_w_adms` were dropped after the runs; the RDS
files were not kept.
Loading
Loading