Skip to content

Standalone wsg_run_one.R writes NULL run_uid and run_label; only the dispatcher path labels rows #281

Description

@NewGraphEnvironment

Problem

study_area_run.sh mints a run_uid and exports it with LNK_RUN_LABEL to
every host, so every campaign row in fresh.log is queryable as a unit and
carries operator text pointing at the work that asked for it. The standalone
path — wsg_run_one.R / wsg_recompute_one.R invoked directly — sets neither,
and nothing says so.

Both fields simply land NULL. The run succeeds, the row looks complete, and the
absence is only visible if you go looking for it.

What changes if we do it: a single-WSG top-up is as traceable as a campaign
WSG, and WHERE run_label LIKE '%<issue>%' reaches every row rather than only
the ones that happened to go through the dispatcher.

What happens if we never do: run_uid and run_label mean "this row came
from study_area_run.sh" rather than what they were built to mean (link#262),
and the fraction of NULLs grows every time someone models one WSG — which is the
cheap, correct path whenever the downstream closure is already current, so it
should get more common, not less.

Measured 2026-09-25

MCGR modelled standalone for floodplains#90 — the legitimate path, since its 12
downstream WSGs were current and the DS-first guard passed in error mode:

LNK_LOAD=loadall Rscript data-raw/wsg_run_one.R MCGR bcfishpass

Resulting fresh.log row: run_uid NULL, run_label NULL. Everything else
present and correct — link 0.50.0, fresh 0.33.0, fresh_sha 7f12d991,
link_dirty = f, config_name bcfishpass, host m1.

Compare with the floodplains#75 rows from study_area_run.sh, all six of which
carry 20260903_173205-67013fc3 and
floodplains-75_north_thompson_UNTH_LNTH_THOM.

Same shape as link#278, inverted

link#278 is a surprising default — a value supplied when none was asked for.
This is a surprising absence — a field the campaign path always populates,
silently empty on the other path to the same table. Both are invisible at the
call site, and in both cases the two paths disagree with nothing to flag it.

Not fixable after the fact

An UPDATE fresh.log SET run_label = ... records a label the run did not carry.
That is provenance laundering, and this table exists precisely to distinguish
what a run did from what someone later said about it. The remedy has to be at
write time.

Sketch, not a decision

  • Default run_label in wsg_run_one.R to something derived and honest — the
    invocation, or standalone-<wsg>-<date> — rather than NULL.
  • Or mint a run_uid per standalone invocation the same way the dispatcher does,
    so one-WSG runs are still a queryable unit of one.
  • Or make LNK_RUN_LABEL required on the standalone path, which is link#278's
    remedy applied to the same file and would compose with that work.

The third is probably right and should be decided alongside link#278, since both
land in wsg_run_one.R's argument handling.

Relationships

  • link#262 — added run_uid; this is the half its dispatcher-side wiring did
    not reach, and the same class of gap that issue was opened to close
  • link#278 — make driver arguments explicit rather than defaulted; same file,
    should be decided together
  • floodplains#90 — the run that surfaced it

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions