Skip to content

Consumer chain: carry id columns (by =) for several series in one table #95

Description

@NewGraphEnvironment

If we do it: a producer can pass several series at once (stations, AOIs) through cd_baseline() → cd_anomaly() → cd_trend(), keeping its id and QA columns. If we never do: every caller splits by id, loops, and re-joins its own columns afterwards. That works, but each package writes it again.

Problem

After #92 (PR #94), the consumer functions reject more than one row per variable, period and year (series_check()), and cd_anomaly() returns a fixed set of columns. A table holding several stations therefore errors, and columns such as station_number, n_days or a data-quality fraction are dropped. The first multi-series producer (NewGraphEnvironment/wet#25, per-station flow in date windows) splits by station and re-joins its columns.

Proposed Solution

  • A by = argument (default NULL, today's behaviour) naming id columns. They are added to the grouping in cd_baseline(), cd_anomaly(), cd_trend() and cd_compare(), and to series_check()'s uniqueness key.
  • cd_anomaly() keeps the by columns in its output. Whether to pass other extra columns through is a separate call.
  • Tests on a two-series table matching two single-series runs.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions