Skip to content

Compute the rate family's windows by an ASOF join, not a frame list (T-632) - #457

Merged
chasers merged 1 commit into
mainfrom
t-632-edge-rollups-asof
Oct 6, 2026
Merged

chasers merged 1 commit into
mainfrom
t-632-edge-rollups-asof

Conversation

@chasers

@chasers chasers commented Oct 6, 2026

Copy link
Copy Markdown
Owner

TL;DR: A pushed rate, irate, increase, delta (and the rest of their family) now costs the same at any window. rate(m[1h]) over 7 d on 2.8 M samples: 22.3 s → 1.9 s, same answers.

Tracker: T-632.

Why

  • The rate family gave every sample a list of every sample in its window (list(...) OVER a RANGE frame).
  • Cost was samples × window. Grafana's $__rate_interval grows with the range, so long ranges paid twice.
  • On the sandbox, sum(rate(m[1h])) over 7 d took 30.7 s, then 134 s. It pushed the query pod to 10 GiB and the liveness probe killed it.
  • The scan was not the cost: count(*) over the same 2.8 M rows takes 0.4 s.

What changed

  • Pushdown.Windows has a second descriptor for the edge rollups: rate, deriv_fast, increase, increase_pure, delta, idelta, irate, ideriv, first_over_time, present_over_time.
  • Each sample unnests into the points where it is the window's last. That part is unchanged.
  • An ASOF join to the first row of each timestamp group finds the window's first sample.
  • Everything else a point needs is a lag/lead on those two rows: the sample before the window, the one after the first, the one before the last, and irate's earlier sample.
  • Cost: samples + series × points, whatever the window.
  • The edge rollups are now pushed past the 32-step gate (rate(m[1d]) at 15 s no longer falls back to Elixir).
  • The whole-window rollups (stddev_over_time, changes, …) keep the list descriptor and the gate.
  • Upgrade note in docs/deployment.md.

Measured

2.8 M synthetic counter samples, 7 d, step 600 s, sum by (op):

query before after same answer
rate(m[615s]) 7.0 s 1.8 s ✅
rate(m[30m]) 11.9 s 1.6 s ✅
rate(m[1h]) 22.3 s 1.9 s ✅
irate(m[1h]) 22.1 s 2.6 s ✅

Tests

  • ✅ Over-the-stack test: the pushed answer equals the Elixir evaluator's, byte for byte. Range, off-grid and instant grids, with duplicate timestamps and counter resets.
  • ✅ That test adds windows of 9 m to 1 h at a 15 s step (more than 32 steps), including a window longer than the data.
  • ✅ windows_test: SQL shape of the ASOF path and of the list path.
  • ✅ plan/2: edge rollups are planned at 1d; whole-window ones are still not.
  • ✅ mix precommit, reach.check --arch --smells.

Review

  • A Fable-model review checked every descriptor column against the old list path, including duplicate timestamps, empty windows, per-series windows, instant queries and a series' first sample. Nothing blocking.
  • Its small findings are applied: an unused window column is gone, the doc numbers match the measurements, and a longer-than-data window was added to the tests.

🤖 Generated with Claude Code

…T-632)

The rollups read off a window's edges (rate, deriv_fast, increase,
increase_pure, delta, idelta, irate, ideriv, first_over_time,
present_over_time) gave every sample `list(...) OVER` its RANGE frame, so
work and memory were samples x window / scrape interval. On the sandbox
`sum(rate(m[1h]))` over 7 d took 30.7 s and 10 GiB for 2.8 M samples.

They now get a descriptor of their own: each sample is unnested into the
points where it is the window's last (or the sample before an empty
window), and an ASOF join to the first row of each timestamp group finds
the window's first sample. Everything else a point needs (the sample before
the window, the one after the first, the one before the last, irate's
earlier) is a lag or lead carried on those two rows. The cost is samples
plus series x points at any window, so these rollups are pushed past the
32-step gate too. The whole-window rollups keep the list descriptor and
the gate.

On 2.8 M synthetic samples, identical answers: rate[615s] 7.0 s -> 1.8 s,
rate[30m] 11.9 s -> 1.6 s, rate[1h] 22.3 s -> 1.9 s, irate[1h] 22.1 s ->
2.6 s. The over-the-stack equivalence test gains windows of 9 m to 1 h at
a 15 s step.
@chasers
chasers merged commit 282a10f into main Oct 6, 2026
7 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant