Repository navigation
Compute the rate family's windows by an ASOF join, not a frame list (T-632) - #457
Merged
Merged
Conversation
…T-632) The rollups read off a window's edges (rate, deriv_fast, increase, increase_pure, delta, idelta, irate, ideriv, first_over_time, present_over_time) gave every sample `list(...) OVER` its RANGE frame, so work and memory were samples x window / scrape interval. On the sandbox `sum(rate(m[1h]))` over 7 d took 30.7 s and 10 GiB for 2.8 M samples. They now get a descriptor of their own: each sample is unnested into the points where it is the window's last (or the sample before an empty window), and an ASOF join to the first row of each timestamp group finds the window's first sample. Everything else a point needs (the sample before the window, the one after the first, the one before the last, irate's earlier) is a lag or lead carried on those two rows. The cost is samples plus series x points at any window, so these rollups are pushed past the 32-step gate too. The whole-window rollups keep the list descriptor and the gate. On 2.8 M synthetic samples, identical answers: rate[615s] 7.0 s -> 1.8 s, rate[30m] 11.9 s -> 1.6 s, rate[1h] 22.3 s -> 1.9 s, irate[1h] 22.1 s -> 2.6 s. The over-the-stack equivalence test gains windows of 9 m to 1 h at a 15 s step.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
TL;DR: A pushed
rate,irate,increase,delta(and the rest of their family) now costs the same at any window.rate(m[1h])over 7 d on 2.8 M samples: 22.3 s → 1.9 s, same answers.Tracker: T-632.
Why
list(...) OVERa RANGE frame).$__rate_intervalgrows with the range, so long ranges paid twice.sum(rate(m[1h]))over 7 d took 30.7 s, then 134 s. It pushed the query pod to 10 GiB and the liveness probe killed it.count(*)over the same 2.8 M rows takes 0.4 s.What changed
Pushdown.Windowshas a second descriptor for the edge rollups:rate,deriv_fast,increase,increase_pure,delta,idelta,irate,ideriv,first_over_time,present_over_time.lag/leadon those two rows: the sample before the window, the one after the first, the one before the last, and irate's earlier sample.rate(m[1d])at 15 s no longer falls back to Elixir).stddev_over_time,changes, …) keep the list descriptor and the gate.docs/deployment.md.Measured
2.8 M synthetic counter samples, 7 d, step 600 s,
sum by (op):rate(m[615s])rate(m[30m])rate(m[1h])irate(m[1h])Tests
windows_test: SQL shape of the ASOF path and of the list path.plan/2: edge rollups are planned at1d; whole-window ones are still not.mix precommit,reach.check --arch --smells.Review
🤖 Generated with Claude Code