Class SelectionIntoTimeSeriesPass
SESSION-001 / DOWNSAMPLE-001).
σ k = c (SESSIONIZE ts GAP … PER k AS s (R)) → SESSIONIZE … PER k AS s (σ k = c (R))
σ k = c (DOWNSAMPLE ts BY '5m' USING AVG PER k (R))
→ DOWNSAMPLE … PER k (σ k = c (R))
Both compute independently per partition — a session boundary never spans a
PER key, and a bucket is keyed by (PER keys, bucket) — which is the
identical soundness argument ADR-0020 already accepted for WINDOW,
TOP and OPTIMIZE. Both are also
MaterializationMode.BAG: they buffer and process
every partition to produce the one the query asked for, so the saving is real
work, not just rows discarded a little later.
The descriptor and the traversal are PartitionPruning's; this class supplies
only the two node shapes.
FOR n ROWS blocks the DOWNSAMPLE push
DownsampleNode's maxRows keeps the n most recent buckets
globally, not per group — DownsampleExecutor sorts all output
rows by bucket descending and takes the first n. So with two groups and
FOR 2 ROWS, filtering afterwards can leave one row while pushing the filter
first would produce two. The rule therefore does not fire at all when
maxRows is present. (The issue did not flag this; it is a genuine
counterexample rather than a conservative choice.)
Out of scope, and noted as a follow-on: a predicate on DOWNSAMPLE's
timestamp column. It is a range restriction on the buckets rather
than a partition selection — a different and more interesting rewrite, which has to
reason about bucket boundaries rather than just about which groups exist.
This class is package-private and stateless; call
apply(RelNode, String, SchemaAnnotations, OptimizationContext) as a static
method.
-
Method Summary