Class RedundantGroupingPass
AGG-001).
Rule
γ cols [] (γ keys, aggs (R)) → γ keys, aggs (R)
The rule fires when the outer AggregationNode
- has no aggregate functions of its own, and
- groups by a key set that is exactly the inner aggregation's output columns (its grouping keys plus the output name of each inner aggregate).
An aggregation emits exactly one row per distinct grouping-key combination, so its output is already distinct on its own output columns. Re-grouping that output by the full set of those columns, without computing any new aggregate, therefore reproduces the inner result row-for-row and column-for-column — the outer grouping is pure redundant work and is removed.
The output-column names are derived exactly as schema inference derives
them: a grouping attribute keeps its name, and an aggregate uses its
alias when present, otherwise the synthetic
operator_attribute name (e.g. sum_amount). Because the rule
needs only the node itself, it requires no SchemaAnnotations.
The comparison is set-based, so the outer keys may list the inner columns in any order; the surviving inner node keeps its own canonical column order. The pass deliberately does not fire when the outer aggregation has its own aggregates, or when its key set differs from the inner output columns (e.g. a strict subset), because those genuinely change the result — they are not redundant.
The pass is bottom-up: each input is rewritten before the rule is
attempted at the current node, so a stack of three or more aggregations
collapses from the inside out within a single application — the node
above sees the already-collapsed input and tests the rule against that. It
therefore needs no re-sweep of its phase
(pinned by PipelineConvergenceTest).
This class is package-private and stateless; call
apply(RelNode, String, SchemaAnnotations, OptimizationContext) as a
static method.
-
Method Summary