Class ColumnPruningPass
PROJ-004) — narrows the rows flowing out of every base
relation to the columns the query actually reads.
Where PROJ-001..003 are local π rewrites, this pass is a
top-down "required columns" walk: each node is told which of its
output columns its parent needs, derives what it therefore needs from each child,
and hands that down. At a base-relation leaf the accumulated requirement is
compared against the leaf's schema, and a narrowing projection is inserted when
the query reads strictly fewer columns than the table has:
γ region, SUM(amount) → total (Sales) -- Sales(rep, region, amount, note)
→ γ region, SUM(amount) → total (π region, amount (Sales))
A projection is also narrowed in place when its parent reads only some of the columns it produces, so the requirement that reaches the leaf is as small as the query allows.
Why it is worth a pass of its own
SqlPushdownPlannerbuilds itsSELECTlist from the schema of a bare scan, so a narrower leaf is fewer columns over the wire, not merely fewer cycles in the engine.- A hash join materializes its build side; narrower rows mean a smaller build.
Sort,Aggregate,Window,FULL OUTER,UNIONandDivisionall buffer rows, and width multiplies straight through every one of them.
Required-column derivation
The requirement is either unconstrained (Optional.empty() —
"every column this node produces may be read") or a set of unqualified,
lowercased column names. The root starts unconstrained; a node that cannot
prove a narrower requirement for a child hands that child an
unconstrained one, which stops pruning at that edge but never below it — a π
further down still starts a fresh requirement of its own.
| Node | What each child is told it must produce |
|---|---|
| π | the attribute expressions' own column references |
| σ | parent's requirement ∪ the predicate's columns |
| τ | parent's requirement ∪ the sort keys' columns |
| λ | parent's requirement, unchanged |
| ρ (pair form) | parent's requirement mapped back through the renames, ∪ every renamed source column |
| γ | the grouping keys' ∪ the aggregate arguments' columns — independent of what the parent reads |
| × ⋈ and the conditional joins | (parent's requirement ∪ the join condition's columns) restricted to that side's schema; ⋈ additionally keeps every common column, because those are its join keys |
| ⋉ ▷ | as above on the left; the right side needs only the join condition's columns, since a semi/anti join emits no right column |
What deliberately does not prune
- δ and the deduplicating set operations (
∪ ∩ − ∆) — they compare whole rows, soπ a (δ R) ≠ δ (π a R): dropping a column before the dedup merges rows that were distinct. Same argument for÷,∘and⊔, whose semantics are defined in terms of the operand headings. ⊎— safe in principle, but its branches are matched positionally, and pruning each branch independently could align them differently. Deferred rather than risked.- The positional form of ρ (
ρ E (a, b, c)) — it is arity-bound, so narrowing its input would break the rename. - μ, and every analytic / recursive / solver operator — they are reached through the default arm, which hands children an unconstrained requirement.
- An open or unresolved leaf schema (schema-on-read sources, ADR-0001;
the
*:ANYplaceholder) — you cannot prune what you cannot enumerate.
The pass runs last, after every pattern-matching phase has
settled. It has to: an inserted π between a σ and its leaf would hide the shape
GEN-001, CLOSURE-001 and their family match on, and running it
before ProjectionPass would let PROJ-002 merge a freshly inserted
leaf projection straight back into the one above it.
This class is package-private and stateless; call
apply(RelNode, String, SchemaAnnotations, OptimizationContext)
as a static method.
-
Method Summary