java.lang.Object
com.darkcollective.relix.optimizer.internal.ColumnPruningPass

public final class ColumnPruningPass extends Object
Column pruning (PROJ-004) — narrows the rows flowing out of every base relation to the columns the query actually reads.

Where PROJ-001..003 are local π rewrites, this pass is a top-down "required columns" walk: each node is told which of its output columns its parent needs, derives what it therefore needs from each child, and hands that down. At a base-relation leaf the accumulated requirement is compared against the leaf's schema, and a narrowing projection is inserted when the query reads strictly fewer columns than the table has:


   γ region, SUM(amount) → total (Sales)     -- Sales(rep, region, amount, note)
     →  γ region, SUM(amount) → total (π region, amount (Sales))
 

A projection is also narrowed in place when its parent reads only some of the columns it produces, so the requirement that reaches the leaf is as small as the query allows.

Why it is worth a pass of its own

  • SqlPushdownPlanner builds its SELECT list from the schema of a bare scan, so a narrower leaf is fewer columns over the wire, not merely fewer cycles in the engine.
  • A hash join materializes its build side; narrower rows mean a smaller build.
  • Sort, Aggregate, Window, FULL OUTER, UNION and Division all buffer rows, and width multiplies straight through every one of them.

Required-column derivation

The requirement is either unconstrained (Optional.empty() — "every column this node produces may be read") or a set of unqualified, lowercased column names. The root starts unconstrained; a node that cannot prove a narrower requirement for a child hands that child an unconstrained one, which stops pruning at that edge but never below it — a π further down still starts a fresh requirement of its own.

Per-node derivation
NodeWhat each child is told it must produce
πthe attribute expressions' own column references
σparent's requirement ∪ the predicate's columns
τparent's requirement ∪ the sort keys' columns
λparent's requirement, unchanged
ρ (pair form)parent's requirement mapped back through the renames, ∪ every renamed source column
γthe grouping keys' ∪ the aggregate arguments' columns — independent of what the parent reads
× ⋈ and the conditional joins(parent's requirement ∪ the join condition's columns) restricted to that side's schema; ⋈ additionally keeps every common column, because those are its join keys
⋉ ▷as above on the left; the right side needs only the join condition's columns, since a semi/anti join emits no right column

What deliberately does not prune

  • δ and the deduplicating set operations (∪ ∩ − ∆) — they compare whole rows, so π a (δ R) ≠ δ (π a R): dropping a column before the dedup merges rows that were distinct. Same argument for ÷, ∘ and ⊔, whose semantics are defined in terms of the operand headings.
  • ⊎ — safe in principle, but its branches are matched positionally, and pruning each branch independently could align them differently. Deferred rather than risked.
  • The positional form of ρ (ρ E (a, b, c)) — it is arity-bound, so narrowing its input would break the rename.
  • μ, and every analytic / recursive / solver operator — they are reached through the default arm, which hands children an unconstrained requirement.
  • An open or unresolved leaf schema (schema-on-read sources, ADR-0001; the *:ANY placeholder) — you cannot prune what you cannot enumerate.

The pass runs last, after every pattern-matching phase has settled. It has to: an inserted π between a σ and its leaf would hide the shape GEN-001, CLOSURE-001 and their family match on, and running it before ProjectionPass would let PROJ-002 merge a freshly inserted leaf projection straight back into the one above it.

This class is package-private and stateless; call apply(RelNode, String, SchemaAnnotations, OptimizationContext) as a static method.