java.lang.Object
com.darkcollective.relix.optimizer.internal.DistinctEliminationPass

public final class DistinctEliminationPass extends Object
Optimization pass for the two δ (DISTINCT) elimination rules: one looks down from the δ (DIST-001), the other looks up from it (DIST-002).

Rules


   δ(R)                 →  R                    when R is already duplicate-free
   γ keys, aggs (δ R)   →  γ keys, aggs (R)     when every aggregate ignores multiplicity
 

"Duplicate-free" is decided by PropertyDeriver, which derives distinctness bottom-up from operator semantics: the output of γ, a set operation (∪/∩/ −/∆), ÷, transitive closure, ∀, or another δ is a set, and distinctness is carried through row-subset operators (σ/τ/λ/sampling/TOP/OPTIMIZE, the left of ⋉/▷) and renames. So δ(γ …), δ(A ∪ B), δ(δ …) and the like collapse to their input.

The derivation is purely structural, so it is unaffected by earlier passes having rewritten the input subtree (a rewritten node carries no schema annotation); this pass therefore needs no SchemaAnnotations.

DIST-002 — the δ its consumer makes irrelevant

An aggregation groups its input, so a δ beneath it is wasted work provided no aggregate counts duplicates. MIN and MAX reduce a multiset the same way they reduce the set beneath it, and a γ with no aggregates at all is pure grouping; SUM/COUNT/AVG/COLLECT genuinely read multiplicity, and ARGMAX/ARGMIN are excluded as well (deliberately conservative: their yield expression is a second reduction the rule would have to reason about). Every aggregate argument must also be deterministic — the rule changes how many rows the argument is evaluated over, so MIN(Rand()) is not invariant under it.

The saving is real rather than notional: δ is MaterializationMode.SET — it buffers a whole hash set — and it sits directly below an operator that is itself blocking.

The rule looks through a chain of operators that map each input row to at most one output row and so cannot turn "duplicates present" into a different answer above: σ, π, ρ and τ. A λ, a sample or a TOP is not transparent this way — how many rows they keep depends on how many arrive, which is exactly what the δ changes.

The pass is bottom-up: each input is rewritten before the rule is attempted at the current node, so δ(δ(R)) collapses fully and a δ exposed as redundant by an inner rewrite is also removed.

This class is package-private and stateless; call apply(RelNode, String, SchemaAnnotations, OptimizationContext) as a static method.