Class RenameEliminationPass
RENAME-001 … RENAME-004) — the cleanup pass that
runs immediately after ViewInliner.
Why it exists
ViewInliner replaces every view reference with
ρ V (inline(V.body)). The wrapper is load-bearing at the moment of inlining —
it re-establishes V as an alias so a qualified V.col still resolves —
but once nothing names V, the ρ is dead weight that survives into the
final plan, sitting between operators that other passes need to see adjacent.
That is not hypothetical: SelectionIntoFixpointPass grew explicit logic to
peel interleaved relation-only renames out of its transparent prefix, and every
future shape-matching pass would pay the same tax.
The two rules
RENAME-001— collapse a ρ chain- A relation-only ρ directly above another ρ absorbs it:
ρ V (ρ W [spec] (R))→ρ V [spec] (R). Unconditional. The outer ρ re-anchors every column's provenance toV, soWis not a resolvable qualifier at the outer node's output either way — collapsing loses a name that was already invisible one level up. RENAME-002— drop an unreferenced relation-only ρρ V (X)→X, when nothing in the query namesV. A relation-only ρ leaves every column name untouched (seeSchemaInferenceVisitor#visit(RenameNode)), so the only thing removal can change is how a relation-qualified reference resolves.RENAME-003— drop identity pairsρ id→id, a→q (A)→ρ a→q (A). Dropping the pairs is unconditional; dropping the node isRENAME-002's decision, since a ρ whose pairs all cancel may still carry a relation name something resolves against. So this rule reduces the node to its relation-only form and hands it straight to the sweep below.RENAME-004— compose stacked pair-form renamesρ b→c (ρ a→b (R))→ρ a→c (R). WhereRENAME-001needs its outer ρ to be relation-only, this one composes two column renames — and is unconditional for the same reason, once the merged node keeps whichever relation name exists. Seecompose(com.darkcollective.relix.ast.RenameNode, java.lang.String, com.darkcollective.relix.optimizer.internal.OptimizationContext)for the one case it declines and for why a rename cycle needs no arm of its own.
The column rules run first at each node: composing is what turns a cycle into the
identity pairs RENAME-003 drops, and reducing a cancelled ρ to relation-only
form is what puts it in front of RENAME-002.
The reference sweep, and why it checks two things
A qualified reference resolves by column provenance and silently falls
back to a bare-name match when the qualifier names nothing
(ArrayRow#get(String)). Removing ρ V reverts the provenance of every
column beneath it from V to whatever the input carried, so two things have to
hold — and a reference can hide in a predicate, a projection, a join condition, a
grouping or sort key, so the sweep covers all of them via OperandWalker:
Vis not referenced. OtherwiseV.colstops resolving and falls back to a bare-name first match — a silent wrong answer.- No relation name the input re-exposes is referenced either. This
is the subtle half. In
(ρ V (Orders)) ⋈ Ordersa reference toOrders.idresolves to exactly one column today, because the left side is stampedV. Drop the ρ and both sides are stampedOrders: the reference becomes ambiguous and falls back to a first match. Nothing aboutVitself would have caught that.
The sweep is deliberately whole-tree rather than "above this ρ". A qualifier used only below the ρ cannot actually be affected by removing it, so treating it as a blocker is over-conservative — but it costs one traversal instead of a position-aware analysis, and an over-conservative descriptor is a no-op, never a wrong answer (ADR-0020, Decision 5).
The same principle governs the node types the sweep understands. It enumerates the operand-carrying core (σ π ρ γ τ λ δ μ, the joins, the set operations); reaching any other node — an analytic, recursive, solver or generator operator whose column references this pass does not know how to enumerate — abandons the whole pass rather than guess. Deriving "unreferenced" from an incomplete sweep is precisely how a rewrite of this shape returns wrong rows.
Where it runs
Not in OptimizationPipeline, but in
QueryOptimizer.inlineThenOptimize between ViewInliner and the schema
re-inference that follows it. Removing a node drops the identity-keyed
SchemaAnnotations entry of every node above
it, so running before the re-inference is what keeps the rule passes seeing a fully
annotated tree.
The positional form (ρ E (a, b, c)) is deliberately out of scope
for both column rules: it renames by position and is arity-bound, so composing it with
a pair form — or with another positional form — is a different question, and one
ViewInliner, the reason this pass exists, never poses.
This class is package-private and stateless; call
apply(RelNode, String, OptimizationContext) as a static method.
-
Method Summary