java.lang.Object
com.darkcollective.relix.optimizer.internal.LateralDecorrelationPass

public final class LateralDecorrelationPass extends Object
Optimization pass that decorrelates an uncorrelated LATERAL join — one whose table-valued-function arguments never reference the outer row (LATERAL-001).

Rule


   L LATERAL f(c…)  →  L × f(c…)      when no argument reads a left column
 

Why it is sound

LateralJoinNode evaluates its arguments against each left row, instantiates the function body with those values, and concatenates the left row with every row the body yields. When the arguments contain no attribute reference they evaluate to the same values for every left row, so every instantiation is the same relation — exactly the relation the leaf RelationFunctionCall denotes. Pairing that one relation with each left row is a Cartesian product, and the two agree on both parts that matter:

  • Schema. Both are left.concat(body) — the lateral's inference arm and binaryConcat produce the identical ordered heading, so positional consumers downstream see no change (the trap JOIN-003 fell into).
  • Row order. × iterates the left input in the outer loop and the materialised right side in the inner one, which is the order the lateral emits its per-row groups in.

Why it is worth doing

The executor re-invokes the planner once per left row — N argument lift-and-bind round trips and N freshly planned bodies for N rows. Decorrelation replaces all of that with one plan and one execution.

The larger win is visibility. RelNode.children() reports only the left input, so the function body is invisible to structural traversal and no other rule can look into or through the operator. Once it is a × over a RelationFunctionCall, the ordinary machinery applies: JOIN-001 can turn a selection above it into a theta join, selections push into the left input, and the planner costs it like any other product.

When it does not fire

Every condition is a way of asking the same question — would running the body once give what running it once per row gives?

  • Any attribute reference in an argument. That is the definition of a correlated lateral, and the test is deliberately blunt: any AttributeOperand anywhere in any argument blocks the rewrite. A lateral's arguments are resolved against the left schema and nothing else, so an attribute that is not a left column is a validation error rather than a decorrelation opportunity — and a false "uncorrelated" would produce wrong answers.
  • A non-deterministic argument. LATERAL f(Rand()) evaluates its argument once per left row; folding it to a single call would evaluate it once in total, which is observable.
  • A non-deterministic function body. The same objection one level down: LATERAL sampleOf(0.5) over a body reading Rand(), drawing an unseeded SAMPLE, or calling a view or nested TVF that does, produces a different relation per left row even with constant arguments. Answered by RelationDeterminism, which is why this rule needs a SymbolTable and therefore runs in QueryOptimizer's per-query preamble beside ViewInliner rather than as an OptimizationPipeline rule — every rule in the pipeline is a pure syntactic rewrite, and this one is not. A function that cannot be resolved is treated as non-deterministic.

The pass is bottom-up: each child is rewritten before the rule is attempted at the current node, so a stack of uncorrelated laterals decorrelates fully in one traversal.

This class is package-private and stateless; call apply(RelNode, SymbolTable, OptimizationContext, String) as a static method.