java.lang.Object
com.darkcollective.relix.embed.Relation

public final class Relation extends Object
A relation — an expression together with the analysis it resolves against.

A relation is a value, not a pending execution. Operators are functions from relations to relations, closed under composition; that is the algebra, and it is what this type models. Composing, optimising, rendering and inspecting a relation are whole uses of the engine in their own right, and none of them runs anything.

Pinned at creation

A relation captures the SemanticModel its session held when it was created. Redefining a source or view afterwards does not change it. Without that rule this type would be a view onto mutable state rather than a value — its meaning would shift under its holder, and it could be neither cached nor shared across threads.

Inspection is staged; execution is not

render() and optimized().render() differ; stream() and optimized().stream() do not. Every execution terminal optimises first, because a library call that quietly skipped a rewrite phase would hand back a correct answer by a worse plan with no signal that it had. optimized() exists so a caller can see and serialise the rewritten form, not to enable it.

asWritten() is how a caller opts out, and the reasoning above is what shapes it: what must not happen silently is skipping the rewriter, so the escape is a method whose name says so at the call site. It makes both trees executable from one relation, which is what comparing them needs — the rewriter preserving the answer is a claim, and running only one of the two trees cannot test it.

Since:
1.0
  • Method Details

    • label

      public Optional<String> label()
      What the script called this query, when it came from one.

      query Open; gives Open; an unnamed query { … } gives <expression 2>, numbered among the script's query statements. Present only on a relation Relix.script(java.lang.String) returned: a relation built by composing carries no label, because the composed expression is not the one the script named.

      It exists because a result needs a heading in every rendering a caller might choose, and inventing one at the point of printing would give a different answer from the one the engine's own feed uses.

      Returns:
      the label, or empty for a relation nobody named
      Since:
      1.0
    • node

      public RelNode node()
      The expression tree this relation denotes.
      Returns:
      the logical expression; never null
      Since:
      1.0
    • model

      public SemanticModel model()
      The analysis this relation is pinned to — the session's state when it was created.
      Returns:
      the semantic model; never null
      Since:
      1.0
    • select

      public Relation select(Predicate predicate)
      σ — the rows satisfying predicate.
      Parameters:
      predicate - the selection predicate; must not be null
      Returns:
      the selected relation
      Since:
      1.0
    • project

      public Relation project(List<ProjectedAttribute> attributes)
      π — the named columns, in the order given.
      Parameters:
      attributes - the projected attributes; must not be null
      Returns:
      the projected relation
      Since:
      1.0
    • project

      public Relation project(String... columns)
      π over bare column names.
      Parameters:
      columns - the column names; must not be null
      Returns:
      the projected relation
      Since:
      1.0
    • rename

      public Relation rename(String relationName)
      ρ — renames the relation.
      Parameters:
      relationName - the new relation name; must not be null
      Returns:
      the renamed relation
      Since:
      1.0
    • rename

      public Relation rename(String relationName, List<String> attributes)
      ρ — renames the relation and its columns positionally.
      Parameters:
      relationName - the new relation name; must not be null
      attributes - the new column names, one per column
      Returns:
      the renamed relation
      Since:
      1.0
    • aggregate

      public Relation aggregate(List<String> groupingColumns, List<AggregateFunction> aggregates)
      γ — grouping keys and aggregates together, because γ is one node.
      Parameters:
      groupingColumns - the bare grouping columns
      aggregates - the aggregate functions
      Returns:
      the aggregated relation
      Since:
      1.0
    • aggregateBy

      public Relation aggregateBy(List<GroupingKey> groupingKeys, List<AggregateFunction> aggregates)
      γ over grouping keys that are full expressions rather than bare columns.
      Parameters:
      groupingKeys - the grouping keys
      aggregates - the aggregate functions
      Returns:
      the aggregated relation
      Since:
      1.0
    • sort

      public Relation sort(List<SortSpecification> sortSpecs)
      τ — sorted by the given keys.
      Parameters:
      sortSpecs - the sort keys, most significant first
      Returns:
      the sorted relation
      Since:
      1.0
    • sort

      public Relation sort(SortSpecification... sortSpecs)
      τ — sorted by the given keys.
      Parameters:
      sortSpecs - the sort keys, most significant first
      Returns:
      the sorted relation
      Since:
      1.0
    • limit

      public Relation limit(long count)
      λ — at most count rows.

      Wraps rather than replaces: λ over λ is well-defined and the tighter bound wins, and rewriting a λ that may not be at the root would be smoothing over the algebra. This is also how an unbounded relation is made finite.

      Parameters:
      count - the maximum number of rows
      Returns:
      the bounded relation
      Since:
      1.0
    • limit

      public Relation limit(long offset, long count)
      λ — at most count rows, after skipping offset.
      Parameters:
      offset - how many rows to skip
      count - the maximum number of rows
      Returns:
      the bounded relation
      Since:
      1.0
    • distinct

      public Relation distinct()
      δ — duplicate rows removed.
      Returns:
      the deduplicated relation
      Since:
      1.0
    • unnest

      public Relation unnest(String column)
      μ — one row per element of the named nested column.
      Parameters:
      column - the array-valued column to unnest
      Returns:
      the unnested relation
      Since:
      1.0
    • unnest

      public Relation unnest(String column, boolean outer, Optional<String> ordinalityColumn)
      μ — unnest, keeping rows whose collection is empty when outer.
      Parameters:
      column - the array-valued column
      outer - whether to keep rows with nothing to unnest
      ordinalityColumn - a column to receive each element's position, if wanted
      Returns:
      the unnested relation
      Since:
      1.0
    • why

      public Relation why()
      ω — each row's lineage reified as a nested provenance column.
      Returns:
      the relation with provenance
      Since:
      1.0
    • forall

      public Relation forall(List<String> groupingAttributes, Predicate predicate)
      ∀ — the groups in which every row satisfies predicate.
      Parameters:
      groupingAttributes - the grouping columns; empty for the no-key form
      predicate - the condition every row must satisfy
      Returns:
      the quantified relation
      Since:
      1.0
    • join

      public Relation join(Relation right)
      ⋈ — natural join on the common columns.
      Parameters:
      right - the right input; must not be null
      Returns:
      the joined relation
      Since:
      1.0
    • join

      public Relation join(Relation right, Predicate condition)
      ⨝ — theta join on an explicit condition.
      Parameters:
      right - the right input; must not be null
      condition - the join condition
      Returns:
      the joined relation
      Since:
      1.0
    • leftJoin

      public Relation leftJoin(Relation right, Predicate condition)
      ⟕ — left outer join.
      Parameters:
      right - the right input; must not be null
      condition - the join condition
      Returns:
      the joined relation
      Since:
      1.0
    • rightJoin

      public Relation rightJoin(Relation right, Predicate condition)
      ⟖ — right outer join.
      Parameters:
      right - the right input; must not be null
      condition - the join condition
      Returns:
      the joined relation
      Since:
      1.0
    • fullJoin

      public Relation fullJoin(Relation right, Predicate condition)
      ⟗ — full outer join.
      Parameters:
      right - the right input; must not be null
      condition - the join condition
      Returns:
      the joined relation
      Since:
      1.0
    • semiJoin

      public Relation semiJoin(Relation right, Predicate condition)
      ⋉ — the left rows having a match.
      Parameters:
      right - the right input; must not be null
      condition - the join condition
      Returns:
      the semi-joined relation
      Since:
      1.0
    • antiJoin

      public Relation antiJoin(Relation right, Predicate condition)
      ▷ — the left rows having no match.
      Parameters:
      right - the right input; must not be null
      condition - the join condition
      Returns:
      the anti-joined relation
      Since:
      1.0
    • forall

      public Relation forall(Relation right, Predicate condition)
      ∀ — the pairwise universal join: left rows matching every right row.
      Parameters:
      right - the right input; must not be null
      condition - the condition every right row must satisfy
      Returns:
      the quantified relation
      Since:
      1.0
    • asOfJoin

      public Relation asOfJoin(Relation right, Predicate condition)
      AS-OF join — each left row matched to the most recent right row.
      Parameters:
      right - the right input; must not be null
      condition - the match condition
      Returns:
      the joined relation
      Since:
      1.0
    • asOfJoin

      public Relation asOfJoin(Relation right, Predicate condition, Optional<Operand> tolerance, boolean inner, TieBreak tieBreak)
      AS-OF join with a tolerance and an explicit tie-break.
      Parameters:
      right - the right input; must not be null
      condition - the match condition
      tolerance - how far back a match may be, if bounded
      inner - whether unmatched left rows are dropped
      tieBreak - which row wins when two are equally recent
      Returns:
      the joined relation
      Since:
      1.0
    • intervalJoin

      public Relation intervalJoin(Relation right, AllenRelation relation, String leftStart, String leftEnd, String rightStart, String rightEnd)
      Interval join — rows whose intervals stand in the given Allen relation.
      Parameters:
      right - the right input; must not be null
      relation - the Allen interval relation
      leftStart - the left interval's start column
      leftEnd - the left interval's end column
      rightStart - the right interval's start column
      rightEnd - the right interval's end column
      Returns:
      the joined relation
      Since:
      1.0
    • lateral

      public Relation lateral(String functionName, Operand... arguments)
      LATERAL — a table function evaluated per left row.
      Parameters:
      functionName - the table function's name
      arguments - its arguments, which may reference this relation's columns
      Returns:
      the joined relation
      Since:
      1.0
    • cross

      public Relation cross(Relation right)
      × — the Cartesian product.
      Parameters:
      right - the right input; must not be null
      Returns:
      the product
      Since:
      1.0
    • union

      public Relation union(Relation right)
      ∪ — set union, duplicates removed.
      Parameters:
      right - the right input; must not be null
      Returns:
      the union
      Since:
      1.0
    • unionAll

      public Relation unionAll(Relation right)
      ⊎ — bag union, duplicates kept.
      Parameters:
      right - the right input; must not be null
      Returns:
      the union
      Since:
      1.0
    • outerUnion

      public Relation outerUnion(Relation right)
      ⊔ — outer union over differing headings.
      Parameters:
      right - the right input; must not be null
      Returns:
      the union
      Since:
      1.0
    • difference

      public Relation difference(Relation right)
      − — the rows of this relation not in right.
      Parameters:
      right - the right input; must not be null
      Returns:
      the difference
      Since:
      1.0
    • intersect

      public Relation intersect(Relation right)
      ∩ — the rows in both.
      Parameters:
      right - the right input; must not be null
      Returns:
      the intersection
      Since:
      1.0
    • divide

      public Relation divide(Relation right)
      ÷ — relational division.
      Parameters:
      right - the divisor; must not be null
      Returns:
      the quotient
      Since:
      1.0
    • symmetricDifference

      public Relation symmetricDifference(Relation right)
      ∆ — the rows in exactly one of the two.
      Parameters:
      right - the right input; must not be null
      Returns:
      the symmetric difference
      Since:
      1.0
    • compose

      public Relation compose(Relation right)
      ∘ — relational composition.
      Parameters:
      right - the right input; must not be null
      Returns:
      the composition
      Since:
      1.0
    • closure

      public Relation closure(String fromColumn, String toColumn)
      CLOSURE — the transitive closure over an edge relation.
      Parameters:
      fromColumn - the edge's source column
      toColumn - the edge's target column
      Returns:
      the closure
      Since:
      1.0
    • closure

      public Relation closure(String fromColumn, String toColumn, boolean reflexive)
      CLOSURE, optionally reflexive.
      Parameters:
      fromColumn - the edge's source column
      toColumn - the edge's target column
      reflexive - whether every node reaches itself
      Returns:
      the closure
      Since:
      1.0
    • closure

      public Relation closure(String fromColumn, String toColumn, boolean undirected, boolean reflexive)
      CLOSURE reading its two columns as an undirected edge, so the relation is followed both ways from one edge set.
      Parameters:
      fromColumn - the first endpoint column
      toColumn - the second endpoint column
      undirected - true to read the edges both ways
      reflexive - true for R*, false for R⁺
      Returns:
      the closure
      Since:
      1.0
    • cluster

      public Relation cluster(String fromColumn, String toColumn, String labelColumn)
      CLUSTER — connected components, labelled.
      Parameters:
      fromColumn - the edge's source column
      toColumn - the edge's target column
      labelColumn - the column to receive each component's label
      Returns:
      the clustered relation
      Since:
      1.0
    • path

      public Relation path(String fromColumn, String toColumn, int minHops, int maxHops, String depthColumn)
      PATH — paths of bounded length between nodes.
      Parameters:
      fromColumn - the edge's source column
      toColumn - the edge's target column
      minHops - the minimum path length
      maxHops - the maximum path length
      depthColumn - the column to receive each path's length
      Returns:
      the paths
      Since:
      1.0
    • path

      public Relation path(String fromColumn, String toColumn, boolean undirected, int minHops, int maxHops, String depthColumn)
      PATH over an edge relation read both ways.
      Parameters:
      fromColumn - the first endpoint column
      toColumn - the second endpoint column
      undirected - true to read the edges both ways
      minHops - the inclusive lower bound on path length
      maxHops - the inclusive upper bound on path length
      depthColumn - the name of the appended shortest-distance column
      Returns:
      the bounded paths
      Since:
      1.0
    • trace

      public Relation trace(String from, String to, String weight, ObjectiveSense sense, String path)
      TRACE — the optimal path between each reachable pair, as an ordered array.
      Parameters:
      from - the edge's source column
      to - the edge's target column
      weight - the edge-weight column
      sense - whether to minimise or maximise
      path - the column to receive the path
      Returns:
      the traced relation
      Since:
      1.0
    • trace

      public Relation trace(String from, String to, boolean undirected, String weight, ObjectiveSense sense, String path)
      TRACE over a weighted edge relation read both ways, each edge traversable in either direction at the same cost.
      Parameters:
      from - the first endpoint column
      to - the second endpoint column
      undirected - true to read the edges both ways
      weight - the edge-weight column
      sense - whether to minimise or maximise the total weight
      path - the name of the appended path-array column
      Returns:
      the optimal paths
      Since:
      1.0
    • fix

      public Relation fix(String name, UnaryOperator<Relation> step)
      FIX — the least fixpoint of step, with this relation as the base.

      The step is a function because it has to refer to the relation it is building, which does not exist until this call makes it: step is handed a relation standing for name — everything derived so far — and returns the rows to add from it. That relation has this one's heading, qualified by name, and means something only inside the step; reading it anywhere else is an error.

      Relation reach = edges.fix("Reach", r ->
              r.join(edges.rename("E", List.of("s", "nxt")),
                              Expr.eq(attr("Reach.dst"), attr("E.s")))
                      .project(List.of(projected(attr("Reach.src")),
                                       projected(attr("E.nxt"), "dst"))));
      
      Parameters:
      name - the name the step refers to the relation by
      step - the recursive step, from the relation derived so far to the rows it adds
      Returns:
      the fixpoint
      Since:
      1.0
    • iterate

      public Relation iterate(String name, UnaryOperator<Relation> step, IterateStop stop)
      ITERATE — step applied round after round, with this relation as the first round and each round replacing the last, until stop says to stop.

      As for fix(java.lang.String, java.util.function.UnaryOperator<com.darkcollective.relix.embed.Relation>), the step is a function of a relation standing for name — here the previous round — with this one's heading qualified by name.

      Parameters:
      name - the name the step refers to the previous round by
      step - the step, from the previous round to the next
      stop - when to stop: AstBuilders.rounds, untilStable or untilConverged
      Returns:
      the last round
      Since:
      1.0
    • window

      public Relation window(WindowFunction function, List<String> partitionKeys, List<SortSpecification> sortSpecs, WindowFrame frame, String outputColumn)
      WINDOW — a window function over ordered partitions.
      Parameters:
      function - the window function
      partitionKeys - the partition columns
      sortSpecs - the ordering within a partition
      frame - the frame
      outputColumn - the column to receive the result
      Returns:
      the windowed relation
      Since:
      1.0
    • top

      public Relation top(List<String> groupingAttributes, List<SortSpecification> sortSpecs, long count)
      TOP — the first count rows of each group, by the given order.
      Parameters:
      groupingAttributes - the grouping columns; empty for the whole relation
      sortSpecs - the ordering
      count - how many rows per group
      Returns:
      the top rows
      Since:
      1.0
    • sessionize

      public Relation sessionize(String orderColumn, Operand threshold, List<String> partitionKeys, String sessionColumn)
      SESSIONIZE — groups consecutive rows into sessions by an inactivity threshold.
      Parameters:
      orderColumn - the column defining order
      threshold - the gap that starts a new session
      partitionKeys - the columns sessions are computed within
      sessionColumn - the column to receive the session identifier
      Returns:
      the sessionized relation
      Since:
      1.0
    • downsample

      public Relation downsample(String tsColumn, String interval, ConsolidationFunction function, List<String> keys)
      DOWNSAMPLE — consolidates rows into fixed time buckets.
      Parameters:
      tsColumn - the timestamp column
      interval - the bucket width
      function - how values within a bucket are consolidated
      keys - the columns bucketed within
      Returns:
      the downsampled relation
      Since:
      1.0
    • pivot

      public Relation pivot(String valueColumn, String keyColumn, List<String> groupKeys)
      PIVOT — turns row values into columns.
      Parameters:
      valueColumn - the column supplying the values
      keyColumn - the column supplying the new column names
      groupKeys - the columns pivoted within
      Returns:
      the pivoted relation
      Since:
      1.0
    • unpivot

      public Relation unpivot(List<String> columns, String nameColumn, String valueColumn)
      UNPIVOT — turns columns into rows.
      Parameters:
      columns - the columns to unpivot
      nameColumn - the column to receive each column's name
      valueColumn - the column to receive its value
      Returns:
      the unpivoted relation
      Since:
      1.0
    • tree

      public Relation tree(String keyColumn, String parentColumn, String childrenColumn)
      TREE — folds an adjacency list into a forest of nested documents.
      Parameters:
      keyColumn - each row's identifier
      parentColumn - its parent's identifier
      childrenColumn - the column to receive each node's children
      Returns:
      the forest
      Since:
      1.0
    • sample

      public Relation sample(double probability)
      SAMPLE — a Bernoulli sample, each row kept with the given probability.
      Parameters:
      probability - the per-row probability
      Returns:
      the sample
      Since:
      1.0
    • sample

      public Relation sample(double probability, long seed)
      SAMPLE with a seed, so the draw is reproducible.
      Parameters:
      probability - the per-row probability
      seed - the seed
      Returns:
      the sample
      Since:
      1.0
    • sampleReservoir

      public Relation sampleReservoir(long count)
      RESERVOIR SAMPLE — exactly count rows, uniformly drawn in one pass.
      Parameters:
      count - how many rows to keep
      Returns:
      the sample
      Since:
      1.0
    • sampleReservoir

      public Relation sampleReservoir(long count, long seed)
      RESERVOIR SAMPLE with a seed.
      Parameters:
      count - how many rows to keep
      seed - the seed
      Returns:
      the sample
      Since:
      1.0
    • solve

      public Relation solve(Operand left, Operand right)
      SOLVE — solves left = right for the unknown, per row.
      Parameters:
      left - the equation's left side
      right - the equation's right side
      Returns:
      the solved relation
      Since:
      1.0
    • optimize

      public Relation optimize(ObjectiveSense sense, Operand objective, List<OptimizeConstraint> constraints, List<String> groupingKeys)
      OPTIMIZE — the subset of rows optimising an objective under constraints.

      Note that this is the operator. Staging the logical rewriter is optimized(), which takes no arguments: the two are unrelated, and the operator keeps the name the reference manual gives it.

      Parameters:
      sense - whether to maximise or minimise
      objective - the quantity being optimised
      constraints - the constraints
      groupingKeys - the columns a separate problem is solved per
      Returns:
      the chosen rows
      Since:
      1.0
    • cover

      public Relation cover(int strength)
      COVER — a covering test suite of the given strength.
      Parameters:
      strength - the interaction strength
      Returns:
      the covering set
      Since:
      1.0
    • cover

      public Relation cover(int strength, boolean exact)
      COVER, optionally exact.
      Parameters:
      strength - the interaction strength
      exact - whether the cover must be minimal
      Returns:
      the covering set
      Since:
      1.0
    • render

      public String render()
      This relation as Relix text.

      Re-parseable: the text is what Relix.relation(String) reads back, so a program can hand its own optimised query to something that only speaks the language. Formatting is normalised rather than preserved, and a literal's spelling may be too — 5.0 renders as it was parsed, not as it was typed.

      Returns:
      the expression as .relix text
      Since:
      1.0
    • renderJson

      public String renderJson()
      This relation's expression tree as JSON, for a program rather than a reader.

      The logical counterpart to explainJson(): one object per operator, each carrying its inputs and the heading the analyser inferred for it. Nothing is planned or run.

      Returns:
      the tree as a JSON document
      Since:
      1.0
    • schema

      public Schema schema()
      The heading this relation produces.
      Returns:
      the output schema
      Throws:
      RelixException - if the expression carries no inferred schema, which happens only when it references a name the analyser could not resolve
      Since:
      1.0
    • optimized

      public Relation optimized()
      The logical rewriter's output — the same relation, rewritten.

      Views are inlined first, so rules optimise across a boundary the author wrote for clarity. The result carries the events of its own rewrite, which is what makes this more than "the tree, but different": events() names each rule that fired.

      Calling it is not what makes execution optimised — every execution terminal runs the rewriter regardless, unless asWritten() says otherwise. This is how a caller sees the result.

      Executing what it returns runs this tree: the rewriter is not run a second time over its own output, so the tree node() shows is the tree that runs, and stream(QueryEventListener) reports the rewrite that produced it.

      Returns:
      the rewritten relation
      Since:
      1.0
    • asWritten

      public Relation asWritten()
      This relation, to be executed as it is written — the rewriter is not run.

      The counterpart of optimized() on the execution side, and the only way to reach the unrewritten tree with rows: every terminal otherwise optimises first, because a library call that quietly ran a worse plan would give a correct answer with no signal that a phase had been skipped. Naming this method is that signal, which is why the capability is here and not a default.

      What it is for is comparing the two — that a rewrite preserves the answer is a claim, and the only way to test it is to run both trees over the same data:

      assert query.asWritten().toList().equals(query.toList());
      

      It is also how a harness measures the rewriter rather than being measured through it: a pushdown that holds for the query as written and not after the rules have run is a pessimisation, and seeing it needs both plans from one relation.

      It settles execution only. render(), explain() and plan() already describe the relation as written, so they are unchanged, and on a relation that came from optimized() this is a no-op — that tree has been rewritten already and nothing can un-rewrite it.

      It survives composition — asWritten().limit(5).stream() still runs the tree as written — because it is an instruction about how this caller wants their query run rather than a fact about one tree. optimized() is the opposite and behaves accordingly: what it marks is a fact about its tree, so a combinator on it goes back to rewriting, the added expression having been through no rule.

      Returns:
      this relation, with the rewriter off for its execution terminals
      Since:
      1.0
    • events

      public List<QueryEvent> events()
      What the rewriter did to produce this relation.

      Empty unless this relation came from optimized() — a relation nobody has asked to rewrite has had nothing done to it. Each event names a rule by its code (SEL-001, PROJ-004) and describes what it rewrote.

      Returns:
      the OPTIMIZE-stage events, in the order the rules fired
      Since:
      1.0
    • rewrites

      public List<TransformationRecord> rewrites()
      What the rewriter did, as records rather than as a feed.

      The same rewrite events() reports, in the form a consumer that wants to render it needs: each record names the rule, the relation it fired on, and what it did. An event says a rule fired; a record says what it fired on, which is the difference between a trace and a report.

      Empty unless this relation came from optimized(), for the reason events() is.

      Returns:
      the transformations, in the order they were applied
      Since:
      1.0
    • explain

      public String explain()
      The physical plan, as --explain prints it: join algorithms and build sides, what was folded into a native query and pushed to its backend, and each node's estimated row count.

      Staged like render(): this plans the relation as written, and optimized().explain() plans the rewritten one. Nothing is executed and no source is read — planning asks the catalog and the cost model, not the data.

      Returns:
      the printed plan
      Since:
      1.0
    • explain

      public String explain(QueryEventListener listener)
      The physical plan, with the planner's decisions reported to listener.

      explain() for a caller that wants the choices as well as the plan — which join algorithm, which build side, what was pushed. The printed plan shows the outcome; the events say what was decided along the way, and a host keeping a feed wants both from one planning pass rather than two.

      Parameters:
      listener - notified of each planning decision; must not be null
      Returns:
      the printed plan
      Since:
      1.0
    • explainJson

      public String explainJson()
      The physical plan as JSON, for a program rather than a reader.
      Returns:
      the plan as a JSON document
      Since:
      1.0
    • plan

      public PlannedQuery plan()
      The physical plan itself, together with the row estimates the planner computed.

      The estimates travel beside the plan rather than inside it: two structurally identical plans costed under different statistics would otherwise be unequal, so PlanEstimates is an identity-keyed side table.

      Returns:
      the plan and its estimates
      Since:
      1.0
    • count

      public long count()
      How many rows this relation has.

      This is the algebra rather than a third execution mode: it is γ COUNT(*) over the relation, an operator that already pushes down, so against a database the backend does the counting and exactly one row crosses the wire.

      Returns:
      the row count
      Throws:
      RelixException - if the relation cannot be executed
      Since:
      1.0
    • toList

      public List<Tuple> toList()
      Every row, in memory.

      The one to reach for first. It drains the row stream and closes it, so the common case — a result that fits in memory — cannot leak the JDBC connection a lazy stream holds open.

      Returns:
      the rows, in the order the query produced them
      Throws:
      UnboundedRelationException - if the relation is provably unbounded, since collecting one would never return
      QueryExecutionException - if the query runs and a source gives way
      RelixException - if the session is closed
      Since:
      1.0
    • toList

      public <T> List<T> toList(Class<T> type)
      Every row, as an instance of type.

      toList() with the mapping written for you. A record's components carry a name and a type, which is everything the mapping needs, so each component is read from the column of the same name through the accessor its type names — with the same refusal semantics Tuple has, since it is Tuple doing the reading.

      record Order(long id, String customer, BigDecimal amount) { }
      
      List<Order> orders = relix.relation("Orders").toList(Order.class);
      

      The binding is checked before any row moves. The heading is known without running anything, so a component naming no column, or one whose column holds another type, is a failure at this call rather than at row 400,000. The exceptions are what the heading cannot answer: a column typed ANY carries no declared type to check against, and whether a value is NULL is not a property of a heading at all — so a NULL read into a primitive component is refused at that row, naming the component and asking for the boxed type.

      Names must match. There is no annotation to say otherwise: the query already renames a column, and π amount → total (…) says it where a reader of the query will look. Matching is case-insensitive, as every column lookup is.

      A component may be a String, a BigDecimal, Long/long, Integer/int, Double/double, Boolean/boolean, an Instant, LocalDate, LocalTime or Duration, a List for an array column, a Map for a struct column, or a Value — which is the one to reach for where a column's type genuinely varies, and the only one that receives a NULL as a NullValue rather than as Java null.

      The record is built through its canonical constructor, reflectively. A record in a named module therefore has to be reachable from here — public in an exported package, or its package opened — and is refused by name when it is not.

      Type Parameters:
      T - the record type
      Parameters:
      type - the record to read each row into; must not be null
      Returns:
      the rows, in the order the query produced them
      Throws:
      RelixException - if type is not a record, if a component names no column, if a column holds a type the component cannot hold, or if the record cannot be constructed from here
      UnboundedRelationException - if the relation is provably unbounded, since collecting one would never return
      QueryExecutionException - if the query runs and a source gives way
      Since:
      1.0
    • stream

      public Stream<Tuple> stream()
      Every row, lazily.

      The deliberate opt-in for a caller who wants the engine's laziness and accepts what comes with it: the stream is a resource and the caller closes it, because it may hold a live database cursor. Use it in a try-with-resources.

      Closing early is how a partial read is cancelled — the engine is pull-based end to end, so nothing further is computed once nobody pulls. That is also why this needs no boundedness guard: lazily consuming a relation that never ends is what a generator is for.

      Returns:
      a lazy stream of rows; the caller must close it
      Throws:
      RelixException - if the session is closed
      Since:
      1.0
    • stream

      public Stream<Tuple> stream(QueryEventListener listener)
      Every row, lazily, with the run reported to listener as it happens.

      run() for a caller that must not collect the rows — a trace over a result too large to hold, most obviously, where what is wanted is the feed and the rows are drained and dropped.

      The rewrite's events arrive first and in full, because the rewriter has already finished by the time there is a stream to hand back; the planner's and the executor's arrive as the stream is pulled. As with stream(), the caller owns the stream and closes it.

      Parameters:
      listener - notified of each event of this run; must not be null
      Returns:
      a lazy stream of rows; the caller must close it
      Throws:
      RelixException - if the session is closed
      Since:
      1.0
    • stream

      public <T> Stream<T> stream(Class<T> type)
      Every row, lazily, as an instance of type.

      toList(Class)'s mapping over stream()'s laziness — the pairing a result too large to hold wants. The binding is still checked before the query runs, and the stream is still the caller's to close.

      Type Parameters:
      T - the record type
      Parameters:
      type - the record to read each row into; must not be null
      Returns:
      a lazy stream of records; the caller must close it
      Throws:
      RelixException - if the binding cannot be made — see toList(Class)
      Since:
      1.0
    • run

      public Rows run()
      Every row, together with what the engine did to produce it.

      toList() with the run's own event feed attached — the rules that fired, the planner's choices, what was pushed to a backend. See Rows.

      Returns:
      the rows and this run's events
      Throws:
      UnboundedRelationException - if the relation is provably unbounded
      QueryExecutionException - if the query runs and a source gives way
      RelixException - if the session is closed
      Since:
      1.0
    • provenance

      public <K> AnnotatedRelation<K> provenance(Semiring<K> semiring)
      The relation annotated over a semiring — where each output row came from, in whatever algebra the semiring names.

      Booleans give presence, ℕ gives multiplicity or a path count, the tropical semiring gives a shortest distance, and the polynomial semiring gives full lineage: the exact input tuples that produced each output tuple, and how they combined.

      This is not why(). That operator reifies lineage into a column of an ordinary relation; this annotates the whole relation over a semiring of the caller's choosing, which is why it cannot ride the row terminals.

      Evaluated from the raw logical tree — no rewriting and no pushdown, since annotation tracking is in-engine, above the federation boundary — and materialised rather than streamed, because a K-relation is canonical.

      Type Parameters:
      K - the annotation type
      Parameters:
      semiring - the annotation semiring; must not be null
      Returns:
      the annotated relation
      Throws:
      RelixException - if the session is closed
      Since:
      1.0
    • provenance

      public <K> AnnotatedRelation<K> provenance(Semiring<K> semiring, String weightColumn)
      As provenance(Semiring), reading each edge's weight from a column.

      What makes a weighted transitive closure mean something: with a tropical semiring and a distance column the annotation of a reachable pair is the shortest route to it. A row lacking the column, or carrying a null or non-numeric value, weighs the semiring's one.

      Type Parameters:
      K - the annotation type
      Parameters:
      semiring - the annotation semiring; must not be null
      weightColumn - the per-edge weight column; must not be null
      Returns:
      the annotated relation
      Throws:
      RelixException - if the session is closed
      Since:
      1.0
    • toString

      public String toString()
      Overrides:
      toString in class Object