java.lang.Object
com.darkcollective.relix.processor.provenance.AnnotatedRelation<K>
Type Parameters:
K - the semiring annotation type

public final class AnnotatedRelation<K> extends Object
A K-relation: a relation in which every distinct tuple carries an annotation drawn from a Semiring (after Green, Karvounarakis & Tannen, PODS 2007). This is the opt-in annotated-relation model — it is built only when a provenance mode is active; the engine's default Stream<Row> path is untouched and pays nothing.

The relation is held in canonical form: each distinct tuple appears exactly once, mapped to the ⊕-combination of every derivation that produced it, and no tuple maps to zero (a zero annotation means "absent"). Tuple identity is the Row's own structural equals/hashCode — the same value-equality the engine's DISTINCT/set operators already use — so equal tuples are merged exactly as a set relation would dedup them, but combining their annotations instead of discarding the duplicate. Insertion order is preserved for stable output.

Choosing the semiring chooses the meaning of the annotation: the boolean semiring reproduces set semantics (present / absent), the ℕ semiring reproduces bag multiplicity, and so on. This class is the carrier model and its defining operations — ProvenanceEvaluator threads the annotations through the positive operators (σ/π/×/⋈/∪): lift a plain relation in, normalise an annotated stream to canonical form, and forget the annotations back out.

  • Method Details

    • lift

      public static <K> AnnotatedRelation<K> lift(Schema schema, Semiring<K> semiring, Stream<Row> rows)
      Lifts a plain relation into a K-relation by annotating every input row with the semiring's one, combining duplicate tuples with ⊕. This is the canonical embedding of ordinary data: under the boolean semiring a tuple becomes simply "present", and under ℕ a tuple's annotation becomes its multiplicity in the input bag.
      Type Parameters:
      K - the annotation type
      Parameters:
      schema - the relation's schema; never null
      semiring - the annotation semiring; never null
      rows - the plain input rows; never null
      Returns:
      the canonical K-relation
    • normalise

      public static <K> AnnotatedRelation<K> normalise(Schema schema, Semiring<K> semiring, Stream<Annotated<K>> annotated)
      Reduces a stream of Annotated rows to canonical form: tuples that are structurally equal have their annotations combined with ⊕ (in encounter order), and any tuple whose combined annotation equals zero is dropped (it is absent). This is the operation that makes an annotated multiset a well-defined K-relation.
      Type Parameters:
      K - the annotation type
      Parameters:
      schema - the relation's schema; never null
      semiring - the annotation semiring; never null
      annotated - the annotated input rows; never null
      Returns:
      the canonical K-relation
    • schema

      public Schema schema()
      Returns the schema shared by every tuple.
      Returns:
      the schema shared by every tuple
    • semiring

      public Semiring<K> semiring()
      Returns the semiring over which this relation's annotations are combined.
      Returns:
      the semiring over which this relation's annotations are combined
    • annotationOf

      public K annotationOf(Row row)
      Returns this tuple's annotation, or the semiring's zero if the tuple is absent from the relation.
      Parameters:
      row - the tuple to look up; never null
      Returns:
      the tuple's annotation, or zero if absent
    • size

      public int size()
      Returns the number of distinct present (non-zero) tuples.
      Returns:
      the number of distinct present (non-zero) tuples
    • isEmpty

      public boolean isEmpty()
      Returns whether the relation has no present tuples.
      Returns:
      whether the relation has no present tuples
    • stream

      public Stream<Annotated<K>> stream()
      Returns a lazy stream of the present tuples paired with their annotations, in insertion order.
      Returns:
      a lazy stream of the present tuples paired with their annotations, in insertion order
    • rows

      public Stream<Row> rows()
      Forgets the annotations, yielding the present tuples as a plain row stream in insertion order — the bridge back to the engine's ordinary Stream<Row> path (and the basis for surfacing provenance without exposing it as a column).
      Returns:
      a lazy stream of the present rows
    • select

      public AnnotatedRelation<K> select(Predicate<Row> predicate)
      Selection (σ): keeps each tuple satisfying predicate with its annotation unchanged, and drops the rest. In semiring terms a kept tuple is multiplied by one (identity) and a dropped tuple by zero (absent), so this is simply a filter — no two surviving tuples can collide, so no ⊕ is needed.
      Parameters:
      predicate - the row predicate; never null
      Returns:
      a K-relation of the matching tuples
    • project

      public AnnotatedRelation<K> project(Schema outputSchema, UnaryOperator<Row> map)
      Projection (π): rewrites each tuple to outputSchema via map, combining the annotations of tuples that collapse to the same output tuple with ⊕ (alternative derivations of one output row). Under ℕ this adds the multiplicities of merged rows; under booleans it is plain duplicate elimination.
      Parameters:
      outputSchema - the projected schema; never null
      map - maps an input row to its projected row; never null
      Returns:
      the projected K-relation
    • product

      public AnnotatedRelation<K> product(AnnotatedRelation<K> other, Schema outputSchema, BinaryOperator<Row> concat)
      Cartesian product (×): every pair of tuples, one from each side, combined by concat and annotated times(k₁, k₂). Equivalent to a join whose match always holds.
      Parameters:
      other - the right relation; must share this relation's semiring
      outputSchema - the combined schema; never null
      concat - combines a left and right row into the output row; never null
      Returns:
      the product K-relation
    • join

      public AnnotatedRelation<K> join(AnnotatedRelation<K> other, BiPredicate<Row,Row> match, Schema outputSchema, BinaryOperator<Row> concat)
      Join (⋈/⨝): pairs of tuples satisfying match, combined by concat and annotated times(k₁, k₂) — the joint-requirement ⊗. Output tuples that coincide have their annotations ⊕-combined. Under booleans this is the ordinary join; under ℕ the result multiplicity is the product of the matched multiplicities.
      Parameters:
      other - the right relation; must share this relation's semiring
      match - whether a left/right row pair joins; never null
      outputSchema - the combined schema; never null
      concat - combines a matched left and right row; never null
      Returns:
      the join K-relation
    • union

      public AnnotatedRelation<K> union(AnnotatedRelation<K> other)
      Union (∪): the union-compatible combination of two K-relations, with a shared tuple's annotations combined by ⊕ (it is derivable from either side). Under booleans this is set union; under ℕ the multiplicities add (bag union). Both relations must have equal schemas and share the semiring.
      Parameters:
      other - the other relation; same schema and semiring
      Returns:
      the union K-relation