Relix

Language reference

Glossary (the vocabulary)

Syntax

R, S            stand for whole relations (or expressions producing one)
a, b, c         stand for attributes — the columns of a relation
σ π ρ γ τ λ δ μ  the unary operators: one relation in, one relation out
⋈ ⨝ ⟕ ⟖ ⟗ ⋉ ▷    the joins: two relations in, one out
∪ ⊎ ∩ − × ÷ ∘ ∆  the set operations
∧ ∨ ¬            the connectives that build a condition

Description

The terms this manual uses, in the order you meet them. Every one links to the page that covers it properly. If you are here before writing anything, read getting started first — it uses most of these words in a single page of script.

The data model

  • Relation — a table: a set of rows that all have the same named columns. It is the only kind of value an operator takes and the only kind it returns, which is why operators nest freely.
  • Tuple, or row — one member of a relation: a value for each column.
  • Attribute, or column — one named, typed field of every row.
  • Heading, or schema — a relation's column names and types. Relix works out the heading of every expression before running it, so an operator that cannot apply to its input is an error at analysis time.
  • Degree — how many columns a relation has. Cardinality — how many rows.
  • Type — what a column holds: NUMBER, STRING, BOOLEAN, the temporal types DATE/TIME/TIMESTAMP/DURATION, a struct, an array, or ANY for data whose shape is only known when it is read.
  • NULL — a missing value, not a zero or an empty string. Comparing it to anything yields neither true nor false, so σ drops the row; test for it with IS NULL.
  • Bag and set — a bag may hold the same row twice, a set may not. Relix keeps duplicates unless you ask for them to go: ∪ deduplicates, ⊎ does not, and δ removes them from anything.
  • Key — a set of columns whose values identify a row uniquely. The engine derives keys where it can and uses them to skip work, such as a δ over rows that are already distinct.
  • Nested relation (NF²) — a relation whose columns may themselves hold structs, arrays, or whole relations. COLLECT builds one and μ flattens it back.

Writing a script

  • Statement — one declaration, assignment or query, terminated by ;.
  • Source — a declaration binding a name to external data: a database table, a CSV or JSON file, an HTTP endpoint.
  • Base relation — a relation the data comes from directly (a source or an inline table), as opposed to one computed from others.
  • View — a name given to an expression with :=. It stores no rows; the optimizer folds the expression into whatever reads it, so naming the steps of a long query is free.
  • Query statement — query { … }, which marks a result you want back. A script with no query computes nothing.
  • Namespace and import — how one script's names are grouped and how another script reuses them.
  • Scalar function — a function over single values, Len(name), either built in or defined in the script.
  • Table-valued function (TVF) — a function returning a relation, callable anywhere a relation is expected.
  • Relationship — a named, bounded link between two relations, declared with relate or a source's references: block. Together they form the schema graph, which is what lets a join be resolved from the names alone.
  • Catalog relation — a read-only relix.* relation describing the script or the engine itself, queried like any other: the introspection stdlib and the observability feed.

Operators and expressions

  • Operator — a function from relations to a relation. Each has a symbol and an equivalent keyword (σ = SELECT); the two parse identically.
  • Predicate — a condition evaluated per row, built from comparisons, IN, LIKE and the connectives ∧ ∨ ¬. It is what goes inside a σ or a join.
  • Selection (σ) keeps rows; projection (π) keeps, reorders, renames and computes columns; rename (ρ) renames a relation or its columns.
  • Aggregate — a function reducing a group of rows to one value (SUM, COUNT). Grouping (γ) splits a relation by key and applies aggregates to each group.
  • Join — combining two relations by matching rows. Natural matches on shared column names, theta on an explicit condition, outer keeps unmatched rows with NULLs, and semi/anti filter one side by whether a match exists rather than combining columns.
  • Set operations — ∪, ∩, − over rows of the same shape; × pairs every row with every row; ÷ answers "for all of these"; ∘ joins and then drops the matched columns.
  • Quantification (∀) — keeps the groups in which every row satisfies a condition, as opposed to σ, which asks about one row at a time.
  • Recursion — CLOSURE walks a graph edge relation to reachability; FIX is the general least-fixpoint form.
  • Window — an aggregate or ranking computed over the rows around each row rather than collapsing them: ROLLING, WINDOW.
  • Provenance, or lineage — which input rows produced an output row. ω reifies it as a queryable column.

How a query runs

  • Analysis — the pass that resolves names, infers every expression's heading and reports errors, before any data is read. Its output is the semantic model the later stages read.
  • Logical plan — the tree of operators the script describes. Physical plan — what the engine will actually execute: join algorithms chosen, scans bound to connectors, work ordered.
  • Optimizer rule — a rewrite that replaces part of the logical plan with a cheaper equivalent, each with a code (SEL-001, PROJ-004) you can see fire. See the optimizer.
  • Pushdown — sending work to the system holding the data, so a filter becomes a SQL WHERE clause or a Mongo $match and the rows never travel.
  • Streaming and blocking — a streaming operator emits a row as soon as it has one; a blocking operator (sort, grouping, the deduplicating set operations) must see all its input first. Materialisation is that buffering.
  • Cardinality estimate — the planner's guess at how many rows an operator will produce, used to choose between plans. It is a guess, and :explain shows it as ~N rows.
  • Shared sub-plan — one sub-expression read by several parts of a plan and evaluated once, rather than recomputed per reader.
  • Boundedness — whether a relation is known to be finite. A generator can be unbounded, and an operator that must see every row cannot run over one.
  • Connector — the component that reads an external system (JDBC, CSV, JSON, HTTP, MongoDB) and, where it can, executes pushed-down work on the engine's behalf.

Examples

Every part of one small query, named:

Orders := [
| order_id | customer_id | amount |
|----------|-------------|--------|
| 100      | 1           | 120    |
| 101      | 1           | 80     |
| 102      | 2           | 45     |
];

BigSpenders := { τ total DESC (γ customer_id, SUM(amount) → total (σ amount ≥ 50 (Orders))) };

query { BigSpenders };

Reading it from the inside out: Orders is a base relation, here an inline table. σ amount ≥ 50 is a selection whose predicate is a comparison. γ customer_id, SUM(amount) → total is a grouping: customer_id is the grouping key, SUM the aggregate, total the output attribute. τ total DESC sorts, and is a blocking operator — it cannot emit its first row until it has seen the last. BigSpenders := names the whole expression as a view, storing nothing, and query { … } asks for it back.

Had Orders been a database source rather than an inline table, the optimizer would have pushed the selection, the grouping and the sort into one SQL statement — pushdown — and the engine would have read only the three result rows.

See Also

getting started, optimizer, repl, assignment & query, relate

Notes

Relix follows the standard relational-algebra vocabulary, so a textbook's "relation, tuple, attribute" mean here exactly what they mean there. Where SQL uses a different word for the same idea — table, row, column — the two are used interchangeably in this manual.