Language reference
Glossary (the vocabulary)
Syntax
R, S stand for whole relations (or expressions producing one)
a, b, c stand for attributes — the columns of a relation
σ π ρ γ τ λ δ μ the unary operators: one relation in, one relation out
⋈ ⨝ ⟕ ⟖ ⟗ ⋉ ▷ the joins: two relations in, one out
∪ ⊎ ∩ − × ÷ ∘ ∆ the set operations
∧ ∨ ¬ the connectives that build a condition
Description
The terms this manual uses, in the order you meet them. Every one links to the page that covers it properly. If you are here before writing anything, read getting started first — it uses most of these words in a single page of script.
The data model
- Relation — a table: a set of rows that all have the same named columns. It is the only kind of value an operator takes and the only kind it returns, which is why operators nest freely.
- Tuple, or row — one member of a relation: a value for each column.
- Attribute, or column — one named, typed field of every row.
- Heading, or schema — a relation's column names and types. Relix works out the heading of every expression before running it, so an operator that cannot apply to its input is an error at analysis time.
- Degree — how many columns a relation has. Cardinality — how many rows.
- Type — what a column holds:
NUMBER,STRING,BOOLEAN, the temporal typesDATE/TIME/TIMESTAMP/DURATION, a struct, an array, orANYfor data whose shape is only known when it is read. - NULL — a missing value, not a zero or an empty string. Comparing it to anything yields neither true nor false, so
σdrops the row; test for it withIS NULL. - Bag and set — a bag may hold the same row twice, a set may not. Relix keeps duplicates unless you ask for them to go:
∪deduplicates,⊎does not, andδremoves them from anything. - Key — a set of columns whose values identify a row uniquely. The engine derives keys where it can and uses them to skip work, such as a
δover rows that are already distinct. - Nested relation (NF²) — a relation whose columns may themselves hold structs, arrays, or whole relations.
COLLECTbuilds one andμflattens it back.
Writing a script
- Statement — one declaration, assignment or
query, terminated by;. - Source — a declaration binding a name to external data: a database table, a CSV or JSON file, an HTTP endpoint.
- Base relation — a relation the data comes from directly (a source or an inline table), as opposed to one computed from others.
- View — a name given to an expression with
:=. It stores no rows; the optimizer folds the expression into whatever reads it, so naming the steps of a long query is free. - Query statement —
query { … }, which marks a result you want back. A script with noquerycomputes nothing. - Namespace and import — how one script's names are grouped and how another script reuses them.
- Scalar function — a function over single values,
Len(name), either built in or defined in the script. - Table-valued function (TVF) — a function returning a relation, callable anywhere a relation is expected.
- Relationship — a named, bounded link between two relations, declared with
relateor a source'sreferences:block. Together they form the schema graph, which is what lets a join be resolved from the names alone. - Catalog relation — a read-only
relix.*relation describing the script or the engine itself, queried like any other: the introspection stdlib and the observability feed.
Operators and expressions
- Operator — a function from relations to a relation. Each has a symbol and an equivalent keyword (
σ=SELECT); the two parse identically. - Predicate — a condition evaluated per row, built from comparisons,
IN,LIKEand the connectives∧ ∨ ¬. It is what goes inside aσor a join. - Selection (
σ) keeps rows; projection (π) keeps, reorders, renames and computes columns; rename (ρ) renames a relation or its columns. - Aggregate — a function reducing a group of rows to one value (
SUM,COUNT). Grouping (γ) splits a relation by key and applies aggregates to each group. - Join — combining two relations by matching rows. Natural matches on shared column names, theta on an explicit condition, outer keeps unmatched rows with NULLs, and semi/anti filter one side by whether a match exists rather than combining columns.
- Set operations —
∪,∩,−over rows of the same shape;×pairs every row with every row;÷answers "for all of these";∘joins and then drops the matched columns. - Quantification (
∀) — keeps the groups in which every row satisfies a condition, as opposed toσ, which asks about one row at a time. - Recursion —
CLOSUREwalks a graph edge relation to reachability;FIXis the general least-fixpoint form. - Window — an aggregate or ranking computed over the rows around each row rather than collapsing them:
ROLLING,WINDOW. - Provenance, or lineage — which input rows produced an output row.
ωreifies it as a queryable column.
How a query runs
- Analysis — the pass that resolves names, infers every expression's heading and reports errors, before any data is read. Its output is the semantic model the later stages read.
- Logical plan — the tree of operators the script describes. Physical plan — what the engine will actually execute: join algorithms chosen, scans bound to connectors, work ordered.
- Optimizer rule — a rewrite that replaces part of the logical plan with a cheaper equivalent, each with a code (
SEL-001,PROJ-004) you can see fire. See the optimizer. - Pushdown — sending work to the system holding the data, so a filter becomes a SQL
WHEREclause or a Mongo$matchand the rows never travel. - Streaming and blocking — a streaming operator emits a row as soon as it has one; a blocking operator (sort, grouping, the deduplicating set operations) must see all its input first. Materialisation is that buffering.
- Cardinality estimate — the planner's guess at how many rows an operator will produce, used to choose between plans. It is a guess, and
:explainshows it as~N rows. - Shared sub-plan — one sub-expression read by several parts of a plan and evaluated once, rather than recomputed per reader.
- Boundedness — whether a relation is known to be finite. A generator can be unbounded, and an operator that must see every row cannot run over one.
- Connector — the component that reads an external system (JDBC, CSV, JSON, HTTP, MongoDB) and, where it can, executes pushed-down work on the engine's behalf.
Examples
Every part of one small query, named:
Orders := [
| order_id | customer_id | amount |
|----------|-------------|--------|
| 100 | 1 | 120 |
| 101 | 1 | 80 |
| 102 | 2 | 45 |
];
BigSpenders := { τ total DESC (γ customer_id, SUM(amount) → total (σ amount ≥ 50 (Orders))) };
query { BigSpenders };
Orders := [
| order_id | customer_id | amount |
|----------|-------------|--------|
| 100 | 1 | 120 |
| 101 | 1 | 80 |
| 102 | 2 | 45 |
];
BigSpenders := { SORT total DESC (GROUP customer_id, SUM(amount) -> total (SELECT amount >= 50 (Orders))) };
query { BigSpenders };
Reading it from the inside out: Orders is a base relation, here an inline table. σ amount ≥ 50 is a selection whose predicate is a comparison. γ customer_id, SUM(amount) → total is a grouping: customer_id is the grouping key, SUM the aggregate, total the output attribute. τ total DESC sorts, and is a blocking operator — it cannot emit its first row until it has seen the last. BigSpenders := names the whole expression as a view, storing nothing, and query { … } asks for it back.
Had Orders been a database source rather than an inline table, the optimizer would have pushed the selection, the grouping and the sort into one SQL statement — pushdown — and the engine would have read only the three result rows.
See Also
getting started, optimizer, repl, assignment & query, relate
Notes
Relix follows the standard relational-algebra vocabulary, so a textbook's "relation, tuple, attribute" mean here exactly what they mean there. Where SQL uses a different word for the same idea — table, row, column — the two are used interchangeably in this manual.