Class CatalogSnapshot
- All Implemented Interfaces:
CatalogProvider
A relation composes, renders and optimises offline already. Metadata, however, genuinely does come from the database, so "works offline" and "works well" pull apart unless it can be supplied some other way. There are three ways over one seam — live introspection, a schema declared in the script, and this: introspect once, serialize, replay.
It is a CatalogProvider, so replaying it is a matter of handing it to the
analyser where a live provider would go. Statistics ride the same seam as schemas, so a
snapshot carries both and a snapshot-backed session reaches the same cost-based
decisions the live one would: candidate keys for merge joins, distinct counts for
selectivity, row counts for the build side of a hash join. The practical shape of that
is tuning or validating a production query from a machine with no access to production.
// Against the live database, once:
CatalogSnapshot snapshot = CatalogSnapshot.capture(liveCatalog, model, Clock.systemUTC());
Files.writeString(path, snapshot.toJson());
// Anywhere, afterwards:
CatalogProvider offline = CatalogSnapshot.parse(Files.readString(path));
It is a log of observations, not a merged picture
Every CatalogSnapshot.Entry records where it came from and when. That is not decoration: a
snapshot is a cache, every cache goes stale, and a plan optimised from statistics that
no longer describe the database is harder to notice than a missing plan, because it is
still a correct answer arrived at badly.
So entries are never merged into one another. A table may carry several — an
introspected one and, later, one recorded from a run that measured the real thing — and
a lookup picks per field: CatalogSnapshot.Origin.OBSERVED ahead of CatalogSnapshot.Origin.INTROSPECTED,
and within one origin the most recent. A measured row count therefore beats a
DatabaseMetaData estimate without discarding the schema that came with the
estimate, and nothing has to invent an origin for an entry assembled from two.
What it is not
Statistics are advisory everywhere in the engine, and a snapshot changes nothing about that: a stale row count costs a worse plan, never a wrong answer. A stale schema is different — a column that has since been dropped resolves here and fails at execution — which is what the capture time is for.
-
Nested Class Summary
Nested ClassesModifier and TypeClassDescriptionstatic final recordOne observation of one table.static enumWhere one entry's knowledge came from. -
Field Summary
Fields inherited from interface com.darkcollective.relix.semantic.CatalogProvider
NONE -
Method Summary
Modifier and TypeMethodDescriptionstatic CatalogSnapshotcapture(CatalogProvider live, SemanticModel model, Clock clock) Introspects every connection-backed tablemodelnames, throughlive, and records what comes back.captured()Returns when this snapshot was assembled.static CatalogSnapshotempty()A snapshot that knows nothing — the shapeCatalogProvider.NONEalready has.entries()Returns every observation, in the order it was recorded.static CatalogSnapshotof(Collection<CatalogSnapshot.Entry> entries, Instant captured) A snapshot over observations the caller already holds.static CatalogSnapshotReads a snapshot written bytoJson().tables(ConnectionDeclaration connection) Returns the names of the tablesconnectionholds, as far as this provider knows them, orOptional.empty()if it cannot enumerate them.tableSchema(ConnectionDeclaration connection, String table) Returns the output schema oftablewithinconnection, orOptional.empty()if it cannot be determined.tableStatistics(ConnectionDeclaration connection, String table) Returns optimizer statistics (row count, keys) fortablewithinconnection, orOptional.empty()if none are available.toJson()This snapshot as JSON.toString()with(CatalogSnapshot.Entry entry) This snapshot withentryadded.
-
Method Details
-
of
A snapshot over observations the caller already holds.- Parameters:
entries- the observations, in any order; must not be nullcaptured- when the snapshot as a whole was assembled- Returns:
- the snapshot
-
empty
A snapshot that knows nothing — the shapeCatalogProvider.NONEalready has. -
capture
Introspects every connection-backed tablemodelnames, throughlive, and records what comes back.The model is what says which tables matter: a source bound to a connection declares one, and a dotted reference the analyser has already resolved has become one. So capturing against a live database yields exactly the metadata the same script will ask for again offline — no more, and nothing guessed at.
A table the live provider says nothing about is not recorded. An entry carrying neither a schema nor statistics is indistinguishable, on replay, from a table the snapshot never heard of.
- Parameters:
live- the provider to introspect through; must not be nullmodel- the analysed model naming the tables; must not be nullclock- the clock the capture time is read from; must not be null- Returns:
- the snapshot
-
with
This snapshot withentryadded.Added rather than merged: an observation is a fact about a moment, and a table that has been both introspected and measured carries both.
- Parameters:
entry- the observation; must not be null- Returns:
- a new snapshot
-
entries
Returns every observation, in the order it was recorded.- Returns:
- every observation, in the order it was recorded
-
captured
Returns when this snapshot was assembled.- Returns:
- when this snapshot was assembled
-
tableSchema
Description copied from interface:CatalogProviderReturns the output schema oftablewithinconnection, orOptional.empty()if it cannot be determined.- Specified by:
tableSchemain interfaceCatalogProvider- Parameters:
connection- the declared connection the table belongs totable- the (possibly schema-qualified) remote table name- Returns:
- the table's schema, or empty if unavailable
-
tableStatistics
Description copied from interface:CatalogProviderReturns optimizer statistics (row count, keys) fortablewithinconnection, orOptional.empty()if none are available.The default returns empty; introspecting providers override this to query the live catalog. Statistics are advisory: a missing or partial value never affects correctness, only the quality of cost-based decisions.
- Specified by:
tableStatisticsin interfaceCatalogProvider- Parameters:
connection- the declared connection the table belongs totable- the (possibly schema-qualified) remote table name- Returns:
- the table's statistics, or empty if unavailable
-
tables
Returns the names of the tablesconnectionholds, as far as this provider knows them, orOptional.empty()if it cannot enumerate them.The analyser uses the list to suggest a table when a reference names one the provider cannot describe, as it already suggests a column. The default returns empty, which suits live introspection: asking a database for every table to correct one typo is not a cost worth paying during analysis. A provider holding a fixed set of tables, such as a
CatalogSnapshot, lists them.A snapshot knows exactly the tables it recorded for the connection, so it always enumerates them, in the order they were first recorded.
- Specified by:
tablesin interfaceCatalogProvider- Parameters:
connection- the declared connection- Returns:
- the table names, as a reference would spell them after the connection's name; or empty if this provider cannot enumerate them
-
toJson
This snapshot as JSON.Round-trips through
parse(String). A type is written structurally rather than as its compact IR code, because a code needs a parser and a parser is a second grammar to keep in step with the first.- Returns:
- the document
-
parse
Reads a snapshot written bytoJson().- Parameters:
json- the document; must not be null- Returns:
- the snapshot
- Throws:
IllegalArgumentException- if the text is not a snapshot this version reads
-
toString
-