java.lang.Object
com.darkcollective.relix.semantic.CatalogSnapshot
All Implemented Interfaces:
CatalogProvider

public final class CatalogSnapshot extends Object implements CatalogProvider
Table metadata captured once and replayed later, so a session composes and costs a query with no database in reach.

A relation composes, renders and optimises offline already. Metadata, however, genuinely does come from the database, so "works offline" and "works well" pull apart unless it can be supplied some other way. There are three ways over one seam — live introspection, a schema declared in the script, and this: introspect once, serialize, replay.

It is a CatalogProvider, so replaying it is a matter of handing it to the analyser where a live provider would go. Statistics ride the same seam as schemas, so a snapshot carries both and a snapshot-backed session reaches the same cost-based decisions the live one would: candidate keys for merge joins, distinct counts for selectivity, row counts for the build side of a hash join. The practical shape of that is tuning or validating a production query from a machine with no access to production.

// Against the live database, once:
CatalogSnapshot snapshot = CatalogSnapshot.capture(liveCatalog, model, Clock.systemUTC());
Files.writeString(path, snapshot.toJson());

// Anywhere, afterwards:
CatalogProvider offline = CatalogSnapshot.parse(Files.readString(path));

It is a log of observations, not a merged picture

Every CatalogSnapshot.Entry records where it came from and when. That is not decoration: a snapshot is a cache, every cache goes stale, and a plan optimised from statistics that no longer describe the database is harder to notice than a missing plan, because it is still a correct answer arrived at badly.

So entries are never merged into one another. A table may carry several — an introspected one and, later, one recorded from a run that measured the real thing — and a lookup picks per field: CatalogSnapshot.Origin.OBSERVED ahead of CatalogSnapshot.Origin.INTROSPECTED, and within one origin the most recent. A measured row count therefore beats a DatabaseMetaData estimate without discarding the schema that came with the estimate, and nothing has to invent an origin for an entry assembled from two.

What it is not

Statistics are advisory everywhere in the engine, and a snapshot changes nothing about that: a stale row count costs a worse plan, never a wrong answer. A stale schema is different — a column that has since been dropped resolves here and fails at execution — which is what the capture time is for.

  • Method Details

    • of

      public static CatalogSnapshot of(Collection<CatalogSnapshot.Entry> entries, Instant captured)
      A snapshot over observations the caller already holds.
      Parameters:
      entries - the observations, in any order; must not be null
      captured - when the snapshot as a whole was assembled
      Returns:
      the snapshot
    • empty

      public static CatalogSnapshot empty()
      A snapshot that knows nothing — the shape CatalogProvider.NONE already has.
    • capture

      public static CatalogSnapshot capture(CatalogProvider live, SemanticModel model, Clock clock)
      Introspects every connection-backed table model names, through live, and records what comes back.

      The model is what says which tables matter: a source bound to a connection declares one, and a dotted reference the analyser has already resolved has become one. So capturing against a live database yields exactly the metadata the same script will ask for again offline — no more, and nothing guessed at.

      A table the live provider says nothing about is not recorded. An entry carrying neither a schema nor statistics is indistinguishable, on replay, from a table the snapshot never heard of.

      Parameters:
      live - the provider to introspect through; must not be null
      model - the analysed model naming the tables; must not be null
      clock - the clock the capture time is read from; must not be null
      Returns:
      the snapshot
    • with

      This snapshot with entry added.

      Added rather than merged: an observation is a fact about a moment, and a table that has been both introspected and measured carries both.

      Parameters:
      entry - the observation; must not be null
      Returns:
      a new snapshot
    • entries

      public List<CatalogSnapshot.Entry> entries()
      Returns every observation, in the order it was recorded.
      Returns:
      every observation, in the order it was recorded
    • captured

      public Instant captured()
      Returns when this snapshot was assembled.
      Returns:
      when this snapshot was assembled
    • tableSchema

      public Optional<Schema> tableSchema(ConnectionDeclaration connection, String table)
      Description copied from interface: CatalogProvider
      Returns the output schema of table within connection, or Optional.empty() if it cannot be determined.
      Specified by:
      tableSchema in interface CatalogProvider
      Parameters:
      connection - the declared connection the table belongs to
      table - the (possibly schema-qualified) remote table name
      Returns:
      the table's schema, or empty if unavailable
    • tableStatistics

      public Optional<RelationStatistics> tableStatistics(ConnectionDeclaration connection, String table)
      Description copied from interface: CatalogProvider
      Returns optimizer statistics (row count, keys) for table within connection, or Optional.empty() if none are available.

      The default returns empty; introspecting providers override this to query the live catalog. Statistics are advisory: a missing or partial value never affects correctness, only the quality of cost-based decisions.

      Specified by:
      tableStatistics in interface CatalogProvider
      Parameters:
      connection - the declared connection the table belongs to
      table - the (possibly schema-qualified) remote table name
      Returns:
      the table's statistics, or empty if unavailable
    • tables

      public Optional<List<String>> tables(ConnectionDeclaration connection)
      Returns the names of the tables connection holds, as far as this provider knows them, or Optional.empty() if it cannot enumerate them.

      The analyser uses the list to suggest a table when a reference names one the provider cannot describe, as it already suggests a column. The default returns empty, which suits live introspection: asking a database for every table to correct one typo is not a cost worth paying during analysis. A provider holding a fixed set of tables, such as a CatalogSnapshot, lists them.

      A snapshot knows exactly the tables it recorded for the connection, so it always enumerates them, in the order they were first recorded.

      Specified by:
      tables in interface CatalogProvider
      Parameters:
      connection - the declared connection
      Returns:
      the table names, as a reference would spell them after the connection's name; or empty if this provider cannot enumerate them
    • toJson

      public String toJson()
      This snapshot as JSON.

      Round-trips through parse(String). A type is written structurally rather than as its compact IR code, because a code needs a parser and a parser is a second grammar to keep in step with the first.

      Returns:
      the document
    • parse

      public static CatalogSnapshot parse(String json)
      Reads a snapshot written by toJson().
      Parameters:
      json - the document; must not be null
      Returns:
      the snapshot
      Throws:
      IllegalArgumentException - if the text is not a snapshot this version reads
    • toString

      public String toString()
      Overrides:
      toString in class Object