java.lang.Object
com.darkcollective.relix.parser.Lexer

public final class Lexer extends Object
Tokenizes a relational algebra expression string into a stream of Tokens consumed by RelAlgebraParser.

The lexer recognises all Unicode RA operator symbols (π σ ρ γ τ λ δ ω ⋈ ⨝ ⟕ ⟖ ⟗ ⋉ ▷ ∪ ⊎ ⊔ − ∩ ÷ ∆ ∘ ∀ × ∧ ∨ ¬ → ⊥ ∈ ∉ ≠ ≤ ≥ ⁺) and their ASCII keyword equivalents (PROJECT, SELECT, RENAME, …) as defined by TokenType. The superscript-plus ⁺ is the Kleene-plus postfix glyph (transitive closure); the plain * stays MULTIPLY and is interpreted as the Kleene-star postfix glyph contextually by the parser when it follows a relation expression. Single-quoted and double-quoted string literals are accepted; escape sequences \\, \", \n, \t, \r are decoded inside string literals.

Comments are SQL's: -- to the end of the line, and /* */ for a block. They are the forms the script grammar uses, which matters because a .relix script hands an embedded expression body to this lexer verbatim: a comment written one way outside a brace and another way inside it would be a rule about braces rather than about comments. // is therefore not a comment but two division operators, and reads as a syntax error wherever it appears.

Line and column tracking is 1-based. The optional startLine and startColumn constructor parameters allow the lexer to continue a line/column count from a mid-file offset, which is needed when the language parser delegates an embedded RA body to this lexer.

The lexer is not thread-safe and is intended for single-use: construct, call next() repeatedly until TokenType.EOF, then discard.

  • Constructor Details

  • Method Details

    • reservedWords

      public static Set<String> reservedWords()
      The reserved words that the lexer maps to a non-identifier token — i.e. every word that cannot be used as a bare name and must be backtick-delimited (or is auto-delimited by a pretty-printer) to appear in name position. Case-insensitive lowercase forms. Exposed so pretty-printers and tooling can agree with the lexer on what needs delimiting without duplicating the map.
      Returns:
      an unmodifiable view of the reserved keyword set (lowercase)
    • input

      public String input()
    • next

      public Token next()