java.lang.Object
com.darkcollective.relix.value.internal.CodePoints

public final class CodePoints extends Object
Strings as sequences of code points — the unit relix counts, slices and orders a string by.

Why not just use String's own methods

Java's are written in char, which is a UTF-16 code unit: a character outside the basic multilingual plane is two of them. So "a😀b".length() is 4, substring can cut an emoji in half, and compareTo sorts every supplementary character before U+E000..U+FFFF because it compares the leading surrogate 0xD83D against a real character's code.

None of that is a decision relix made. It is an encoding detail of the JVM reaching the language surface — nobody asking how long a name is means "how many UTF-16 code units" — and it is the one thing that separated relix's answer from every SQL database's, all of which count characters and order by code point. Counting the same way is both the more defensible answer and what lets those functions be handed to a backend at all.

For a string of only basic-plane characters — which is most text, and all of ASCII — every method here agrees exactly with its String counterpart. The difference is confined to the strings the String version is wrong about.

  • Method Summary

    Modifier and Type
    Method
    Description
    static int
    Compares two strings by code point — the order every SQL binary collation uses, and the order UTF-8 bytes already sort in.
    static int
    first(String text)
    The first code point of text.
    static int
    indexOf(String text, String sought, int from)
    The position of sought in text, counted in code points from from, or -1 when it does not occur.
    static String
    last(String text, int count)
    The last count code points, or all of them when there are fewer.
    static int
    length(String text)
    The number of code points in text — its length as a user would count it.
    static String
    substring(String text, int start)
    The substring from start code points in, to the end.
    static String
    substring(String text, int start, int count)
    The substring of at most count code points, starting start code points in.

    Methods inherited from class java.lang.Object

    clone, equals, finalize, getClass, hashCode, notify, notifyAll, toString, wait, wait, wait
  • Method Details

    • length

      public static int length(String text)
      The number of code points in text — its length as a user would count it.
      Parameters:
      text - the string; must not be null
      Returns:
      the count
    • substring

      public static String substring(String text, int start)
      The substring from start code points in, to the end.
      Parameters:
      text - the string; must not be null
      start - the number of code points to skip; clamped to the string's length
      Returns:
      the remainder
    • substring

      public static String substring(String text, int start, int count)
      The substring of at most count code points, starting start code points in. Both bounds are clamped, so this never raises for an over-long request — the callers ask for "up to n" and mean it.
      Parameters:
      text - the string; must not be null
      start - the number of code points to skip; clamped to the string's length
      count - the maximum number of code points to take; clamped to what remains
      Returns:
      the slice
    • last

      public static String last(String text, int count)
      The last count code points, or all of them when there are fewer.
    • indexOf

      public static int indexOf(String text, String sought, int from)
      The position of sought in text, counted in code points from from, or -1 when it does not occur.
      Parameters:
      text - the string to search; must not be null
      sought - the string to find; must not be null
      from - the code-point position to start at
      Returns:
      the code-point index of the first occurrence, or -1
    • first

      public static int first(String text)
      The first code point of text.
      Parameters:
      text - a non-empty string; must not be null
      Returns:
      the code point
    • compare

      public static int compare(String a, String b)
      Compares two strings by code point — the order every SQL binary collation uses, and the order UTF-8 bytes already sort in.

      String.compareTo does not: it compares UTF-16 code units, so a supplementary character sorts before U+E000..U+FFFF rather than after it. That is the only case in which the two disagree.

      Parameters:
      a - the first string; must not be null
      b - the second; must not be null
      Returns:
      negative, zero or positive as a sorts before, with, or after b