Class CodePoints
Why not just use String's own methods
Java's are written in char, which is a UTF-16 code unit: a character
outside the basic multilingual plane is two of them. So "a😀b".length() is 4,
substring can cut an emoji in half, and compareTo sorts every
supplementary character before U+E000..U+FFFF because it compares the
leading surrogate 0xD83D against a real character's code.
None of that is a decision relix made. It is an encoding detail of the JVM reaching the language surface — nobody asking how long a name is means "how many UTF-16 code units" — and it is the one thing that separated relix's answer from every SQL database's, all of which count characters and order by code point. Counting the same way is both the more defensible answer and what lets those functions be handed to a backend at all.
For a string of only basic-plane characters — which is most text, and all of ASCII —
every method here agrees exactly with its String counterpart. The difference is
confined to the strings the String version is wrong about.
-
Method Summary
Modifier and TypeMethodDescriptionstatic intCompares two strings by code point — the order every SQL binary collation uses, and the order UTF-8 bytes already sort in.static intThe first code point oftext.static intThe position ofsoughtintext, counted in code points fromfrom, or-1when it does not occur.static StringThe lastcountcode points, or all of them when there are fewer.static intThe number of code points intext— its length as a user would count it.static StringThe substring fromstartcode points in, to the end.static StringThe substring of at mostcountcode points, startingstartcode points in.
-
Method Details
-
length
The number of code points intext— its length as a user would count it.- Parameters:
text- the string; must not be null- Returns:
- the count
-
substring
The substring fromstartcode points in, to the end.- Parameters:
text- the string; must not be nullstart- the number of code points to skip; clamped to the string's length- Returns:
- the remainder
-
substring
The substring of at mostcountcode points, startingstartcode points in. Both bounds are clamped, so this never raises for an over-long request — the callers ask for "up to n" and mean it.- Parameters:
text- the string; must not be nullstart- the number of code points to skip; clamped to the string's lengthcount- the maximum number of code points to take; clamped to what remains- Returns:
- the slice
-
last
The lastcountcode points, or all of them when there are fewer. -
indexOf
The position ofsoughtintext, counted in code points fromfrom, or-1when it does not occur.- Parameters:
text- the string to search; must not be nullsought- the string to find; must not be nullfrom- the code-point position to start at- Returns:
- the code-point index of the first occurrence, or -1
-
first
The first code point oftext.- Parameters:
text- a non-empty string; must not be null- Returns:
- the code point
-
compare
Compares two strings by code point — the order every SQL binary collation uses, and the order UTF-8 bytes already sort in.String.compareTodoes not: it compares UTF-16 code units, so a supplementary character sorts beforeU+E000..U+FFFFrather than after it. That is the only case in which the two disagree.- Parameters:
a- the first string; must not be nullb- the second; must not be null- Returns:
- negative, zero or positive as
asorts before, with, or afterb
-