
Unicode character database access, case conversion, classification and script detection via Codepoint wrapper; generated lookup tables, surrogate handling, CharSequence and Appendable extensions.
Lightweight Unicode code-point handling for Kotlin Multiplatform – the Character API you've been missing in common code, without depending on ICU.
Kodepoint brings Unicode Character Database queries – case conversion, character classification, script and category lookup, and correct code-point iteration – to every Kotlin target through a single, allocation-free Codepoint value class.
val emoji = Codepoint(0x1F600) // 😀
emoji.isLetter() // false
emoji.getUnicodeScript() // COMMON
Codepoint('A'.code).toLowerCase().asString() // "a"Kotlin's common standard library has no equivalent of Java's Character. In shared code you can't ask whether a character is a letter, uppercase it, or safely walk a string by code point – and Kotlin's Char is only 16 bits, so emoji, CJK extensions, and other supplementary characters (anything above U+FFFF) silently break naive per-Char logic.
Kodepoint fills that gap:
commonMain.Codepoint is a @JvmInline value class wrapping a single Int. No boxing, no wrapper objects on the hot path.java.lang.Character across all 1,114,112 code points (see JVM Compatibility).java.lang.Character.Note: This project is a stopgap until KT-23251 (Extend Unicode support in Kotlin common) and KT-24908 (CodePoint inline class) land in the Kotlin standard library.
Reach for Kodepoint whenever you're in Kotlin Multiplatform / commonMain and need something the common standard library doesn't offer:
Codepoint(c).isLetter(), .isDigit(), .isUpperCase() – the java.lang.Character equivalents, but multiplatform.codepoint.toUpperCase() / .toLowerCase().text.forEachCodepoint { ... } (surrogate-safe, unlike for (c in text)).forEachCodepoint instead of .length.codepoint.getUnicodeScript() / .getCategory().StringBuilder?" → sb.appendCodePoint(codepoint)..isJavaIdentifierStart(), .isUnicodeIdentifierPart(), etc.Kodepoint covers code-point-level UCD properties. It is not an ICU replacement – for locale-aware collation, Unicode normalization (NFC/NFD), bidi, or grapheme-cluster segmentation, use a dedicated text library.
Kodepoint is published to Maven Central. Add it to your commonMain dependencies:
// build.gradle.kts
dependencies {
implementation("me.zolotov.kodepoint:kodepoint:4.0.0")
}import me.zolotov.kodepoint.*
// Create a code point from an Int...
val grinning = Codepoint(0x1F600) // 😀
// ...or from a Char.
val a = Codepoint('A'.code)
// Query Unicode properties.
a.isLetter() // true
a.isUpperCase() // true
a.toLowerCase() // Codepoint(0x61) -> "a"
a.getCategory() // Category.UPPERCASE_LETTER
grinning.getUnicodeScript() // UnicodeScript.COMMON
// Convert back to text.
grinning.asString() // "😀"Walking a string with for (c in text) splits emoji and other supplementary
characters into broken surrogate halves. forEachCodepoint gives you whole
characters:
val text = "Hi 👋🏽 世界"
text.forEachCodepoint { cp ->
println("U+%04X %s".format(cp.codepoint, cp.asString()))
}
// U+0048 H
// U+0069 i
// U+0020
// U+1F44B 👋
// U+1F3FD 🏽
// U+4E16 世
// U+754C 界// Count code points, not UTF-16 chars ("😀".length == 2).
var count = 0
"a😀b".forEachCodepoint { count++ } // 3
// Append a supplementary code point to any Appendable.
val sb = StringBuilder()
sb.appendCodePoint(Codepoint(0x1F44D)) // 👍
sb.appendCodePoint('!'.code)
sb.toString() // "👍!"A value class wrapping an Int code point.
| Member | Returns | Description |
|---|---|---|
Codepoint(Int) |
Codepoint |
Construct from a code-point value. |
Codepoint.fromChars(high, low) |
Codepoint |
Combine a high/low surrogate pair. |
codepoint |
Int |
The raw code-point value. |
charCount |
Int |
UTF-16 units needed (1 or 2). |
asString() |
String |
Encode as a String. |
Classification
| Method | Description |
|---|---|
isLetter() |
Letter (Lu, Ll, Lt, Lm, Lo). |
isDigit() |
Decimal digit (Nd). |
isLetterOrDigit() |
Letter or digit. |
isUpperCase() / isLowerCase()
|
Upper- / lowercase letter. |
isSpaceChar() |
Unicode space (Zs, Zl, Zp). |
isWhitespace() |
Whitespace (see differences). |
isIdeographic() |
Ideographic character. |
isISOControl() |
Control character (Cc). |
isIdentifierIgnorable() |
Ignorable in identifiers. |
isUnicodeIdentifierStart() / isUnicodeIdentifierPart()
|
Unicode identifier rules. |
isJavaIdentifierStart() / isJavaIdentifierPart()
|
Java identifier rules. |
Conversion & metadata
| Method | Returns | Description |
|---|---|---|
toUpperCase() / toLowerCase()
|
Codepoint |
Case mapping (unchanged if none). |
getCategory() |
Category |
General category (e.g. UPPERCASE_LETTER). |
getUnicodeScript() |
UnicodeScript |
Unicode script (e.g. LATIN, HAN). |
| Extension | Returns | Description |
|---|---|---|
forEachCodepoint { cp -> } |
Unit |
Iterate code points left-to-right (surrogate-safe, inline). |
forEachCodepointReversed { cp -> } |
Unit |
Iterate right-to-left. |
codePointAt(index) |
Codepoint |
Code point starting at index. |
codePointBefore(index) |
Codepoint |
Code point ending before index. |
codepoints(offset, direction) |
Iterator<Codepoint> |
Lazy iterator (Direction.FORWARD / BACKWARD). |
| Extension | Description |
|---|---|
appendCodePoint(Int) |
Append a code point (as its surrogate pair when needed). |
appendCodePoint(Codepoint) |
Same, taking a Codepoint. |
| Unicode version | 18.0.0 on non-JVM targets; the running JDK's on the JVM – see Unicode version |
| Kotlin API/language version | 2.1+ |
| Minimum consumer Kotlin | 2.1 on the JVM, 2.2 for all other targets (built with Kotlin 2.2.20) |
| JVM bytecode target | 11 |
| Correctness | Non-JVM output validated against java.lang.Character for all 1,114,112 code points, up to the documented Unicode version delta |
| Runtime dependencies | None |
| Platform | Targets |
|---|---|
| JVM | jvm |
| JavaScript |
js, wasmJs
|
| WASI | wasmWasi |
| iOS |
iosArm64, iosSimulatorArm64, iosX64
|
| macOS |
macosArm64, macosX64
|
| tvOS |
tvosArm64, tvosSimulatorArm64, tvosX64
|
| watchOS |
watchosArm64, watchosSimulatorArm64, watchosX64
|
| Linux |
linuxArm64, linuxX64
|
| Windows | mingwX64 |
Note: Only JVM and WasmJS targets are actively tested in CI. WasmWasi and the other targets compile and should work correctly, but have not been thoroughly validated.
Kodepoint aims for consistent Unicode behavior across all platforms. On the JVM it delegates to java.lang.Character; elsewhere it uses generated Unicode Character Database lookup tables. Non-JVM output is validated against the JVM implementation for all 1,114,112 code points, with the expected differences between the two Unicode versions checked exactly (see Unicode version).
These produce identical results to java.lang.Character for every code point assigned in the JDK's Unicode version:
isLetter(), isDigit(), isLetterOrDigit()
isUpperCase(), isLowerCase()
toLowerCase(), toUpperCase()
isSpaceChar()isIdeographic()isIdentifierIgnorable()isISOControl()isJavaIdentifierStart(), isJavaIdentifierPart() – derived from the Unicode general category following the java.lang.Character contractThe lookup tables follow Unicode 18.0.0, while java.lang.Character follows the JDK: JDK 24 and 25 implement Unicode 16.0.0. On the JVM Kodepoint returns the JDK's answers, so until a JDK ships Unicode 18 the two differ exactly by the 16.0 → 18.0 delta:
getCategory() = UNASSIGNED, getUnicodeScript() = UNKNOWN, isLetter() = false, …); the tables classify them.LOWERCASE_LETTER to OTHER_LETTER.toUpperCase()/toLowerCase() map them only on non-JVM targets.UnicodeScript gains BERIA_ERFE, JURCHEN, PROTO_CUNEIFORM, SEAL, SIDETIC, TAI_YO and TOLONG_SIKI. An exhaustive when over the enum needs new branches or an else.The complete per-function list is generated by ./gradlew :unicode:generateUcdDiff into unicode/build/generated/resources/ucd-diff/ucd-diff.txt; the JVM validation test requires the observed differences to match it exactly. Once a JDK implementing Unicode 18 is used, the delta is empty and the platforms agree again.
A few further functions intentionally differ from JVM behavior to follow the Unicode standard more closely.
This library uses Unicode's White_Space property, which differs from Character.isWhitespace():
| Codepoint | Character | Unicode White_Space | Java isWhitespace |
|---|---|---|---|
| U+001C | File Separator | false |
true |
| U+001D | Group Separator | false |
true |
| U+001E | Record Separator | false |
true |
| U+001F | Unit Separator | false |
true |
| U+0085 | Next Line (NEL) | true |
false |
| U+00A0 | No-Break Space | true |
false |
| U+2007 | Figure Space | true |
false |
| U+202F | Narrow No-Break Space | true |
false |
Java excludes non-breaking spaces from isWhitespace() and includes control characters that Unicode does not classify as whitespace.
JVM includes U+2E2F (VERTICAL TILDE) for backward compatibility, but this character is not in Unicode's ID_Start or ID_Continue properties. This library follows the Unicode standard.
Kodepoint is split into three modules:
lib (me.zolotov.kodepoint:kodepoint) – the public API: the Codepoint value class plus CharSequence/Appendable extensions, with platform-specific implementations.unicode – generated Unicode property lookup tables used by non-JVM targets.common – the UnicodeScript enum shared across modules.For a deep dive into data storage, lookup-table layout, and platform implementations, see ARCHITECTURE.md.
Benchmark results – including history and comparisons against java.lang.Character – are published to the dashboard. Commands, categories, and the reporting pipeline are documented in benchmarks/README.md.
# Build all modules
./gradlew build
# Run tests
./gradlew allTestsReleases and the changelog workflow are described in RELEASING.md.
Issues and pull requests are welcome. If you hit a code point that behaves differently from java.lang.Character (outside the documented differences), please open an issue.
Licensed under the Apache License 2.0.
Lightweight Unicode code-point handling for Kotlin Multiplatform – the Character API you've been missing in common code, without depending on ICU.
Kodepoint brings Unicode Character Database queries – case conversion, character classification, script and category lookup, and correct code-point iteration – to every Kotlin target through a single, allocation-free Codepoint value class.
val emoji = Codepoint(0x1F600) // 😀
emoji.isLetter() // false
emoji.getUnicodeScript() // COMMON
Codepoint('A'.code).toLowerCase().asString() // "a"Kotlin's common standard library has no equivalent of Java's Character. In shared code you can't ask whether a character is a letter, uppercase it, or safely walk a string by code point – and Kotlin's Char is only 16 bits, so emoji, CJK extensions, and other supplementary characters (anything above U+FFFF) silently break naive per-Char logic.
Kodepoint fills that gap:
commonMain.Codepoint is a @JvmInline value class wrapping a single Int. No boxing, no wrapper objects on the hot path.java.lang.Character across all 1,114,112 code points (see JVM Compatibility).java.lang.Character.Note: This project is a stopgap until KT-23251 (Extend Unicode support in Kotlin common) and KT-24908 (CodePoint inline class) land in the Kotlin standard library.
Reach for Kodepoint whenever you're in Kotlin Multiplatform / commonMain and need something the common standard library doesn't offer:
Codepoint(c).isLetter(), .isDigit(), .isUpperCase() – the java.lang.Character equivalents, but multiplatform.codepoint.toUpperCase() / .toLowerCase().text.forEachCodepoint { ... } (surrogate-safe, unlike for (c in text)).forEachCodepoint instead of .length.codepoint.getUnicodeScript() / .getCategory().StringBuilder?" → sb.appendCodePoint(codepoint)..isJavaIdentifierStart(), .isUnicodeIdentifierPart(), etc.Kodepoint covers code-point-level UCD properties. It is not an ICU replacement – for locale-aware collation, Unicode normalization (NFC/NFD), bidi, or grapheme-cluster segmentation, use a dedicated text library.
Kodepoint is published to Maven Central. Add it to your commonMain dependencies:
// build.gradle.kts
dependencies {
implementation("me.zolotov.kodepoint:kodepoint:4.0.0")
}import me.zolotov.kodepoint.*
// Create a code point from an Int...
val grinning = Codepoint(0x1F600) // 😀
// ...or from a Char.
val a = Codepoint('A'.code)
// Query Unicode properties.
a.isLetter() // true
a.isUpperCase() // true
a.toLowerCase() // Codepoint(0x61) -> "a"
a.getCategory() // Category.UPPERCASE_LETTER
grinning.getUnicodeScript() // UnicodeScript.COMMON
// Convert back to text.
grinning.asString() // "😀"Walking a string with for (c in text) splits emoji and other supplementary
characters into broken surrogate halves. forEachCodepoint gives you whole
characters:
val text = "Hi 👋🏽 世界"
text.forEachCodepoint { cp ->
println("U+%04X %s".format(cp.codepoint, cp.asString()))
}
// U+0048 H
// U+0069 i
// U+0020
// U+1F44B 👋
// U+1F3FD 🏽
// U+4E16 世
// U+754C 界// Count code points, not UTF-16 chars ("😀".length == 2).
var count = 0
"a😀b".forEachCodepoint { count++ } // 3
// Append a supplementary code point to any Appendable.
val sb = StringBuilder()
sb.appendCodePoint(Codepoint(0x1F44D)) // 👍
sb.appendCodePoint('!'.code)
sb.toString() // "👍!"A value class wrapping an Int code point.
| Member | Returns | Description |
|---|---|---|
Codepoint(Int) |
Codepoint |
Construct from a code-point value. |
Codepoint.fromChars(high, low) |
Codepoint |
Combine a high/low surrogate pair. |
codepoint |
Int |
The raw code-point value. |
charCount |
Int |
UTF-16 units needed (1 or 2). |
asString() |
String |
Encode as a String. |
Classification
| Method | Description |
|---|---|
isLetter() |
Letter (Lu, Ll, Lt, Lm, Lo). |
isDigit() |
Decimal digit (Nd). |
isLetterOrDigit() |
Letter or digit. |
isUpperCase() / isLowerCase()
|
Upper- / lowercase letter. |
isSpaceChar() |
Unicode space (Zs, Zl, Zp). |
isWhitespace() |
Whitespace (see differences). |
isIdeographic() |
Ideographic character. |
isISOControl() |
Control character (Cc). |
isIdentifierIgnorable() |
Ignorable in identifiers. |
isUnicodeIdentifierStart() / isUnicodeIdentifierPart()
|
Unicode identifier rules. |
isJavaIdentifierStart() / isJavaIdentifierPart()
|
Java identifier rules. |
Conversion & metadata
| Method | Returns | Description |
|---|---|---|
toUpperCase() / toLowerCase()
|
Codepoint |
Case mapping (unchanged if none). |
getCategory() |
Category |
General category (e.g. UPPERCASE_LETTER). |
getUnicodeScript() |
UnicodeScript |
Unicode script (e.g. LATIN, HAN). |
| Extension | Returns | Description |
|---|---|---|
forEachCodepoint { cp -> } |
Unit |
Iterate code points left-to-right (surrogate-safe, inline). |
forEachCodepointReversed { cp -> } |
Unit |
Iterate right-to-left. |
codePointAt(index) |
Codepoint |
Code point starting at index. |
codePointBefore(index) |
Codepoint |
Code point ending before index. |
codepoints(offset, direction) |
Iterator<Codepoint> |
Lazy iterator (Direction.FORWARD / BACKWARD). |
| Extension | Description |
|---|---|
appendCodePoint(Int) |
Append a code point (as its surrogate pair when needed). |
appendCodePoint(Codepoint) |
Same, taking a Codepoint. |
| Unicode version | 18.0.0 on non-JVM targets; the running JDK's on the JVM – see Unicode version |
| Kotlin API/language version | 2.1+ |
| Minimum consumer Kotlin | 2.1 on the JVM, 2.2 for all other targets (built with Kotlin 2.2.20) |
| JVM bytecode target | 11 |
| Correctness | Non-JVM output validated against java.lang.Character for all 1,114,112 code points, up to the documented Unicode version delta |
| Runtime dependencies | None |
| Platform | Targets |
|---|---|
| JVM | jvm |
| JavaScript |
js, wasmJs
|
| WASI | wasmWasi |
| iOS |
iosArm64, iosSimulatorArm64, iosX64
|
| macOS |
macosArm64, macosX64
|
| tvOS |
tvosArm64, tvosSimulatorArm64, tvosX64
|
| watchOS |
watchosArm64, watchosSimulatorArm64, watchosX64
|
| Linux |
linuxArm64, linuxX64
|
| Windows | mingwX64 |
Note: Only JVM and WasmJS targets are actively tested in CI. WasmWasi and the other targets compile and should work correctly, but have not been thoroughly validated.
Kodepoint aims for consistent Unicode behavior across all platforms. On the JVM it delegates to java.lang.Character; elsewhere it uses generated Unicode Character Database lookup tables. Non-JVM output is validated against the JVM implementation for all 1,114,112 code points, with the expected differences between the two Unicode versions checked exactly (see Unicode version).
These produce identical results to java.lang.Character for every code point assigned in the JDK's Unicode version:
isLetter(), isDigit(), isLetterOrDigit()
isUpperCase(), isLowerCase()
toLowerCase(), toUpperCase()
isSpaceChar()isIdeographic()isIdentifierIgnorable()isISOControl()isJavaIdentifierStart(), isJavaIdentifierPart() – derived from the Unicode general category following the java.lang.Character contractThe lookup tables follow Unicode 18.0.0, while java.lang.Character follows the JDK: JDK 24 and 25 implement Unicode 16.0.0. On the JVM Kodepoint returns the JDK's answers, so until a JDK ships Unicode 18 the two differ exactly by the 16.0 → 18.0 delta:
getCategory() = UNASSIGNED, getUnicodeScript() = UNKNOWN, isLetter() = false, …); the tables classify them.LOWERCASE_LETTER to OTHER_LETTER.toUpperCase()/toLowerCase() map them only on non-JVM targets.UnicodeScript gains BERIA_ERFE, JURCHEN, PROTO_CUNEIFORM, SEAL, SIDETIC, TAI_YO and TOLONG_SIKI. An exhaustive when over the enum needs new branches or an else.The complete per-function list is generated by ./gradlew :unicode:generateUcdDiff into unicode/build/generated/resources/ucd-diff/ucd-diff.txt; the JVM validation test requires the observed differences to match it exactly. Once a JDK implementing Unicode 18 is used, the delta is empty and the platforms agree again.
A few further functions intentionally differ from JVM behavior to follow the Unicode standard more closely.
This library uses Unicode's White_Space property, which differs from Character.isWhitespace():
| Codepoint | Character | Unicode White_Space | Java isWhitespace |
|---|---|---|---|
| U+001C | File Separator | false |
true |
| U+001D | Group Separator | false |
true |
| U+001E | Record Separator | false |
true |
| U+001F | Unit Separator | false |
true |
| U+0085 | Next Line (NEL) | true |
false |
| U+00A0 | No-Break Space | true |
false |
| U+2007 | Figure Space | true |
false |
| U+202F | Narrow No-Break Space | true |
false |
Java excludes non-breaking spaces from isWhitespace() and includes control characters that Unicode does not classify as whitespace.
JVM includes U+2E2F (VERTICAL TILDE) for backward compatibility, but this character is not in Unicode's ID_Start or ID_Continue properties. This library follows the Unicode standard.
Kodepoint is split into three modules:
lib (me.zolotov.kodepoint:kodepoint) – the public API: the Codepoint value class plus CharSequence/Appendable extensions, with platform-specific implementations.unicode – generated Unicode property lookup tables used by non-JVM targets.common – the UnicodeScript enum shared across modules.For a deep dive into data storage, lookup-table layout, and platform implementations, see ARCHITECTURE.md.
Benchmark results – including history and comparisons against java.lang.Character – are published to the dashboard. Commands, categories, and the reporting pipeline are documented in benchmarks/README.md.
# Build all modules
./gradlew build
# Run tests
./gradlew allTestsReleases and the changelog workflow are described in RELEASING.md.
Issues and pull requests are welcome. If you hit a code point that behaves differently from java.lang.Character (outside the documented differences), please open an issue.
Licensed under the Apache License 2.0.