
Streaming voice-activity detection with embedded Silero ONNX model; frame-based API accepts Short/Float/Byte PCM, keeps internal state, supports reset and simple integration.
Silero VAD (Voice Activity Detector) for Kotlin Multiplatform, powered by onnxruntime-kmp. The ONNX model is embedded into the library with the cn.rtast.kembeddable Gradle plugin.
| Platform | Backend |
|---|---|
| JVM (desktop) |
com.microsoft.onnxruntime:onnxruntime via onnxruntime-kmp |
| Android |
com.microsoft.onnxruntime:onnxruntime-android via onnxruntime-kmp |
| macOS arm64 / x64 | prebuilt libonnxruntime via onnxruntime-kmp cinterop |
| Linux x64 | prebuilt libonnxruntime via onnxruntime-kmp cinterop |
| Windows x64 | prebuilt onnxruntime.dll via onnxruntime-kmp cinterop |
// settings.gradle.kts
dependencyResolutionManagement {
repositories {
mavenCentral()
// kembeddable runtime artifacts
maven("https://repo.maven.rtast.cn/releases")
}
}
// module build.gradle.kts
kotlin {
sourceSets {
commonMain.dependencies {
implementation("cn.enaium:silero-vad-kmp:1.0.0")
}
}
}dependencies {
implementation("cn.enaium:silero-vad-kmp-jvm:1.0.0")
}dependencies {
implementation("cn.enaium:silero-vad-kmp-android:1.0.0")
}The model is embedded via the kembeddable plugin (no extra configuration needed).
The libonnxruntime shared library must be shipped next to the final binary (see
onnxruntime-kmp for the artifact/table):
| Platform | Runtime file to ship |
|---|---|
| macOS arm64 | libonnxruntime.1.dylib |
| macOS x64 | libonnxruntime.1.23.2.dylib |
| Linux x64 | libonnxruntime.so.1 |
| Windows x64 | onnxruntime.dll |
import cn.enaium.silero.vad.SileroVad
import cn.enaium.silero.vad.config.SampleRate
val vad = SileroVad(
sampleRate = SampleRate.SAMPLE_RATE_16K, // 8K or 16K
minSpeechDurationMs = 100,
minSilenceDurationMs = 100,
)
// 16 kHz -> 512 samples (32 ms) per frame, 8 kHz -> 256 samples.
val speech = vad.isSpeech(shortSamples) // ShortArray / FloatArray / ByteArray (16-bit LE PCM)
vad.close()isSpeech(ShortArray) / isSpeech(FloatArray) / isSpeech(ByteArray) (little-endian 16-bit PCM).reset() clears the internal state (call before a new utterance/stream).The examples/basic module shares a Compose Multiplatform UI;
only the microphone reader differs per platform:
javax.sound TargetDataLine: run with ./gradlew :examples:basic:run.AudioRecord: host the examples/basic AAR in an app (contains the MainActivity).AVFAudio AVAudioEngine: run with ./gradlew :examples:basic:runDebugExecutableMacosArm64
(console output; the shared Compose UI is JVM-only, so the native example is a CLI app).examples/waveform records the microphone and draws the raw signal and the
VAD-gated signal side by side with Dear ImGui
and ImPlot:
# JVM
./gradlew :examples:waveform:jvmRun
./gradlew :examples:waveform:jvmRun --args="--frames 300" # exit after 300 frames
# Native executable (macosArm64 shown; macosX64, linuxX64 and mingwX64 build the same way)
./gradlew :examples:waveform:runDebugExecutableMacosArm64
SILERO_VAD_KMP_FRAMES=300 ./examples/waveform/build/bin/linuxX64/releaseExecutable/waveform.kexeThe controls at the top are the transport and the settings SileroVad is
constructed with:
min speech at 0 ms reports speech on the first frame
above the threshold instead of waiting for that much of it; window
(100 ms - 10 s) sets how much of the recording the plots show and how much
Play plays.off - 1 s) replays the audio that was written as silence
into the gated output when the gate opens: the model only reports speech
after min speech of it, so the onset of an utterance is already gated off by
then. With 250 ms the 250 ms of raw audio that precede every detected
onset are copied over that silence, which brings the start of the utterance
back (the plot stays aligned, the frames are overwritten where they are).Both plots share one auto-scaled amplitude axis, so the detector's effect is directly visible: the output plot is the raw signal while the model reports speech and 0 everywhere else - silence is filled with zeros instead of a copy of the input. The readouts name the capture device, the model frame size (512 samples / 32 ms at 16 kHz), how much of the recording was speech and the peak level of each signal in dBFS.
The example builds for the desktop targets the library publishes (jvm,
macosArm64, macosX64, linuxX64, mingwX64) and for Android:
./gradlew :examples:waveform:android:assembleDebug
adb install -r examples/waveform/android/build/outputs/apk/debug/android-debug.apkOn Android everything runs through the JVM APIs of the dependencies - imgui-kmp
and sdl-kmp ship their per-ABI JNI libraries in their AARs, audio-io-kmp uses
AudioRecord/AudioTrack, silero-vad-kmp uses onnxruntime-android - so the
APK needs no NDK build. The activity extends SDL's own SDLActivity (from
sdl-kmp's sdl-kmp-android-jvm AAR), returns "sdl_jni" from
getLibraries() - the one shared object that carries SDL3, the JNI bridge and
SDL's Android Java layer - and replaces the entry point SDL would call in a
native SDL_main with the shared Kotlin runWaveformExample(), which leaves
SDL's surface, input and lifecycle handling in place. The example runs in
sensor landscape, hides the system bars, and scales the whole UI - font, style
and fixed sizes - by the display density (dpiScale=3.0 on a 480 dpi phone), so
the controls are touch sized; the APK ships arm64-v8a and x86_64
(defaultConfig.ndk.abiFilters, each ABI costs ~55 MB of JNI libraries).
macOS + JVM: SDL has to own the first thread, so an IDE run configuration needs
-XstartOnFirstThread --enable-native-access=ALL-UNNAMEDin its VM options.:examples:waveform:jvmRunand the native executables already run on the first thread; without the flag SDL cannot open a window and the example fails with that message instead of rendering invisibly.
Capture and playback go through audio-io-kmp
(JavaSound on the JVM, Core Audio / ALSA / WASAPI natively, AudioRecord/AudioTrack on Android)
and the UI is rendered through the SDL3 backends published by
imgui-kmp.
./gradlew publishToMavenLocal # requires onnxruntime-kmp 1.0.0 in ~/.m2
./gradlew jvmTest # JVM tests (embedded model over a 60 s test WAV)
./gradlew macosArm64Test # native test (copies libonnxruntime next to the binary)
./gradlew :examples:basic:run # JVM desktop example
./gradlew :examples:basic:runDebugExecutableMacosArm64 # macOS native example (AVAudioEngine)
./gradlew :examples:waveform:jvmRun # JVM waveform exampleMIT, see LICENSE. The embedded silero_vad.onnx model is from the
snakers4/silero-vad project (MIT).
Silero VAD (Voice Activity Detector) for Kotlin Multiplatform, powered by onnxruntime-kmp. The ONNX model is embedded into the library with the cn.rtast.kembeddable Gradle plugin.
| Platform | Backend |
|---|---|
| JVM (desktop) |
com.microsoft.onnxruntime:onnxruntime via onnxruntime-kmp |
| Android |
com.microsoft.onnxruntime:onnxruntime-android via onnxruntime-kmp |
| macOS arm64 / x64 | prebuilt libonnxruntime via onnxruntime-kmp cinterop |
| Linux x64 | prebuilt libonnxruntime via onnxruntime-kmp cinterop |
| Windows x64 | prebuilt onnxruntime.dll via onnxruntime-kmp cinterop |
// settings.gradle.kts
dependencyResolutionManagement {
repositories {
mavenCentral()
// kembeddable runtime artifacts
maven("https://repo.maven.rtast.cn/releases")
}
}
// module build.gradle.kts
kotlin {
sourceSets {
commonMain.dependencies {
implementation("cn.enaium:silero-vad-kmp:1.0.0")
}
}
}dependencies {
implementation("cn.enaium:silero-vad-kmp-jvm:1.0.0")
}dependencies {
implementation("cn.enaium:silero-vad-kmp-android:1.0.0")
}The model is embedded via the kembeddable plugin (no extra configuration needed).
The libonnxruntime shared library must be shipped next to the final binary (see
onnxruntime-kmp for the artifact/table):
| Platform | Runtime file to ship |
|---|---|
| macOS arm64 | libonnxruntime.1.dylib |
| macOS x64 | libonnxruntime.1.23.2.dylib |
| Linux x64 | libonnxruntime.so.1 |
| Windows x64 | onnxruntime.dll |
import cn.enaium.silero.vad.SileroVad
import cn.enaium.silero.vad.config.SampleRate
val vad = SileroVad(
sampleRate = SampleRate.SAMPLE_RATE_16K, // 8K or 16K
minSpeechDurationMs = 100,
minSilenceDurationMs = 100,
)
// 16 kHz -> 512 samples (32 ms) per frame, 8 kHz -> 256 samples.
val speech = vad.isSpeech(shortSamples) // ShortArray / FloatArray / ByteArray (16-bit LE PCM)
vad.close()isSpeech(ShortArray) / isSpeech(FloatArray) / isSpeech(ByteArray) (little-endian 16-bit PCM).reset() clears the internal state (call before a new utterance/stream).The examples/basic module shares a Compose Multiplatform UI;
only the microphone reader differs per platform:
javax.sound TargetDataLine: run with ./gradlew :examples:basic:run.AudioRecord: host the examples/basic AAR in an app (contains the MainActivity).AVFAudio AVAudioEngine: run with ./gradlew :examples:basic:runDebugExecutableMacosArm64
(console output; the shared Compose UI is JVM-only, so the native example is a CLI app).examples/waveform records the microphone and draws the raw signal and the
VAD-gated signal side by side with Dear ImGui
and ImPlot:
# JVM
./gradlew :examples:waveform:jvmRun
./gradlew :examples:waveform:jvmRun --args="--frames 300" # exit after 300 frames
# Native executable (macosArm64 shown; macosX64, linuxX64 and mingwX64 build the same way)
./gradlew :examples:waveform:runDebugExecutableMacosArm64
SILERO_VAD_KMP_FRAMES=300 ./examples/waveform/build/bin/linuxX64/releaseExecutable/waveform.kexeThe controls at the top are the transport and the settings SileroVad is
constructed with:
min speech at 0 ms reports speech on the first frame
above the threshold instead of waiting for that much of it; window
(100 ms - 10 s) sets how much of the recording the plots show and how much
Play plays.off - 1 s) replays the audio that was written as silence
into the gated output when the gate opens: the model only reports speech
after min speech of it, so the onset of an utterance is already gated off by
then. With 250 ms the 250 ms of raw audio that precede every detected
onset are copied over that silence, which brings the start of the utterance
back (the plot stays aligned, the frames are overwritten where they are).Both plots share one auto-scaled amplitude axis, so the detector's effect is directly visible: the output plot is the raw signal while the model reports speech and 0 everywhere else - silence is filled with zeros instead of a copy of the input. The readouts name the capture device, the model frame size (512 samples / 32 ms at 16 kHz), how much of the recording was speech and the peak level of each signal in dBFS.
The example builds for the desktop targets the library publishes (jvm,
macosArm64, macosX64, linuxX64, mingwX64) and for Android:
./gradlew :examples:waveform:android:assembleDebug
adb install -r examples/waveform/android/build/outputs/apk/debug/android-debug.apkOn Android everything runs through the JVM APIs of the dependencies - imgui-kmp
and sdl-kmp ship their per-ABI JNI libraries in their AARs, audio-io-kmp uses
AudioRecord/AudioTrack, silero-vad-kmp uses onnxruntime-android - so the
APK needs no NDK build. The activity extends SDL's own SDLActivity (from
sdl-kmp's sdl-kmp-android-jvm AAR), returns "sdl_jni" from
getLibraries() - the one shared object that carries SDL3, the JNI bridge and
SDL's Android Java layer - and replaces the entry point SDL would call in a
native SDL_main with the shared Kotlin runWaveformExample(), which leaves
SDL's surface, input and lifecycle handling in place. The example runs in
sensor landscape, hides the system bars, and scales the whole UI - font, style
and fixed sizes - by the display density (dpiScale=3.0 on a 480 dpi phone), so
the controls are touch sized; the APK ships arm64-v8a and x86_64
(defaultConfig.ndk.abiFilters, each ABI costs ~55 MB of JNI libraries).
macOS + JVM: SDL has to own the first thread, so an IDE run configuration needs
-XstartOnFirstThread --enable-native-access=ALL-UNNAMEDin its VM options.:examples:waveform:jvmRunand the native executables already run on the first thread; without the flag SDL cannot open a window and the example fails with that message instead of rendering invisibly.
Capture and playback go through audio-io-kmp
(JavaSound on the JVM, Core Audio / ALSA / WASAPI natively, AudioRecord/AudioTrack on Android)
and the UI is rendered through the SDL3 backends published by
imgui-kmp.
./gradlew publishToMavenLocal # requires onnxruntime-kmp 1.0.0 in ~/.m2
./gradlew jvmTest # JVM tests (embedded model over a 60 s test WAV)
./gradlew macosArm64Test # native test (copies libonnxruntime next to the binary)
./gradlew :examples:basic:run # JVM desktop example
./gradlew :examples:basic:runDebugExecutableMacosArm64 # macOS native example (AVAudioEngine)
./gradlew :examples:waveform:jvmRun # JVM waveform exampleMIT, see LICENSE. The embedded silero_vad.onnx model is from the
snakers4/silero-vad project (MIT).