VAD (Voice Activity Detection)¶
Overview¶
Stateful per-chunk speech detection. Drives the gate between mic capture and Whisper / TTS, or slices a longer recording into speech regions in one batch call (see Segment Extraction under Usage Guides below).
VAD is reached through the singleton (JSON) path on both Kotlin and Flutter — the response shape is identical.
Supported Models¶
Model |
HF Repo |
Notes |
|---|---|---|
Silero VAD |
|
Stateful LSTM, 512-sample chunks @ 16 kHz |
VAD runs on ORT-CPU on Android (it is a small stateful LSTM that does not benefit from the Hexagon NPU).
API Reference¶
Init & Infer (single-chunk)¶
Kotlin:
// build.gradle.kts: implementation(files("libs/TheStageCore.aar"))
import ai.thestage.qlip.TheStageAI
// Inside a coroutine (initialize / start_model / infer are suspend).
TheStageAI.registerContext(context)
TheStageAI.initialize(api_token = "your-api-token")
TheStageAI.start_model(
model_name = "vad",
engines_path = "TheStageAI/silero-vad"
)
// Process audio in 512-sample chunks (32 ms @ 16 kHz).
val result = TheStageAI.infer(
model_name = "vad",
input_json = mapOf("audio" to audio_chunk) // FloatArray
)
val probability = result[0]["probability"] as Double
if (probability > 0.5) {
println("Speech detected!")
}
Flutter:
import 'package:thestage_android_sdk/thestage_android_sdk.dart';
import 'dart:typed_data';
await TheStageFlutterSDK.initialize(api_token: 'your-api-token');
await TheStageFlutterSDK.start_model(
model_name: 'vad',
engines_path: 'TheStageAI/silero-vad',
);
// audio_chunk: Float32List, 16 kHz mono, exactly 512 samples.
final result = await TheStageFlutterSDK.infer(
model_name: 'vad',
input_json: {'audio': audio_chunk},
);
final probability = result[0]['probability'] as double;
if (probability > 0.5) {
print('Speech detected!');
}
Inputs / Outputs (single-chunk mode)¶
Direction |
Type |
Description |
|---|---|---|
input |
|
16 kHz mono PCM, exactly 512 samples (32 ms). |
input |
|
Reset the LSTM state between independent utterances. |
output |
|
Speech probability in |
The single-chunk path returns just the probability — apply your own threshold and hysteresis.
Audio contract¶
16 kHz mono
FloatArray, samples in[-1.0, 1.0].Chunk size: exactly 512 samples per
infercall. Smaller chunks are zero-padded to 512 internally; larger chunks are rejected.Stateful. The model keeps an LSTM hidden state across calls. Pass
"reset_state": truebetween independent utterances (or callreset_state()on the directSileroVADAPI).Internal context. A 64-sample carry-over from the previous chunk is prepended automatically — you don’t need to overlap your capture yourself.
See TheStage Android SDK (Audio I/O Contract) for the shared format used across VAD / ASR / TTS.
Cleanup¶
Kotlin:
TheStageAI.stop_model(model_name = "vad")
Flutter:
await TheStageFlutterSDK.stop_model(model_name: 'vad');
Usage Guides¶
Real-Time Usage Pattern¶
Kotlin:
val threshold = 0.5
val speechBuffer = ArrayList<Float>()
for (chunk in microphoneStream) { // 512 samples each
val result = TheStageAI.infer(
model_name = "vad",
input_json = mapOf("audio" to chunk)
)
val probability = result[0]["probability"] as Double
if (probability > threshold) {
speechBuffer.addAll(chunk.toList())
} else if (speechBuffer.isNotEmpty()) {
// End of utterance — send to ASR.
val transcript = TheStageAI.infer(
model_name = "stt",
input_json = mapOf("audio" to speechBuffer.toFloatArray())
)
speechBuffer.clear()
}
}
Flutter:
const threshold = 0.5;
final speechBuffer = <double>[];
await for (final Float32List chunk in microphoneStream) { // 512 samples each
final result = await TheStageFlutterSDK.infer(
model_name: 'vad',
input_json: {'audio': chunk},
);
final probability = result[0]['probability'] as double;
if (probability > threshold) {
speechBuffer.addAll(chunk);
} else if (speechBuffer.isNotEmpty) {
final pcm = Float32List.fromList(speechBuffer);
await TheStageFlutterSDK.infer(
model_name: 'stt',
input_json: {'audio': pcm},
);
speechBuffer.clear();
}
}
For a production speech gate (onset/offset frames, pre-roll, max accumulation,
speculative ASR) use TheStageVoiceAgent — its VAD node already implements
all of this, including segment endpointing with hysteresis. See
Voice Agent.
Segment Extraction¶
For a whole recording, hand the entire buffer to infer with
extract_segments: true and read back the speech spans as sample indices —
the SDK runs the model over the buffer and applies onset/offset hysteresis +
padding for you.
Kotlin:
val segments = TheStageAI.infer(
model_name = "vad",
input_json = mapOf(
"audio" to recording, // FloatArray, 16 kHz mono
"extract_segments" to true,
"min_silence_duration_ms" to 100,
)
)
for (seg in segments) {
val start = seg["start"] as Int // sample indices
val end = seg["end"] as Int
val speech = recording.copyOfRange(start, end)
}
Flutter:
final segments = await TheStageFlutterSDK.infer(
model_name: 'vad',
input_json: {
'audio': recording, // Float32List, 16 kHz mono
'extract_segments': true,
},
);
for (final seg in segments) {
final start = seg['start'] as int; // sample indices
final end = seg['end'] as int;
}
Param |
Type |
Default |
Meaning |
|---|---|---|---|
|
|
|
Switch |
|
|
|
Onset probability threshold. |
|
|
auto |
Offset threshold; |
|
|
|
Drop spans shorter than this. |
|
|
|
Silence needed to end a span. |
|
|
|
Padding grown around each span. |
Output is one { "start": Int, "end": Int } per span (sample indices into
the input buffer). From Kotlin you can also call
SileroVAD.extract_segments(...) directly for a List<SpeechSegment>.
For a streaming, turn-level speech gate (onset/offset frames, pre-roll, max
accumulation, speculative ASR) use TheStageVoiceAgent instead — its VAD
node does live endpointing. WhisperPipeline also runs an internal Silero
pre-pass before transcribing.
Tuning the Threshold¶
The threshold decides how aggressively a chunk is classified as speech. In
single-chunk mode you apply it yourself on probability; in segment mode
pass it as the threshold param (the offset neg_threshold follows it
automatically unless you override it).
Higher (0.7–0.8): fewer false positives — triggers only on clear, confident speech. Good in noisy environments, but may miss quiet or distant speakers.
Lower (0.3–0.4): catches soft or distant voices, but also triggers more on ambient noise, keyboard clicks, or music.
Default (0.5): a balanced starting point for most environments.
Starting points by environment:
Environment |
Threshold |
|---|---|
Quiet office |
|
Noisy café / car |
|
Distant speakers (conference room) |
|
Push-to-talk (user intends to speak) |
|
The model is at its best on clean or moderate-noise audio. If a high-noise
environment still produces false triggers at 0.8, filter short bursts
with min_speech_duration_ms in segment mode, or add a minimum-duration
check before acting on single-chunk probabilities.
Resetting State Between Utterances¶
The LSTM hidden state carrying across infer calls is intentional — it
gives the model temporal context over a continuous stream. But when you
process independent audio (different files, a restarted recording
session, any gap in capture), pass "reset_state": true on the first
chunk of the new source. Otherwise state bleeds over from the previous
audio — typically inflated probabilities that read silence as speech.
Kotlin:
// Clip A — state builds up across its chunks.
for (chunk in clipA) {
TheStageAI.infer(
model_name = "vad",
input_json = mapOf("audio" to chunk)
)
}
// Clip B — reset on the first chunk only.
clipB.forEachIndexed { i, chunk ->
TheStageAI.infer(
model_name = "vad",
input_json = mapOf(
"audio" to chunk,
"reset_state" to (i == 0)
)
)
}
Flutter:
for (var i = 0; i < clipB.length; i++) {
await TheStageFlutterSDK.infer(
model_name: 'vad',
input_json: {
'audio': clipB[i],
'reset_state': i == 0,
},
);
}
The classic bug: skipping the reset makes the first few chunks of the new
clip report high probability even when they are silence. The mirror-image
caveat: right after a reset the model “warms up” — the first few 512-sample
chunks (~100 ms) can be less sensitive while the LSTM rebuilds context.
Lower your threshold for those first chunks, or rely on
TheStageVoiceAgent’s pre-roll so speech onsets are not clipped.
Troubleshooting¶
Symptom |
Cause / Fix |
|---|---|
|
Chunks larger than 512 samples are rejected — feed exactly 512 samples (32 ms @ 16 kHz). Smaller chunks are zero-padded to 512 internally. |
Unexpected probabilities on the first chunks of a new clip |
The LSTM hidden state carries across calls — pass
|
Flutter audio glitches / NaNs |
|