ten_vad
Source Location
How It Works
- Incoming LINEAR16 bytes are converted to
int16samples - Samples are processed in fixed 256-sample frames (16 ms each)
- Each frame produces a speech probability score
- Speech onset/offset is tracked with the same hysteresis logic as Silero — speech ends only when probability drops below
threshold - 0.15 - Same packet emission pattern:
InterruptionPacketon onset,VadSpeechActivityPacketheartbeats during speech
Parameters
Shared Library
TEN VAD does not use an ONNX model. It requires thelibten_vad shared library at both build and runtime.
Docker: The Dockerfile copies the library from the source tree: