ManySpeech.AliFsmnVad
16k Universal VAD Model: Can be used to detect the start and end time points of effective speech in long speech segments FSMN Monochrome VAD is an efficient speech endpoint detection model proposed by the Speech Team of Damo Institute. It is used to detect the start and end time information of valid speech in input audio, and input the detected valid audio segments into the recognition engine for recognition, reducing recognition errors caused by invalid speech.
Activity
- Latest release
- 11mo ago
- Total releases
- 11
- Cadence
- ~daily
- Last 12 months
- 1
Details
- First release
- May 13, 2025
Releases
| Version | Released | |
|---|---|---|
1.1.4
patch
| ||
1.1.3
patch
| ||
1.1.2
patch
| ||
1.1.1
patch
| ||
1.1.0
minor
| ||
1.0.9
patch
| ||
1.0.8
patch
| ||
1.0.7
patch
| ||
1.0.6
patch
| ||
1.0.5
patch
| ||
1.0.4
initial
|