FunASR
FunASR is a fundamental end-to-end speech recognition toolkit. It offers industrial-grade speech recognition, being 170x faster than Whisper, supporting over 50 languages, and integrating features like speaker diarization, emotion detection, and streaming.
FunASR is currently grouped under Voice / Speech, which makes it easier to evaluate through workflow fit instead of isolated features alone. Based on the available data, it leans most heavily toward Extremely fast (170x faster than Whisper) and Meeting transcription with speaker labels, timestamps, and punctuation. The listed license is MIT, which is useful when adoption constraints matter. It also shows measurable community traction with 19.3k GitHub stars.
Features
Why choose it
Trade-offs
Compatibility
Quick start
Use cases
How it compares
Alternatives
Related searches
Comments
- ?usr_seed_0574May 3, 2026
Language coverage beyond English is a meaningful differentiator.
- ?usr_seed_0101Apr 3, 2026
Active development from Alibaba's speech research team, keeps improving.
- ?usr_seed_0992Mar 26, 2026
170x realtime speech recognition across 50+ languages is genuinely industrial-grade.
- ?usr_seed_0361Mar 24, 2026
Good for teams building speech-enabled AI applications that need production ASR.