Hacker Timesnew | past | comments | ask | show | jobs | submitlogin

Whisper small/tiny/base are almost four years old (they were not updated for Whisper v2 or v3). Is there really nothing better to benchmark against by now?


There are many [0], you can search and filter by streaming and open weight only as well.

Looks like Voxtral and Nvidia's Nemotron are best.

[0] https://artificialanalysis.ai/speech-to-text/non-streaming


There's tons, Parakeet was the last I remember seeing which seemed to gain traction (independent lightweight implementations etc).


I have tried everything (that will run on a 12GB RTX 4070) and I have yet to find anything with better accuracy than Whisper V2 Large for my dataset (discord audio from TTRPG sessions, isolated per-speaker, mostly non-American accents)


Same, for my English-only podcast


Voxtral to me what better


Not v3?


Nvidia's Nemotron subsumes their older Parakeet model now even for real time streaming.


Parakeet is way faster (on Nvidia hardware) but not quite as accurate in my experience.


It's also super fast on CPU.


Parakeet isn't as good as whisper large.




Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: