Earlier quoted context omitted.
Also, this test is English-only, while a strong point of other models is to understand different languages without first having to say which one (so you don't need 3 different keyboard shortcuts if you wanna dictate in 3 languages day-to-day)
Does anyone have any experience with Mandarin STT? What's a good model for this? The use-case I have is subtitling of Mandarin speech.
[0] https://huggingface.co/OpenMOSS-Team/MOSS-Transcribe-Diarize
[1] https://huggingface.co/spaces/OpenMOSS-Team/MOSS-transcribe-...