Live data from Hacker News

Meta AI announces Massive Multilingual Speech code, models for 1000+ languages

github.com

31–40 of 232 posts

Re: Meta AI announces Massive Multilingual Speech code, models for 1000+ languages

#31
post #23

Earlier quoted context omitted.

Based on the availability of STT, TTS, and translation models available to download in that github repo, a real-life babelfish is 'only' some glue code away. Wild times we're living in.

Nah, full-on babelfish is simply not possible. The meaning of the beginning of a sentence can be modified retroactively by the end of the sentence. This means that the Babelfish must either be greatly delayed or awkwardly correct itself every once in a while.

That’s why the babelfish translates brainwaves rather than sounds, which is especially important for communicating with our nonverbal alien neighbors.

Re: Meta AI announces Massive Multilingual Speech code, models for 1000+ languages

#34

Earlier quoted context omitted.

according to their blog post[1], MMS achieves ~half the error rate on words, while supporting 11x more languages. pretty impressive. [1] https://ai.facebook.com/blog/multilingual-model-speech-recog...

I wonder what the performance is on English specifically. Edit: Just checked the paper, it seems to be worse[1][2] but feel free to correct me. I feel like they should've just taken the Whipser architecture, scaled it, and scaled the dataset as they did. [1] Page: https://i.imgur.com/bq15Tno.png [2] Paper: https://scontent.fcai19-5.fna.fbcdn.net/v/t39.8562-6/3488279...

It's worse on English and a lot of other common languages (see Appendix C of the paper). It does better on less common languages like Latvian or Tajik, though.

Re: Meta AI announces Massive Multilingual Speech code, models for 1000+ languages

#35
post #23

Earlier quoted context omitted.

Based on the availability of STT, TTS, and translation models available to download in that github repo, a real-life babelfish is 'only' some glue code away. Wild times we're living in.

Nah, full-on babelfish is simply not possible. The meaning of the beginning of a sentence can be modified retroactively by the end of the sentence. This means that the Babelfish must either be greatly delayed or awkwardly correct itself every once in a while.

Close enough for practical value. Yes, big downside. But personally I’d use it.

Re: Meta AI announces Massive Multilingual Speech code, models for 1000+ languages

#37

meta is doing more for open ai than openai

For OpenAI, AI is the revenue source. For Meta, it's a tool used to build products.

If OpenAI gives their stuff away, they lose customers. If Meta does it, they can have community around it, and have joint effort about improving tools that then they'll use for their internal products.

OpenAI is modern (and most likely - very short lived) Microsoft in the AI space, while Meta tries to replicate Linux in the AI space.

Re: Meta AI announces Massive Multilingual Speech code, models for 1000+ languages

#38

Does it do American Sign Language, US fifth largest language? I didn’t think so.

American Sign Language, ISO-693-3 code 'ase', does not have a formal written or spoken grammar. Obviously, the list of MMS supported languages[0] does not include support for 'ase' because MMS is a tool for spoken language and, as an intermediate state, written language.

[0] https://dl.fbaipublicfiles.com/mms/misc/language_coverage_mm...

Re: Meta AI announces Massive Multilingual Speech code, models for 1000+ languages

#39
post #36

The problem with all these model releases is they have no demos or even video of it working. It’s all just download it and run it, like it’s an app.

I'd argue that having a "download and run" approach is so much better than videos or demos. Why do you think this is a problem?
Post reply on HN