Live data from Hacker News

High-fidelity simultaneous speech-to-speech translation

arxiv.org

41–50 of 59 posts

Re: High-fidelity simultaneous speech-to-speech translation

#41
post #38

This is so cool. The future is cool! I wonder how it will work on languages that have different grammatical structure than french/english? Like Finno-Ugric languages which have sort of a Yoda speech to them. Edit: In Finno-Ugric languages words later on in a sentence can completely change the meaning. Will be interesting to look at. It's considerate of them to name it after my favourite whisky.

If Finnish is not widely known, German is more familiar, and there you can put the "nicht" at the very end of a sentence, reversing its meaning. Also, the verb may come close to the end, after an extended description of the subject / object; in English, you want the verb early. Human translators somehow handle that; machines would likely exhibit a similar delay.

Vaguely related anecdote: have you ever dictated a number to a French speaker? When you say “forty-two” or “seventy-six”, an English speaker will start writing the 4 or the 7 the moment they hear the “forty” or the “seventy”. The French speaker will also write the 4 the moment they hear the “quarante” in “quarante-deux” (40+2), but when you say “soixante-seize” (60+16), they will (without thinking about it!) only start writing 76 at the end of the whole thing, because after only hearing the “soixante” they can’t tell if they’ll need to write a 6 or a 7.

Re: High-fidelity simultaneous speech-to-speech translation

#43
Soniox also supports real-time speech-to-text translation with 60 languages. You can hook that to a TTS and you have Speech-to-Speech translation. That failed Google I/O real-time translation demo? With Soniox it just works.

You can try it out here (select translation instead of transcription) https://soniox.com/

Disclaimer: I work at Soniox.

Re: High-fidelity simultaneous speech-to-speech translation

#44

Earlier quoted context omitted.

I don't know if you're multilingual, but some concepts are just legitimately easier to express in some languages; and the different grammatical structures that languages have can be useful for emphasising certain things, or to express subtle relationships between concepts. I'm not a particularly fluent speaker of Japanese and Russian, but I still find it helpful to drop into them sometimes when speaking with someone…

I have to second this. I study Japanese myself and the entire way the Japanese communicate is reflected so deeply in the language. There is so so much nuance to pretty much every sentence they speak and there are certain grammar points that carry more meaning in three syllables than what can be expressed in English or German in a full sentence. And ok turn this way of communicating shapes their culture too I believe.…

I’ve tried to learn Mandarin and failed because of lack of memory and practice. mostly i’m shocked at how ambiguous it appears to an english-trained mind - you have to fill in a lot of fine article/pronoun detail from custom and common understanding. which is why i think a lot of automatic translations are poor.

Re: High-fidelity simultaneous speech-to-speech translation

#45

This is so cool. The future is cool! I wonder how it will work on languages that have different grammatical structure than french/english? Like Finno-Ugric languages which have sort of a Yoda speech to them. Edit: In Finno-Ugric languages words later on in a sentence can completely change the meaning. Will be interesting to look at. It's considerate of them to name it after my favourite whisky.

even in regular languages with similar structure, sometimes the ending of a sentence forces you to change how you would say the whole sentence. Human synchronous translators usually correct themselves in such cases, which is a trade-off of having better latency in most cases, at the cost of having to correct yourself once in a while.

Re: High-fidelity simultaneous speech-to-speech translation

#46
post #24

Earlier quoted context omitted.

Translators sure, interpreters no. Interpreters also have to factor in cultural context and customs, ensuring that meaning is conveyed without offence being given in formal contexts.

I don't see why software couldn't do that, if you give them the context.

The end-user is unlikely to know which part of the context is relevant, and it may also change from moment to moment depending on who is speaking to whom. Of course you could imagine an AI interpreter that has cameras for situational awareness and asks for clarification if anything important is unclear while smoothing over minor stuff without interrupting, but you could equally easily imagine an AGI, so it's not clear that this could be built to a reasonable quality standard with current technology.

Re: High-fidelity simultaneous speech-to-speech translation

#47

For anyone else looking for examples: https://huggingface.co/spaces/kyutai/hibiki-samples

The high fidelity examples (see CFG-10 in the page) where the translated version has a very heavy French accent is kind of impressive (not that it is really useful, but impressive indeed).

Re: High-fidelity simultaneous speech-to-speech translation

#48
post #38

Earlier quoted context omitted.

If Finnish is not widely known, German is more familiar, and there you can put the "nicht" at the very end of a sentence, reversing its meaning. Also, the verb may come close to the end, after an extended description of the subject / object; in English, you want the verb early. Human translators somehow handle that; machines would likely exhibit a similar delay.

Vaguely related anecdote: have you ever dictated a number to a French speaker? When you say “forty-two” or “seventy-six”, an English speaker will start writing the 4 or the 7 the moment they hear the “forty” or the “seventy”. The French speaker will also write the 4 the moment they hear the “quarante” in “quarante-deux” (40+2), but when you say “soixante-seize” (60+16), they will (without thinking about it!) only sta…

Belgian have figured this correctly

Re: High-fidelity simultaneous speech-to-speech translation

#49

Earlier quoted context omitted.

Translators sure, interpreters no. Interpreters also have to factor in cultural context and customs, ensuring that meaning is conveyed without offence being given in formal contexts.

That seems like something LLMs could eventually get good at

They'll just push everyone to use corporate wooden language and then they won't have to worry about tone and implied meanings :)

Re: High-fidelity simultaneous speech-to-speech translation

#50
post #37

Earlier quoted context omitted.

I think it'll greatly increase cultural learning, by increasing the opportunity to interact with people. I've traveled to a lot of countries, and never learned more than a handful of words in each, primarily related to basic service interactions. I enjoyed talking to locals when they spoke English. I couldn't interact in any meaningful way with the vast majority of people, though. Learning languages is great. If you…

Thanks. Wonderful take and optimistic. You are correct I think.

He's not, because those locals will stop being able to speak English in a few generations. Either you'll have battery and signal or you'll point at things and make monkey noises.
Post reply on HN