Live data from Hacker News

High-fidelity simultaneous speech-to-speech translation

arxiv.org

51–59 of 59 posts

Re: High-fidelity simultaneous speech-to-speech translation

#52
post #15

All these Japanese project names and no Japanese support (ToT)

I wonder why it's so popular to use Japanese words for random software projects. Bonus points if the project's application of the loanword is off-target from the word's usual meaning/usage, or if it's completely unrelated to the project.

Re: High-fidelity simultaneous speech-to-speech translation

#53

This is why I wonder about the value of language learning for reasons other than “I’m really passionate about it.” We are so close to interfaces that reduce the language barrier by a lot…

Well, the value is obvious for romantically involved people not sharing a language, when the batteries run empty. :)

Re: High-fidelity simultaneous speech-to-speech translation

#54
post #22

This is why I wonder about the value of language learning for reasons other than “I’m really passionate about it.” We are so close to interfaces that reduce the language barrier by a lot…

It's not personal but I can't help myself to think that's such a sad post here. Reducing learning a different culture through language by plugging in an earbud. Is the battery is gone or your phone is stolen you realize you can't automate anything and that you've learned nothing. It's not about the tech if it works it's amazing it's like babelfish but it's so shallow to assume everything has some direct and simple "v…

I think you’re reading a sense of cultural reductionism in my comment that I didn’t intend.

There’s more to learning a culture than the language. And having a real-time translator makes it possible to enjoy a huge range of cultures much more directly than before. The fact is, I’m not going to learn Chinese and Swahili and Japanese. So my choices are to go through a human translator or nothing if I want to talk to those people.

How is it sad that a technology is going to allow me to directly talk to a huge number of people that I never could have before?

Re: High-fidelity simultaneous speech-to-speech translation

#55

Earlier quoted context omitted.

Change is hard, but diversity is good, and certainly better than monoculture (of language).

What's the point of diversity if people can't communicate with each other, or if only educated elites within each subculture can do so? Diversity should bring different people together, not divide them artificially.

The irony here, is that diversity is actually extremely aligned with conservative values: freedom of expression and the ability to do what one wants (without regard to others).

Freedom of diversity allows for the flourishing of unique ideas and perspectives, which in turn, has many benefits, in terms of the creation of new value in unexpected ways. Diversity, in a sense, can be a synonym for independence and freedom.

Re: High-fidelity simultaneous speech-to-speech translation

#57
post #38

This is so cool. The future is cool! I wonder how it will work on languages that have different grammatical structure than french/english? Like Finno-Ugric languages which have sort of a Yoda speech to them. Edit: In Finno-Ugric languages words later on in a sentence can completely change the meaning. Will be interesting to look at. It's considerate of them to name it after my favourite whisky.

If Finnish is not widely known, German is more familiar, and there you can put the "nicht" at the very end of a sentence, reversing its meaning. Also, the verb may come close to the end, after an extended description of the subject / object; in English, you want the verb early. Human translators somehow handle that; machines would likely exhibit a similar delay.

>If Finnish is not widely known, German is more familiar, and there you can put the "nicht" at the very end of a sentence,

I've never heard of an english speaker doing that.. .. .. NOT!

Re: High-fidelity simultaneous speech-to-speech translation

#58
post #43

Soniox also supports real-time speech-to-text translation with 60 languages. You can hook that to a TTS and you have Speech-to-Speech translation. That failed Google I/O real-time translation demo? With Soniox it just works. You can try it out here (select translation instead of transcription) https://soniox.com/ Disclaimer: I work at Soniox.

I didn’t see any mention of running a model locally on device like the Hibiki abstract states. Is this available?

Re: High-fidelity simultaneous speech-to-speech translation

#59
post #50
post #37

Earlier quoted context omitted.

Thanks. Wonderful take and optimistic. You are correct I think.

He's not, because those locals will stop being able to speak English in a few generations. Either you'll have battery and signal or you'll point at things and make monkey noises.

What do we care what happens in a few generations? We'll all be dead, and the people alive will probably have universal translators implanted in their brains at birth. We absolutely won't need a "signal" to translate on a device anymore (that'll happen in just a few years, forget about generations), and there won't be anyplace on the entire planet that doesn't have network connectivity (that will also happen in just a few years; it's already reality with Starlink cellular).
Post reply on HN