Live data from Hacker News

Voicebox: Generative AI model for speech that generalizes across tasks

ai.facebook.com

41–50 of 128 posts

Re: Voicebox: Generative AI model for speech that generalizes across tasks

#41
post #9

I am mostly excited for cheaper audiobooks with consistent voices for different characters.

I was wondering if it would be possible to build something like this right now.

Use text to speech and chatgpt to tag the character text and timestamps.

Then use a speech to speech to change the character voices or even the whole reader.

But as a product I feel like theres some legal hurdles to figure out.

Re: Voicebox: Generative AI model for speech that generalizes across tasks

#42
post #5

Is Meta just operating in hope and AI moonshots now ? Everyone one of their products is just garbage to me and becoming less relevant by the day. When do they actually starting building something useful again ? Honestly Apple seems to be using “AI” much more successfully and actually seamlessly integrating it into their existing products to improve them. My theory is Mark is hoping the meta verse will pop out if Yan’…

I'm no fan of Meta but I'm glad there are companies investing in big moonshots like AI and yes, even Metaverse.

Re: Voicebox: Generative AI model for speech that generalizes across tasks

#43
post #4

I think that the "star trek" use case of a live translation is super exciting. I think that this also will force people to have pass phrases that they use to authenticate phone calls with. I normally downplay when people bring up everyone signing everything with a public/private key (impractical for normal users) but clearly there will be a need for authentication protocols as AI proliferates

How "live" can translation ever really be? Properly translating anything from one language to another requires context.

Re: Voicebox: Generative AI model for speech that generalizes across tasks

#44
post #24

> As with other powerful new AI innovations, we recognize that this technology brings the potential for misuse and unintended harm. In our paper, we detail how we built a highly effective classifier that can distinguish between authentic speech and audio generated with Voicebox to mitigate these possible future risks. We believe it is important to be open about our work so the research community can build on it and t…

Assuming 10 hours a piece, 6k books feels a very achievable dataset. Even Librivox claims 18k books (with many duplicates and hugely varying quality levels). If you wanted to get expansive, you could dig into the podcast archives of BBC, NPR, etc which could potentially yield millions of hours.

[0] https://librivox.org/

Re: Voicebox: Generative AI model for speech that generalizes across tasks

#46

That little stinger at the end was not as surprising as they thought it was :P It's very cool tech, but it's far from transparent. It has a very obvious "autotune" like sound to it that jumps right out. when they edited that one word it was obvious it had been edited. Again, super cool tech, just not going to replace voice actors or anything.

For video narration elocution, I'd say it was most of the way there.

When. Narrating. Videos. One. Tends. To. Speak. Differently.

Or, the more important case -- if I'm listening to audio-version-of-X, is it sufficiently human-like that I can forget that it's synthesized voice?

To me, yes.

Easy to tell if you're specifically listening for it, but to use an analogy one doesn't typically read novels and parse closely for grammar, does one? Your attention is elsewhere, on the content and plot.

Re: Voicebox: Generative AI model for speech that generalizes across tasks

#47

Is this really better than eleven labs?

Just listening to it, it's subjectively not better, but if it's > 10x faster/cheaper, I would use it anyway -- it's good enough to be listenable. Eleven Labs is the first voice synthesis that is good enough that I'd listen to an audiobook generated from it, but pricing is such that it would cost $100 to synthesize a 10 hour audiobook. A little too expensive. If they could get it down to $10 I'd cancel my Audible subs…

> So if I can get a locally running voicebox model and just leave it running on my laptop over night transcribing an audiobook, that's even better.

This is basically my dream for local AI... locals models trained on my own data/code/styles. Even if they're slow, as long as they work (V/RAM) and are of high enough quality then I'm happy to wait!

Re: Voicebox: Generative AI model for speech that generalizes across tasks

#48
post #9

I am mostly excited for cheaper audiobooks with consistent voices for different characters.

Why would they make it cheaper when they can make even more profit by not having to pay a voice actor?

Because it would be extremely easy to produce it cheaper or record it independently.

This opens up for non-signed authors to release audio books.

Re: Voicebox: Generative AI model for speech that generalizes across tasks

#49

Is this really better than eleven labs?

Just listening to it, it's subjectively not better, but if it's > 10x faster/cheaper, I would use it anyway -- it's good enough to be listenable. Eleven Labs is the first voice synthesis that is good enough that I'd listen to an audiobook generated from it, but pricing is such that it would cost $100 to synthesize a 10 hour audiobook. A little too expensive. If they could get it down to $10 I'd cancel my Audible subs…

Have you tried tortoisetts? I believe eleven labs basically forked that and made improvements on voice quality and speed there

Re: Voicebox: Generative AI model for speech that generalizes across tasks

#50
post #4

I think that the "star trek" use case of a live translation is super exciting. I think that this also will force people to have pass phrases that they use to authenticate phone calls with. I normally downplay when people bring up everyone signing everything with a public/private key (impractical for normal users) but clearly there will be a need for authentication protocols as AI proliferates

like... a pin?

In the US some phone companies have been using voice recognition to authenticate when their customers call. This will definitely have to see its end.
Post reply on HN