Live data from Hacker News

Voicebox: Generative AI model for speech that generalizes across tasks

ai.facebook.com

91–100 of 128 posts

Re: Voicebox: Generative AI model for speech that generalizes across tasks

#91
post #71

Earlier quoted context omitted.

And authoritarians everywhere will rejoice (and they will give out the means to duplicate these IDs to a select few in case they need to generate evidence that 'you' have offended the state).

The only sensible approach to this problem (assuming it is a real problem) is trusted individuals certifying others as human. There are alternatives to using the state for this, but they are difficult and fraught with UX issues. Perhaps a decentralised web of trust or some sort of blockchain based registrar of trust that can trace trust routes between mutually distrusting individuals. Unless such a system is in place…

I'll also add that some of the state proposals are in general quite nice in where their priorities are. Such as prioritizing working in the open, both in sharing research, working with standards groups, and open source tooling.

It would be nice if we could come to a thorough solution that actually does cover all bases, rather than all these companies trying to create their own digital ID services that just encourage us to instead do silly things like photograph your ID front and back.

I mean hell it's taken like 20 years for privacy by design to become an ISO standard? That sort of timeline is not something we can really tolerate as more and more people continue relying on online services and in turn wind up trusting horribly outdated techniques/general malaise about data.

Re: Voicebox: Generative AI model for speech that generalizes across tasks

#92
post #9

I am mostly excited for cheaper audiobooks with consistent voices for different characters.

Why would they make it cheaper when they can make even more profit by not having to pay a voice actor?

Get a text copy of the book and then pass it through the tool yourself, right? Who cares how much the publisher wants.

Re: Voicebox: Generative AI model for speech that generalizes across tasks

#93
post #24

> As with other powerful new AI innovations, we recognize that this technology brings the potential for misuse and unintended harm. In our paper, we detail how we built a highly effective classifier that can distinguish between authentic speech and audio generated with Voicebox to mitigate these possible future risks. We believe it is important to be open about our work so the research community can build on it and t…

Assuming 10 hours a piece, 6k books feels a very achievable dataset. Even Librivox claims 18k books (with many duplicates and hugely varying quality levels). If you wanted to get expansive, you could dig into the podcast archives of BBC, NPR, etc which could potentially yield millions of hours. [0] https://librivox.org/

older BBC material when the 'standard bbc' voice reigned supreme might make a great training set and could/should be publicly accessible?

Re: Voicebox: Generative AI model for speech that generalizes across tasks

#94
post #4

I think that the "star trek" use case of a live translation is super exciting. I think that this also will force people to have pass phrases that they use to authenticate phone calls with. I normally downplay when people bring up everyone signing everything with a public/private key (impractical for normal users) but clearly there will be a need for authentication protocols as AI proliferates

Coming to an android and iphone near you voice authentication. Audio streams have inaudible data produced by encrypting a mutually known but changeable token like the current time with your private key embedded in the stream for example in frequencies you can't hear. Your phone app queries the service with the time of call and the data received and if they are also a subscriber it is able to discern their identity wi…

> If this feature is standardized and built in it could be paired with a token like a yubikey which is on users keychain and authenticated even if they were using someone else's phone.

Remember the standard tech bait-and-switch: if this feature is built in, it'll not be good enough to function for purposes you describe, but it will be good enough to track you by your voice for advertising purposes.

Re: Voicebox: Generative AI model for speech that generalizes across tasks

#95
post #56

Earlier quoted context omitted.

This is a huge issue in the voice acting community. Been frequent recent discussions over at https://www.reddit.com/r/VoiceActing/ . For what its worth, most of the cost of audiobooks doesn't come from paying talent. For intermediate level actors, the going rate is around $50-$100 per finished hour (PFH) and experienced actors it can be around $250-$300. This page does a decent job of laying out pay structures for au…

There’s also a different possible take on this. Voice acting seems to be really bad career, so eliminating that job is desired, if you can deliver same quality/better product for cheaper to customers, without requiring employees to be underpaid. I know it sucks for people in that industry, but technical progress always eliminates jobs. Calculator used to be a job, now it’s a device.

Caveat though:

> if you can deliver same quality/better product for cheaper to customers, without requiring employees to be underpaid.

This almost never happens. Cheaper? Yes. Same or better quality? Not a chance. Automated solutions tend to allow reducing quality way below what humans workers would want to do, or even could cheaply and reliably (i.e. doing worse job than a careless one takes actual effort/skill). Like with every other case of automation replacing humans, expect the quality to be pushed down to minimum tolerable levels, as this is the point that maximizes revenue.

Re: Voicebox: Generative AI model for speech that generalizes across tasks

#96

Earlier quoted context omitted.

There’s also a different possible take on this. Voice acting seems to be really bad career, so eliminating that job is desired, if you can deliver same quality/better product for cheaper to customers, without requiring employees to be underpaid. I know it sucks for people in that industry, but technical progress always eliminates jobs. Calculator used to be a job, now it’s a device.

Caveat though: > if you can deliver same quality/better product for cheaper to customers, without requiring employees to be underpaid. This almost never happens. Cheaper? Yes. Same or better quality? Not a chance. Automated solutions tend to allow reducing quality way below what humans workers would want to do, or even could cheaply and reliably (i.e. doing worse job than a careless one takes actual effort/skill). Li…

Perhaps it's not the bar the GP was applying, but I think "good enough at 1/10 the price" is quite empowering for consumers. Consider all of the people that can't currently afford Audible, but who would like to listen to audio books while they commute to work, for example.

And of course, nothing stops you from paying what we pay now for human voice actors if there continues to be a quality differential that customers care about. (Though perhaps Baumol's Cost Disease would push the price up for today's human-generated quality.)

Extrapolating further -- if the commoditized version of audio books is AI generated voice, perhaps the new job for voice actors is human narrating/acting of AI-generated content for personalized stories ('Ractives from Stephenson's book "The Diamond Age"). Who knows, human voice actors could become more in demand, not less. To be clear I wouldn't forecast this as the most likely outcome, just pointing out that there are many possible outcomes.

Re: Voicebox: Generative AI model for speech that generalizes across tasks

#97
post #96

Earlier quoted context omitted.

Caveat though: > if you can deliver same quality/better product for cheaper to customers, without requiring employees to be underpaid. This almost never happens. Cheaper? Yes. Same or better quality? Not a chance. Automated solutions tend to allow reducing quality way below what humans workers would want to do, or even could cheaply and reliably (i.e. doing worse job than a careless one takes actual effort/skill). Li…

Perhaps it's not the bar the GP was applying, but I think "good enough at 1/10 the price" is quite empowering for consumers. Consider all of the people that can't currently afford Audible, but who would like to listen to audio books while they commute to work, for example. And of course, nothing stops you from paying what we pay now for human voice actors if there continues to be a quality differential that customers…

> I think "good enough at 1/10 the price" is quite empowering for consumers.

That's the thing though - it's not as empowering as it seems longer-term, because the "good enough" quickly drops to "barely fit for purpose"/"if it were any worse, it would be illegal to market or sell". This has been the case with most established classes of products I can think of, including pretty much anything that's been fully commoditized.

And so

> nothing stops you from paying what we pay now for human voice actors if there continues to be a quality differential that customers care about. (Though perhaps Baumol's Cost Disease would push the price up for today's human-generated quality.)

Nothing stops me today. But even if the quality differential exists, the dropping price on the low-quality version will reduce demand on the moderate-quality version, pushing its prices up and reducing number of suppliers (here, voice actors). The end result seems to always be a bifurcation: there is not enough demand to sustain a business doing decent quality work for a reasonable price, so all companies move to providing either low quality work cheaply, or high quality work at a hefty premium. The middle disappears.

In the specific context of this thread, the middle in question is the current quality of audiobooks with voice acting. The quality level available to most consumers will be below that, and the next step up will be niche recordings at high cost.

Re: Voicebox: Generative AI model for speech that generalizes across tasks

#98

That little stinger at the end was not as surprising as they thought it was :P It's very cool tech, but it's far from transparent. It has a very obvious "autotune" like sound to it that jumps right out. when they edited that one word it was obvious it had been edited. Again, super cool tech, just not going to replace voice actors or anything.

> It has a very obvious "autotune" To me it has a very obvious "Hindi is my native language" accent. I mean after literally the first sentence: "The research team at Meta is excited to share our work..." . Ouch. The "our work": just ouch. I was wondering why it wasn't a native english speaker presenting the video when the video is precisely about generating speech. The first seven seconds are particularly bad. Don't…

I'm surprised you had such a negative reaction to the Hindi accent! To me, it was no more difficult to understand than my colleagues who speak English as a second language.

To me, this is a style choice for the demo. Not evidence that they "fucked" it up. Accents are common - everyone has one! It's nice to see the model can support your personal voice even if it's not completely neutral English.

Re: Voicebox: Generative AI model for speech that generalizes across tasks

#99
post #9

I am mostly excited for cheaper audiobooks with consistent voices for different characters.

I've been watching the text-to-speech space for a while, waiting/hoping for something both open and better than CoquiTTS. ElevenLabs sounds amazing but is super expensive for something like a book, and tortoiseTTS is so slow as to be unusable.

I wrote a quick python script to read an ebook using coqui and the end result sounds pretty good. It's come in especially handy for books I want to listen to while doing yard work and stuff around the house.

https://github.com/aedocw/epub2tts

Re: Voicebox: Generative AI model for speech that generalizes across tasks

#100
post #98

Earlier quoted context omitted.

> It has a very obvious "autotune" To me it has a very obvious "Hindi is my native language" accent. I mean after literally the first sentence: "The research team at Meta is excited to share our work..." . Ouch. The "our work": just ouch. I was wondering why it wasn't a native english speaker presenting the video when the video is precisely about generating speech. The first seven seconds are particularly bad. Don't…

I'm surprised you had such a negative reaction to the Hindi accent! To me, it was no more difficult to understand than my colleagues who speak English as a second language. To me, this is a style choice for the demo. Not evidence that they "fucked" it up. Accents are common - everyone has one! It's nice to see the model can support your personal voice even if it's not completely neutral English.

> It's nice to see the model can support your personal voice even if it's not completely neutral English

There is no such thing as "neutral" English.

Post reply on HN