Live data from Hacker News

Voicebox: Generative AI model for speech that generalizes across tasks

ai.facebook.com

31–40 of 128 posts

Re: Voicebox: Generative AI model for speech that generalizes across tasks

#31
post #10

And the source code for this one has.... ...not been open sourced and cannot be found. Sorry AI bros. Better read the paper this time.

> There are many exciting use cases for generative speech models, but because of the risks of misuse, we are not making the Voicebox model or code publicly available at this time. While we believe it is important to be open with the AI community and to share our research to advance the state of the art in AI, it’s also necessary to strike the right balance between openness with responsibility.

Yeah... I mean there definitely are ways to misuse this (especially the style transfer!) I don't think you're going to do anything except delay the inevitable Facebook.

Re: Voicebox: Generative AI model for speech that generalizes across tasks

#32
Anyone know what tool is being used to create AI singing voices/renditions of Mariah Carey singing Whitney Houston to other popular songs?

Here's a bunch of results on YouTube and some are really good

https://www.youtube.com/results?search_query=mariah+carey+ai...

Re: Voicebox: Generative AI model for speech that generalizes across tasks

#35
post #9

I am mostly excited for cheaper audiobooks with consistent voices for different characters.

Why would they make it cheaper when they can make even more profit by not having to pay a voice actor?

If they make it better for the reader, they can potentially raise the price. If they can make it cheaper to produce, they can potentially increase their profit without raising the price.

Usually on balance this falls somewhere in between -- more value for less money for the consumer, and more profit on each marginal unit of production for the producer, which is how technology progresses across most consumer goods.

Re: Voicebox: Generative AI model for speech that generalizes across tasks

#37
post #27

That little stinger at the end was not as surprising as they thought it was :P It's very cool tech, but it's far from transparent. It has a very obvious "autotune" like sound to it that jumps right out. when they edited that one word it was obvious it had been edited. Again, super cool tech, just not going to replace voice actors or anything.

For me its more like I wake up and check if humans have been replaced yet. Oh good, it's another day that I don't have to share one time pads with my mother to ensure that I'm talking to her and not a simulant performing fraud on a massive scale.

[flagged]

Re: Voicebox: Generative AI model for speech that generalizes across tasks

#38

Is this really better than eleven labs?

Just listening to it, it's subjectively not better, but if it's > 10x faster/cheaper, I would use it anyway -- it's good enough to be listenable.

Eleven Labs is the first voice synthesis that is good enough that I'd listen to an audiobook generated from it, but pricing is such that it would cost $100 to synthesize a 10 hour audiobook. A little too expensive. If they could get it down to $10 I'd cancel my Audible subscription and just synthesize audio from ebook text.

So if I can get a locally running voicebox model and just leave it running on my laptop over night transcribing an audiobook, that's even better.

Re: Voicebox: Generative AI model for speech that generalizes across tasks

#39
post #27

Earlier quoted context omitted.

For me its more like I wake up and check if humans have been replaced yet. Oh good, it's another day that I don't have to share one time pads with my mother to ensure that I'm talking to her and not a simulant performing fraud on a massive scale.

[flagged]

[flagged]

Re: Voicebox: Generative AI model for speech that generalizes across tasks

#40
post #10

And the source code for this one has.... ...not been open sourced and cannot be found. Sorry AI bros. Better read the paper this time.

> There are many exciting use cases for generative speech models, but because of the risks of misuse, we are not making the Voicebox model or code publicly available at this time. While we believe it is important to be open with the AI community and to share our research to advance the state of the art in AI, it’s also necessary to strike the right balance between openness with responsibility. Yeah... I mean there de…

They don't really care about misuse. They just don't want to say openly that they like to keep their shiny new tech and make money with it. No idea why, most people wouldn't bat an eye if you stated from the get go: We built it, we'll use it.
Post reply on HN