Live data from Hacker News

Voicebox: Generative AI model for speech that generalizes across tasks

ai.facebook.com

81–90 of 128 posts

Re: Voicebox: Generative AI model for speech that generalizes across tasks

#81
post #17

That little stinger at the end was not as surprising as they thought it was :P It's very cool tech, but it's far from transparent. It has a very obvious "autotune" like sound to it that jumps right out. when they edited that one word it was obvious it had been edited. Again, super cool tech, just not going to replace voice actors or anything.

I agree with you that v1 isn't a suitable replacement... right now. But what v2? v3?

[deleted]

Re: Voicebox: Generative AI model for speech that generalizes across tasks

#82

That little stinger at the end was not as surprising as they thought it was :P It's very cool tech, but it's far from transparent. It has a very obvious "autotune" like sound to it that jumps right out. when they edited that one word it was obvious it had been edited. Again, super cool tech, just not going to replace voice actors or anything.

> It has a very obvious "autotune" To me it has a very obvious "Hindi is my native language" accent. I mean after literally the first sentence: "The research team at Meta is excited to share our work..." . Ouch. The "our work": just ouch. I was wondering why it wasn't a native english speaker presenting the video when the video is precisely about generating speech. The first seven seconds are particularly bad. Don't…

Maybe they deliberately chose an accent that wasn't native English to demonstrate the style transfer capability. I think the ability of the system to output accented voices is a strength not a weakness, so long as it can do other accents too.

Re: Voicebox: Generative AI model for speech that generalizes across tasks

#83
post #73

Earlier quoted context omitted.

Yes, but they're not going to pass on costs from not using a voice actor. They're just going to charge what they normally would have, and not worry about having to give so and so a cut.

Sure, and then we'll get more audiobook options, as it becomes economically viable to make more niche stuff into audiobooks.

Niche stuff has always been economically viable. Niche stuff also tends to get publisher support. Realistically these days the only reason why books don't have an audiobook format is due to not wanting to support Amazon/Apple, or because they just don't want it. Voice actors literally are not expensive for audiobooks. You can absolutely afford one if your tiny $9 book hits more than 200 copies.

Unless you're talking about illegal niche, in which fair but I highly doubt stores are going to accept those. All generation helps with is content that will be free.

Re: Voicebox: Generative AI model for speech that generalizes across tasks

#84

So are they releasing it or not? It's a nice PR statement but "we are not making the Voicebox model or code publicly available at this time". Phantom release?

They are holding it closed for the safety of the children and their future profit margins.

Re: Voicebox: Generative AI model for speech that generalizes across tasks

#85
post #40

Earlier quoted context omitted.

> There are many exciting use cases for generative speech models, but because of the risks of misuse, we are not making the Voicebox model or code publicly available at this time. While we believe it is important to be open with the AI community and to share our research to advance the state of the art in AI, it’s also necessary to strike the right balance between openness with responsibility. Yeah... I mean there de…

They don't really care about misuse. They just don't want to say openly that they like to keep their shiny new tech and make money with it. No idea why, most people wouldn't bat an eye if you stated from the get go: We built it, we'll use it.

I agree. Though in fairness they have opened other models e.g. for speech recognition.

Re: Voicebox: Generative AI model for speech that generalizes across tasks

#86

That little stinger at the end was not as surprising as they thought it was :P It's very cool tech, but it's far from transparent. It has a very obvious "autotune" like sound to it that jumps right out. when they edited that one word it was obvious it had been edited. Again, super cool tech, just not going to replace voice actors or anything.

> It has a very obvious "autotune" To me it has a very obvious "Hindi is my native language" accent. I mean after literally the first sentence: "The research team at Meta is excited to share our work..." . Ouch. The "our work": just ouch. I was wondering why it wasn't a native english speaker presenting the video when the video is precisely about generating speech. The first seven seconds are particularly bad. Don't…

The accent was obvious enough that I wonder if they might have not been trying to hide it at all? Maybe they just happened to pick somebody from the team with a very mild accent.

Re: Voicebox: Generative AI model for speech that generalizes across tasks

#88
post #4

I think that the "star trek" use case of a live translation is super exciting. I think that this also will force people to have pass phrases that they use to authenticate phone calls with. I normally downplay when people bring up everyone signing everything with a public/private key (impractical for normal users) but clearly there will be a need for authentication protocols as AI proliferates

Coming to an android and iphone near you voice authentication. Audio streams have inaudible data produced by encrypting a mutually known but changeable token like the current time with your private key embedded in the stream for example in frequencies you can't hear. Your phone app queries the service with the time of call and the data received and if they are also a subscriber it is able to discern their identity wi…

In a fake kidnapping scheme, it doesn’t seem like that much of a stretch to say “I’ve been kidnapped, and they took my keys, so I don’t have my yubikey” or something along those lines.

Re: Voicebox: Generative AI model for speech that generalizes across tasks

#89
post #52
post #9

I am mostly excited for cheaper audiobooks with consistent voices for different characters.

I'm excited about them making it faster to produce. I finished the most recently published audiobook in a series this weekend. The author posts unpublished chapters to a site called Royal Road. I listen to books while running and driving, so it's a non-starter to visually read them. It would be nice to have that pipeline accelerated. Now, I just want to talk about my little weekend project... I spent a couple of hour…

Yeah, it isn’t so much that I want publishers to have a cheaper way of making an audiobook that avoids the (apparently minimal) cost of employing a voice actor.

I don’t want to wait for the publisher to decide they want to do an audiobook.

Re: Voicebox: Generative AI model for speech that generalizes across tasks

#90
post #27

Earlier quoted context omitted.

For me its more like I wake up and check if humans have been replaced yet. Oh good, it's another day that I don't have to share one time pads with my mother to ensure that I'm talking to her and not a simulant performing fraud on a massive scale.

Cryptographic proof of personhood is going to be a thing, is it not? Outside of BigTech, Signal is as poised as WorldCoin to be just that.

I’m just not convinced that anything not tied directly to a government issued ID is going to be strong enough.
Post reply on HN