Live data from Hacker News

Voicebox: Generative AI model for speech that generalizes across tasks

ai.facebook.com

121–128 of 128 posts

Re: Voicebox: Generative AI model for speech that generalizes across tasks

#121
This is not close to " state of the art " in TTS. The output is clicky , low bitrate and lacks vocal nuance. Its novel for using a " flow matching " approach in its architecture and being suited to cloud - based translation. Have a listen to U Washington , Google " sound storm " instead !

Re: Voicebox: Generative AI model for speech that generalizes across tasks

#123
post #52
post #9

I am mostly excited for cheaper audiobooks with consistent voices for different characters.

I'm excited about them making it faster to produce. I finished the most recently published audiobook in a series this weekend. The author posts unpublished chapters to a site called Royal Road. I listen to books while running and driving, so it's a non-starter to visually read them. It would be nice to have that pipeline accelerated. Now, I just want to talk about my little weekend project... I spent a couple of hour…

Just FYI since you ended up using an external tts anyway -

https://beta.elevenlabs.io/speech-synthesis

is vastly better, especially for fiction.

Also worth trying is: https://speechify.com/

Re: Voicebox: Generative AI model for speech that generalizes across tasks

#124
post #73

Earlier quoted context omitted.

Sure, and then we'll get more audiobook options, as it becomes economically viable to make more niche stuff into audiobooks.

Niche stuff has always been economically viable. Niche stuff also tends to get publisher support. Realistically these days the only reason why books don't have an audiobook format is due to not wanting to support Amazon/Apple, or because they just don't want it. Voice actors literally are not expensive for audiobooks. You can absolutely afford one if your tiny $9 book hits more than 200 copies. Unless you're talking…

I think you’re underestimating how much work goes into hiring and producing something. Making it more self-service will go a long way toward making it more accessible to small creators. And my impression is that many books don’t sell 200 copies.

Re: Voicebox: Generative AI model for speech that generalizes across tasks

#125

can an expert in the field comment on whether the results are more or less impressive than Soundstorm?

I've got the same question; found some of the researchers for both projects on twitter and will see if I can get an opinion from one of them. Just waiting on verification to pm them. Will reply here if I hear back unless you have a contact with notifications.

Re: Voicebox: Generative AI model for speech that generalizes across tasks

#126

Earlier quoted context omitted.

Coming to an android and iphone near you voice authentication. Audio streams have inaudible data produced by encrypting a mutually known but changeable token like the current time with your private key embedded in the stream for example in frequencies you can't hear. Your phone app queries the service with the time of call and the data received and if they are also a subscriber it is able to discern their identity wi…

In a fake kidnapping scheme, it doesn’t seem like that much of a stretch to say “I’ve been kidnapped, and they took my keys, so I don’t have my yubikey” or something along those lines.

A lot of these scams evaporate like a soap bubble when someone questions it. All it can take is calling the persons phone to verify that they aren't actually kidnapped.

Re: Voicebox: Generative AI model for speech that generalizes across tasks

#127
post #111

Earlier quoted context omitted.

> It's nice to see the model can support your personal voice even if it's not completely neutral English There is no such thing as "neutral" English.

>Nonetheless, a form of speech known to linguists as General American is perceived by many Americans to be "accent-less", meaning a person who speaks in such a manner does not appear to be from anywhere in particular. The region of the United States that most resembles this is the central Midwest, specifically eastern Nebraska (including Omaha and Lincoln), southern and central Iowa (including Des Moines), parts of M…

> Nonetheless, a form of speech known to linguists as General American is perceived by many Americans to be "accent-less"

TLDR: "neutral English" is like "neutral water temperature" - it feels neither hot not cold because it matches ones body temperature. It's subjective, and terming it "temperatureless water" is even less accurate.

I'd put emphasis on "perceived" and "American" in that statement, and also note that this is limited to regional accents: General American is unambiguously American. Similar to General American, many countries have developed a "Newscaster" accent, e.g. Received Pronunciation for Britain, but it's not considered neutral as it is the "upper class" accent.

In every language I've known well enough to distinguish accents, I've realized newscasters adopt a distinct accent/cadence that's not commonly used. But I wouldn't call it "accentless" - it's just another accent that may/may not have evolved from a culturally dominant regional accent (or dominant figure from a specific region.)

Re: Voicebox: Generative AI model for speech that generalizes across tasks

#128
post #125

can an expert in the field comment on whether the results are more or less impressive than Soundstorm?

I've got the same question; found some of the researchers for both projects on twitter and will see if I can get an opinion from one of them. Just waiting on verification to pm them. Will reply here if I hear back unless you have a contact with notifications.

got a response here https://github.com/lucidrains/soundstorm-pytorch/discussions...
Post reply on HN