Live data from Hacker News

AI speech generator 'reaches human parity' – but it's too dangerous to release

livescience.com

51–60 of 68 posts

Re: AI speech generator 'reaches human parity' – but it's too dangerous to release

#51
post #4

Speech generation has gotten really good, but there's simply no way to faithfully recreate someone's vocal idiosyncracies and cadence with just "a few seconds" of real audio. That's where the models tend to fall short.

Few seconds means less than a minute. That’s not nothing. Look at a clock and talk for a minute — it’s longer than you might think.

Do you think you could give a recording of a minute of someone talking to a talented impressionist and they could impersonate that person to some degree? It doesn’t seem that far fetched to me.

Re: AI speech generator 'reaches human parity' – but it's too dangerous to release

#52

What is the point of them trying to create this? That something like this would mostly be used to create disinformation and create chaos is easily understood before making something like this. Truly irresponsible

There are legitimate uses of this tech, such as preserving voices of people losing them such as Stephen Hawking, or making it better for blind/low vision people to follow text and interact with devices. For that latter case having a more natural voice that is also accurate is a good thing.

I use TTS to listen to articles and stories that don't have access to an audiobook narrator. I've used some of the voices based on MBROLA tech, but those can grate after a while.

The more recent voice models are a lot higher quality and emotive (without the jarring pitch transitions of things like Cepstral) so are better to listen to. However, the recent models can clip/skip text, have prolonged silence, have long/warped/garbled words, etc. that make them harder to use longer term.

Re: AI speech generator 'reaches human parity' – but it's too dangerous to release

#53

Ah, the old "it's too dangerous to release" marketing move. Why even tell us about it?

Not everything is a conspiracy. It is very plausible they're speaking genuinely.

This isn't the first time someone has spoken like this. It won't be the last. It's 100% marketing.

Re: AI speech generator 'reaches human parity' – but it's too dangerous to release

#54

What is the point of them trying to create this? That something like this would mostly be used to create disinformation and create chaos is easily understood before making something like this. Truly irresponsible

Agreed.

On the one hand, I would love this kind of tech to be available for entertainment purposes. An RPG with convincing NPCs that are able to provide a novel experience for every player? Sounds great.

On the other: this is fraught with ethical problems, not to mention an ideal tool for fraud. At worst, it could be used as a weapon for total asymmetrical warfare on concepts like media integrity and an ideal tool for character assassination; disinformation, propaganda, etc.

I would happy welcome a world where this stuff is nerfed across the board, where videogames and porn are just chock full of AI voice-acting artifacts. We'll adjust and accept that as just a part of the experience, as we have with low fidelity media of the past. But my more cynical side tells me that's not what people in power are concerned about.

Re: AI speech generator 'reaches human parity' – but it's too dangerous to release

#56

Ah, the old "it's too dangerous to release" marketing move. Why even tell us about it?

Not everything is a conspiracy. It is very plausible they're speaking genuinely.

I have a bridge I'm selling....

Re: AI speech generator 'reaches human parity' – but it's too dangerous to release

#57

A gun can help rob a bank. A speech generator can help rob 1000 banks.

it might help if the dumb dumbs at the bank would stop trying to make me say "my voice is my password". I've been careful to only say "no fcuk off you fcuking numpty who came up with this idea after voice cloning hit the mainstream".

Re: AI speech generator 'reaches human parity' – but it's too dangerous to release

#58
post #53

Earlier quoted context omitted.

Not everything is a conspiracy. It is very plausible they're speaking genuinely.

This isn't the first time someone has spoken like this. It won't be the last. It's 100% marketing.

Based on what proof?

Re: AI speech generator 'reaches human parity' – but it's too dangerous to release

#59
post #39

Earlier quoted context omitted.

Not everything is a conspiracy. It is very plausible they're speaking genuinely.

And it's even more plausible this is just a marketing play to build hype. Take, for example, you just made some new super-pathogen in your basement lab. It could kill everyone on the planet. This is obviously pretty dangerous, so do you: A) silently dispose of it and hope nobody else ever makes the mistake of creating it. or B) keep it in the freezer and hold a press release about how you made it but it's too dangero…

No, it's not more plausible.

Re: AI speech generator 'reaches human parity' – but it's too dangerous to release

#60
post #38

I really want something that can do a voice change and match the emotion and articulation of a voice clip that I provide. I don't care (or want) it to be based off a real person and the manners in which they would tend to articulate a sentence. Are there any decent open models out there?

Try StyleTTS2. You will still have to experiment with the settings a little to get the right level of adherence to the reference speaker’s voice and the emotion content.

Without looking at this, are you sure that this can do speech to speech? Maybe my flaw in searching has been disregarding anything that's called "text to speech" as not also "speech to speech"?
Post reply on HN