Live data from Hacker News

Deep-learning text-to-speech tool for generating voices of various characters

15.ai

31–40 of 88 posts

Re: Deep-learning text-to-speech tool for generating voices of various characters

#31

Welp, after messing around with a few voices I was completely impressed with Glados's. This is really cool because I have no idea how the character's voice was synthesized, but apparently ML can do it for me so props to that.

GLaDOS was voiced by a real person [0]. Her voice had some effects added but mostly just her trying to sound like a computer.

[0] http://ellenmclain.net/

Re: Deep-learning text-to-speech tool for generating voices of various characters

#32
post #17
post #6

From the about section: > How much does maintaining the servers cost? > It depends on the amount of traffic, but the minimum baseline is around several thousands of US dollars every month. This is expected as inference is very GPU intensive and a sufficient number of instances need to be spun up to handle thousands of requests coming in every minute. Everything is paid out of pocket. Wow, impressive commitment for so…

You just sort of assume that this is correct? The person[1] running this comes across as a severely unstable character, that number is probably hyperbole. [1] https://twitter.com/fifteenai

It seems like one could get to those numbers pretty easily given the prices for GPU instances on AWS. Even just one decent-sized instance would be thousands of dollars per month.

Re: Deep-learning text-to-speech tool for generating voices of various characters

#33
post #28

Earlier quoted context omitted.

> people should be paying more attention the market implications of products like this than to the social implications. People will absolutely suffer harm from this tech, but hey, think about the dollars that could be made! No, we should absolutely be paying more attention to the social implications.

Eh, this technology currently falls very squarely into the category of "almost good enough that I could use it for a creative project, but not nearly good enough that you're going to be able to convince me that the results aren't generated." I'm not primarily interested about the dollars, I'm interested in allowing communities to do creative things. I think people are looking at this tech like it's only going to be u…

I might be misunderstanding you, but there are no real-world voices on the site? All of them are of characters.

Re: Deep-learning text-to-speech tool for generating voices of various characters

#34
post #18
post #17

Earlier quoted context omitted.

You just sort of assume that this is correct? The person[1] running this comes across as a severely unstable character, that number is probably hyperbole. [1] https://twitter.com/fifteenai

Not a hyperbole – I can provide proof if you'd like.

Separate question - is this English only? It looks like you can feed in phonemes but I assume this has been trained with English audio.

Re: Deep-learning text-to-speech tool for generating voices of various characters

#35

Earlier quoted context omitted.

Eh, this technology currently falls very squarely into the category of "almost good enough that I could use it for a creative project, but not nearly good enough that you're going to be able to convince me that the results aren't generated." I'm not primarily interested about the dollars, I'm interested in allowing communities to do creative things. I think people are looking at this tech like it's only going to be u…

I might be misunderstanding you, but there are no real-world voices on the site? All of them are of characters.

I see a pretty linear drop in quality from Glados to Spongebob to Twilight Sparkle to the narrator from Stanley Parable to the 10th Doctor.

It seems to struggle more and more as the voices get less cartoony/exaggerated.

Re: Deep-learning text-to-speech tool for generating voices of various characters

#36

Earlier quoted context omitted.

I might be misunderstanding you, but there are no real-world voices on the site? All of them are of characters.

I see a pretty linear drop in quality from Glados to Spongebob to Twilight Sparkle to the narrator from Stanley Parable to the 10th Doctor. It seems to struggle more and more as the voices get less cartoony/exaggerated.

I'm not too sure about that. From my testing, Fluttershy, Applejack, Twilight, Chrysalis, Rise, and Kyu (and a bunch of other characters that I'm surely forgetting) seem to perform phenomenally well. Especially Chrysalis, her emotions are extremely believable, and Fluttershy/Applejack/Rise/Kyu have almost zero noise for every generation. This might be the most impressive site I've ever seen.

Oh, I somehow forgot all of the TF2 characters. Some of them do struggle (Medic the most, I think) but everyone else seems incredibly good.

And the Daria characters, too. Honestly, the vast majority of characters are already near-perfect.

Re: Deep-learning text-to-speech tool for generating voices of various characters

#38
post #28

Earlier quoted context omitted.

> people should be paying more attention the market implications of products like this than to the social implications. People will absolutely suffer harm from this tech, but hey, think about the dollars that could be made! No, we should absolutely be paying more attention to the social implications.

Eh, this technology currently falls very squarely into the category of "almost good enough that I could use it for a creative project, but not nearly good enough that you're going to be able to convince me that the results aren't generated." I'm not primarily interested about the dollars, I'm interested in allowing communities to do creative things. I think people are looking at this tech like it's only going to be u…

You are looking at the current implementation and not thinking about the implication.

One, this tech absolutely could be used to fool someone. Not everyone will be listening with a critical ear. Played back over a phone or injecting a phrase or two in otherwise spoken samples will fool many people.

I guarantee you someone will be using this to make their own MLP episodes on YouTube specifically designed to scare children or get them to do awful things.

Models presumably get better over time. It really won't be too much longer until people will be able to fake celebrities, politicians, exes, authority figures, etc. As a fairly benign example, if I had this in high school you better believe I could have called to excuse some of my absences.

I agree, I love the idea of generating some decent voice lines for my own games projects, but this also introduces issues of the rights of the original voice actors.

If you train a model to mimic a performance given by an actor, then use that model and fire the actor, isn't that potentially really problematic? (Also, it draws parallels to the Luddites who were not anti technology, but wanted to ensure that technology wasn't used in a way that reduced worker quality of life.)

And yes, I think there are helpful ways this could be deployed. I'm gender fluid, and I'd love to be able to adjust my voice digitally, but we need to be thinking about how this could cause harm first.

Re: Deep-learning text-to-speech tool for generating voices of various characters

#39

Earlier quoted context omitted.

I see a pretty linear drop in quality from Glados to Spongebob to Twilight Sparkle to the narrator from Stanley Parable to the 10th Doctor. It seems to struggle more and more as the voices get less cartoony/exaggerated.

I'm not too sure about that. From my testing, Fluttershy, Applejack, Twilight, Chrysalis, Rise, and Kyu (and a bunch of other characters that I'm surely forgetting) seem to perform phenomenally well. Especially Chrysalis, her emotions are extremely believable, and Fluttershy/Applejack/Rise/Kyu have almost zero noise for every generation. This might be the most impressive site I've ever seen. Oh, I somehow forgot all…

Hrm. Well, I can't really argue with that beyond that my standards on perfect might be different.

I think some of the best voices they have are characters like Twilight, she shows a ton of promise. But as it stands right now, I would still at least hesitate to use Twilight's voice in a project unless I didn't have other options. Chrysalis's voice is good, but again, is an exaggerated cartoon character with a large amount of inflection. I would not use her voice in her current state without a lot of post-processing. Someone like the Spy I would consider to be unusable, it sounds to me like the character needs to clear their throat or something, it's got a lot of strange artifacts. I definitely would consider the 10th Doctor unusable, even for just a hobby project or a voice assistant.

But... I don't know, maybe this is subjective. I can't just tell you that what you're hearing is wrong, if you like the results then you like the results :)

And again, I don't want to detract from how impressive they are. They are incredibly impressive, particularly because of how characters like Chrysalis emote. Extremely promising. But I still think there's a difference between 'impressive' and 'believable deepfake'.

Re: Deep-learning text-to-speech tool for generating voices of various characters

#40
post #28

Earlier quoted context omitted.

> people should be paying more attention the market implications of products like this than to the social implications. People will absolutely suffer harm from this tech, but hey, think about the dollars that could be made! No, we should absolutely be paying more attention to the social implications.

Eh, this technology currently falls very squarely into the category of "almost good enough that I could use it for a creative project, but not nearly good enough that you're going to be able to convince me that the results aren't generated." I'm not primarily interested about the dollars, I'm interested in allowing communities to do creative things. I think people are looking at this tech like it's only going to be u…

It really doesnt have to be perfect to trick someone. You're expecting this site to be fake so you're listening carefully. If you weren't expecting anything and you were in the middle of a busy day at work, you are much much less likely to notice any discripencies.

We already have stories like https://www.forbes.com/sites/jessedamiani/2019/09/03/a-voice...

That said, as far as harms go, i dont think this is all that bad that it should preclude creative uses of this technology.

Post reply on HN