Live data from Hacker News

Deep-learning text-to-speech tool for generating voices of various characters

15.ai

21–30 of 88 posts

Re: Deep-learning text-to-speech tool for generating voices of various characters

#22

will this be open source eventually?

https://twitter.com/fifteenai/status/1342304487474606081 found an answer. "There's no point in releasing a poorly done model, and to do so for the sake of popularity would be despicable. My goal is to achieve indistinguishability, which I certainly know is possible. Anything short of near-perfection is unacceptable. "

Releasing a poorly done intermediate result would give either competitors or colleagues a leg up in the race, depending on whether one sees them as competitors or colleagues.

Re: Deep-learning text-to-speech tool for generating voices of various characters

#23

will this be open source eventually?

https://twitter.com/fifteenai/status/1342304487474606081 found an answer. "There's no point in releasing a poorly done model, and to do so for the sake of popularity would be despicable. My goal is to achieve indistinguishability, which I certainly know is possible. Anything short of near-perfection is unacceptable. "

Megalomania, always a great excuse.

AI and ML users are massively benefiting from open source but too often refuse to release their data. It's like we're back in the middle ages and alchemy is back in style.

Re: Deep-learning text-to-speech tool for generating voices of various characters

#25

will this be open source eventually?

https://twitter.com/fifteenai/status/1342304487474606081 found an answer. "There's no point in releasing a poorly done model, and to do so for the sake of popularity would be despicable. My goal is to achieve indistinguishability, which I certainly know is possible. Anything short of near-perfection is unacceptable. "

I'm afraid this tweet is taken out of context. I had written this in response to complaints about the release date being delayed because I wanted to make sure that the released model (that is currently on the site) was the best it could be.

I do plan to compile and publish my findings in the future, but nothing is set in stone yet. I know that the model can be improved even further, and I'd prefer to be as comprehensive as possible.

Re: Deep-learning text-to-speech tool for generating voices of various characters

#26

Earlier quoted context omitted.

https://twitter.com/fifteenai/status/1342304487474606081 found an answer. "There's no point in releasing a poorly done model, and to do so for the sake of popularity would be despicable. My goal is to achieve indistinguishability, which I certainly know is possible. Anything short of near-perfection is unacceptable. "

Megalomania, always a great excuse. AI and ML users are massively benefiting from open source but too often refuse to release their data. It's like we're back in the middle ages and alchemy is back in style.

Judging by how the model and site are put together, I think this is some software engineer's hobby project. Not wanting to spill their secrets doesn't make them a megalomaniac for the same reason being a magician doesn't make one a megalomaniac.

Re: Deep-learning text-to-speech tool for generating voices of various characters

#27
I don't usually expect much from demos like this, but I'm kind of surprised how impressive the results currently are. They're definitely not perfect, you're definitely getting some odd clipping and noise, but this shows a large amount of promise.

Being able to generate voices for games would enable a lot of interesting indie projects. IMO people should be paying more attention the market implications of products like this than to the social implications. There are a lot of projects that just aren't really feasible right now that could be if this kind of technology was more polished and generally available for commercial/self-hosted use. And in those cases, you don't even need to do inference, makers will likely be willing to mark up their scripts themselves.

Anyway I digress. Congrats, this is really cool!

Re: Deep-learning text-to-speech tool for generating voices of various characters

#28

I don't usually expect much from demos like this, but I'm kind of surprised how impressive the results currently are. They're definitely not perfect, you're definitely getting some odd clipping and noise, but this shows a large amount of promise. Being able to generate voices for games would enable a lot of interesting indie projects. IMO people should be paying more attention the market implications of products like…

> people should be paying more attention the market implications of products like this than to the social implications.

People will absolutely suffer harm from this tech, but hey, think about the dollars that could be made! No, we should absolutely be paying more attention to the social implications.

Re: Deep-learning text-to-speech tool for generating voices of various characters

#29
post #17
post #6

From the about section: > How much does maintaining the servers cost? > It depends on the amount of traffic, but the minimum baseline is around several thousands of US dollars every month. This is expected as inference is very GPU intensive and a sufficient number of instances need to be spun up to handle thousands of requests coming in every minute. Everything is paid out of pocket. Wow, impressive commitment for so…

You just sort of assume that this is correct? The person[1] running this comes across as a severely unstable character, that number is probably hyperbole. [1] https://twitter.com/fifteenai

I’ve worked with deep learning models enough to know the cost of running GPU inference, and if the live queue stats published on the website are accurate, then thousands of dollars per month is certainly plausible.

I have no reason to disbelieve it.

Re: Deep-learning text-to-speech tool for generating voices of various characters

#30
post #28

I don't usually expect much from demos like this, but I'm kind of surprised how impressive the results currently are. They're definitely not perfect, you're definitely getting some odd clipping and noise, but this shows a large amount of promise. Being able to generate voices for games would enable a lot of interesting indie projects. IMO people should be paying more attention the market implications of products like…

> people should be paying more attention the market implications of products like this than to the social implications. People will absolutely suffer harm from this tech, but hey, think about the dollars that could be made! No, we should absolutely be paying more attention to the social implications.

Eh, this technology currently falls very squarely into the category of "almost good enough that I could use it for a creative project, but not nearly good enough that you're going to be able to convince me that the results aren't generated."

I'm not primarily interested about the dollars, I'm interested in allowing communities to do creative things. I think people are looking at this tech like it's only going to be used for deepfakes, and they're underestimating the extent it's going to be used to create voice-acted game mods, animations, anonymization tools, and other creative/helpful projects.

If you're really worried about this stuff though, you can take some comfort in the fact that by far the worst examples on the site are of real-world voices. This is currently technology that as far as I can see is far more suited for generating new voices or voicing cartoon characters with well-defined patterns/inflections than it is for imitating the president.

Post reply on HN