Live data from Hacker News

StyleTTS2 – open-source Eleven-Labs-quality Text To Speech

github.com

201–210 of 245 posts

Re: StyleTTS2 – open-source Eleven-Labs-quality Text To Speech

#201

Earlier quoted context omitted.

How hard on your end does the task of making the chatbot converse naturally look? Specifically I'm thinking about interruptions, if it's talking too long I would like to be able to start talking and interrupt it like in a normal conversation, or if I'm saying something it could quickly interject something. Once you've got the extremely high speed, theoretically faster than real time, you can start doing that stuff ri…

Yes, I implemented the ability to interrupt the chatbot while it is talking. It wasn't too hard, although it does require you to wear headphones so the bot doesn't hear itself and get interrupted. The other way around (bot interrupting the user) is hard. Currently the bot starts processing a response after every word that the voice recognition outputs, to reduce latency. When new words come in before the response is…

Instead of sacrificing flexibility by building one monolith model that does Audio to audio in one go, wouldn't it be better to train a model that handles conversing with the user (knows when the user is done talking, when it's hearing itself, etc) and leave the thinking to other, more generic models?

Re: StyleTTS2 – open-source Eleven-Labs-quality Text To Speech

#202
post #200
post #199

This is really harmful and unethical work. It will be used to hurt millions of elderly people with scams. That's the real application that will happen 100x more than anything else. It's unethical and harmful to release tools that will be overwhelmingly used to hurt elderly people. What they should do about it is: Stop releasing models. Only release a service so that scammers will not use it. Also, only released audio…

Millions of elderly people are already getting scammed by overseas call centers so unless we do something more significant this tech will not make one iota of a difference.

That's not really true, most scammers have a male voice with a heavy accent. When they have tools that easily disguise their voice, scammers can reach many more elderly people.

Re: StyleTTS2 – open-source Eleven-Labs-quality Text To Speech

#203

Earlier quoted context omitted.

Yes, I implemented the ability to interrupt the chatbot while it is talking. It wasn't too hard, although it does require you to wear headphones so the bot doesn't hear itself and get interrupted. The other way around (bot interrupting the user) is hard. Currently the bot starts processing a response after every word that the voice recognition outputs, to reduce latency. When new words come in before the response is…

Instead of sacrificing flexibility by building one monolith model that does Audio to audio in one go, wouldn't it be better to train a model that handles conversing with the user (knows when the user is done talking, when it's hearing itself, etc) and leave the thinking to other, more generic models?

You don't lose flexibility with an end to end model. You lose controllability. But there are ways to mitigate that.

Re: StyleTTS2 – open-source Eleven-Labs-quality Text To Speech

#205
post #199

This is really harmful and unethical work. It will be used to hurt millions of elderly people with scams. That's the real application that will happen 100x more than anything else. It's unethical and harmful to release tools that will be overwhelmingly used to hurt elderly people. What they should do about it is: Stop releasing models. Only release a service so that scammers will not use it. Also, only released audio…

Just imagine if this line of thinking was used elsewhere.

This tech is already out of the bag and I thank the author(s) for the contribution to humanity. The correct solution here is not to shove your head in the sand and ignore reality, but to get your government to penalize any country or company that facilitates this crime. If they can force severe penalties for other financial crimes and funding terrorism, they can do the same here.

Re: StyleTTS2 – open-source Eleven-Labs-quality Text To Speech

#206
post #202
post #200

Earlier quoted context omitted.

Millions of elderly people are already getting scammed by overseas call centers so unless we do something more significant this tech will not make one iota of a difference.

That's not really true, most scammers have a male voice with a heavy accent. When they have tools that easily disguise their voice, scammers can reach many more elderly people.

That might have been true about a year ago, but I've been getting calls from well-spoken native-level scammers for about two months now. They are so frequent that I can put them on speaker during family gatherings to raise awareness.

Sample sizes of 1 are never representative but they definitely have full access to native speakers or tech that can generate very passable speech.

Re: StyleTTS2 – open-source Eleven-Labs-quality Text To Speech

#207
post #199

This is really harmful and unethical work. It will be used to hurt millions of elderly people with scams. That's the real application that will happen 100x more than anything else. It's unethical and harmful to release tools that will be overwhelmingly used to hurt elderly people. What they should do about it is: Stop releasing models. Only release a service so that scammers will not use it. Also, only released audio…

Just imagine if this line of thinking was used elsewhere. This tech is already out of the bag and I thank the author(s) for the contribution to humanity. The correct solution here is not to shove your head in the sand and ignore reality, but to get your government to penalize any country or company that facilitates this crime. If they can force severe penalties for other financial crimes and funding terrorism, they c…

it's funny because just yesterday I posted:

> soon as it's out, a whole bunch of extremely privileged ML people will throw their hands up and say, "oh well, cats out of the bag."

https://news.ycombinator.com/context?id=38324742

Re: StyleTTS2 – open-source Eleven-Labs-quality Text To Speech

#208
post #179

Earlier quoted context omitted.

See my previous comment about this point. ElevenLabs are based on Tortoise-TTS which was already pre-trained on millions of hours of data, but this one was only trained on LibriTTS which was 500 hours at best. XTTS was also trained with probably millions of speakers in more than 20 languages. If you have seen millions of voices, there are definitely gonna be some of them that sound like you. It is just a matter of tr…

What's your basis for the claim that they are based on TorToiSe? I have seen this claim made (and rebutted) many times.

Very similar features, quite slow inference speed, and various rumors.

Re: StyleTTS2 – open-source Eleven-Labs-quality Text To Speech

#209
post #199

This is really harmful and unethical work. It will be used to hurt millions of elderly people with scams. That's the real application that will happen 100x more than anything else. It's unethical and harmful to release tools that will be overwhelmingly used to hurt elderly people. What they should do about it is: Stop releasing models. Only release a service so that scammers will not use it. Also, only released audio…

Cars actually kill over a million of people per year. Not saying this is good, just that all technology has its tradeoffs.
Post reply on HN