Live data from Hacker News

StyleTTS2 – open-source Eleven-Labs-quality Text To Speech

github.com

241–245 of 245 posts

Re: StyleTTS2 – open-source Eleven-Labs-quality Text To Speech

#241

Earlier quoted context omitted.

How hard on your end does the task of making the chatbot converse naturally look? Specifically I'm thinking about interruptions, if it's talking too long I would like to be able to start talking and interrupt it like in a normal conversation, or if I'm saying something it could quickly interject something. Once you've got the extremely high speed, theoretically faster than real time, you can start doing that stuff ri…

Yes, I implemented the ability to interrupt the chatbot while it is talking. It wasn't too hard, although it does require you to wear headphones so the bot doesn't hear itself and get interrupted. The other way around (bot interrupting the user) is hard. Currently the bot starts processing a response after every word that the voice recognition outputs, to reduce latency. When new words come in before the response is…

I installed this on my system, having a little issue like you mention below where the bot hears itself so it got into a loop of talking to itself and reply to itself.

Could the issue be that I am using a pair of bluetooth headphones and the microphone is built into that - what is the optimum setup, should I be listening on the headphones and using a different mic input instead of the headphone mic?

This is pretty intermeeting I would love to get it to work. Running a 3060.

Thanks.

Re: StyleTTS2 – open-source Eleven-Labs-quality Text To Speech

#242

Earlier quoted context omitted.

How hard on your end does the task of making the chatbot converse naturally look? Specifically I'm thinking about interruptions, if it's talking too long I would like to be able to start talking and interrupt it like in a normal conversation, or if I'm saying something it could quickly interject something. Once you've got the extremely high speed, theoretically faster than real time, you can start doing that stuff ri…

Yes, I implemented the ability to interrupt the chatbot while it is talking. It wasn't too hard, although it does require you to wear headphones so the bot doesn't hear itself and get interrupted. The other way around (bot interrupting the user) is hard. Currently the bot starts processing a response after every word that the voice recognition outputs, to reduce latency. When new words come in before the response is…

I installed this on my system, having a little issue like you mention below where the bot hears itself so it got into a loop of talking to itself and replying to itself, it was pretty funny.

Could the issue be that I am using a pair of bluetooth headphones and the microphone is built into that - what is the optimum setup, should I be listening on the headphones and using a different mic input instead of the headphone mic?

This is pretty intermeeting I would love to get it to work. Running a 3060.

Thanks.

Re: StyleTTS2 – open-source Eleven-Labs-quality Text To Speech

#243

Earlier quoted context omitted.

Yes, I implemented the ability to interrupt the chatbot while it is talking. It wasn't too hard, although it does require you to wear headphones so the bot doesn't hear itself and get interrupted. The other way around (bot interrupting the user) is hard. Currently the bot starts processing a response after every word that the voice recognition outputs, to reduce latency. When new words come in before the response is…

I installed this on my system, having a little issue like you mention below where the bot hears itself so it got into a loop of talking to itself and replying to itself, it was pretty funny. Could the issue be that I am using a pair of bluetooth headphones and the microphone is built into that - what is the optimum setup, should I be listening on the headphones and using a different mic input instead of the headphone…

That's weird, Bluetooth headphones should work I think. Maybe there's something wrong with them. You can check if the mic can hear the speakers using Windows sound recorder. Try a different mic if you have one, the only important thing is the mic can't hear the headphones.

Re: StyleTTS2 – open-source Eleven-Labs-quality Text To Speech

#245
post #211

Earlier quoted context omitted.

That might have been true about a year ago, but I've been getting calls from well-spoken native-level scammers for about two months now. They are so frequent that I can put them on speaker during family gatherings to raise awareness. Sample sizes of 1 are never representative but they definitely have full access to native speakers or tech that can generate very passable speech.

It seems quite possible that the change you've seen in these last two months is because some have started using these models. More likely than a sudden huge shift in either the country of origin or English skills of the scammers.

My point is that these models were already out there before StyleTTS2 was released. Plugging your ears and demanding their regulation in your country will not make them disappear.
Post reply on HN