https://github.com/152334H/tortoise-tts-fast
The developer of tortoise tts fast was hired by Eleven labs.
61–70 of 168 posts
https://github.com/152334H/tortoise-tts-fast
The developer of tortoise tts fast was hired by Eleven labs.
From the example: "Oh no, I'm really sorry to hear you're having trouble with your new device. That sounds frustrating." Being patronized by a machine when you just want help is going to feel absolutely terrible. Not looking forward to this future.
"I can help you get a replacement. Here let me pull up a totally hallucinated order number and a link that goes nowhere. Did that solve your problem?"
We’re using elevenlabs in a new prototype, and it gets confused by its own voice which my mic picks up. Unless I wear headphones, it thinks I’m talking, and it gets into a loop. I hope this release fixes that bug!
On your client you need to implement some form of echo cancellation.
We’re using elevenlabs in a new prototype, and it gets confused by its own voice which my mic picks up. Unless I wear headphones, it thinks I’m talking, and it gets into a loop. I hope this release fixes that bug!
>Public API for Eleven v3 (alpha) is coming soon.
There is zero use for this without an API endpoint. At least is coming.
Unfortunately voice actors will be replaced by someThing like this hopefully they will find someThing else To do
Audible has ruined their catalog listings with their "Virtual voice" thing and no option to filter them out. They're mostly low quality books narrated by subpar AI voice that don't sell at all, while making it extremely difficult to find quality new books to listen to.
What's the state of open source tts? I'm a heavy TTS user, anything that can run at 3x-4x speed off enthusiast hardware?
dialogue like notebooklm: https://github.com/nari-labs/dia
Earlier quoted context omitted.
Ouch. Professional Voice Actor here.
Just here to say the oposite. It is astonshing how far away it still is from a professional voice actor while being really good. Emotion is completely missing. Instead it seems to try to hard to express exactly that. I cant really put my finger on it. It feels predictable, flat and the timing is strange.
Earlier quoted context omitted.
If you edit the text so that laugh makes sense in the context it should be much more natural like this one: https://x.com/elevenlabsio/status/1930689782331412811
The first laugh in that " Hey, Dr. Von Fusion" is a dedicated laugh section, which the model does extremely well, but it works because that's a natural place to laugh before actually speaking the following words. Skip ahead to "...robot chuckle. Jessica: I know right!" and you get an awkwardly time/toned light chuckle completely separated from the "I know" you'd naturally continue saying while making that chuckle. Yo…