That little stinger at the end was not as surprising as they thought it was :P It's very cool tech, but it's far from transparent. It has a very obvious "autotune" like sound to it that jumps right out. when they edited that one word it was obvious it had been edited. Again, super cool tech, just not going to replace voice actors or anything.
Voicebox: Generative AI model for speech that generalizes across tasks
51–60 of 128 posts
Re: Voicebox: Generative AI model for speech that generalizes across tasks
#52I am mostly excited for cheaper audiobooks with consistent voices for different characters.
Now, I just want to talk about my little weekend project... I spent a couple of hours scraping Royal Road and trying to get TTS working. Eventually, I settled on:
1. `wget --recursive` filtering only the chapters 2. A python script to strip extraneous html like advertisements and the headers. 3. Pipe into pandoc emitting plain text. 4. Copy it to my phone for TTS: https://f-droid.org/packages/com.danefinlay.ttsutil/
I really wanted to use all local tools, but I just couldn't get any of the Linux tools to sound as good or work as fast as Google TTS services. Also, the TTS paid services I found were just too expensive to justify (20hr book for ~$70).
I'm more than happy to additionally purchase the audiobook when it is published. I just don't want to wait.
Re: Voicebox: Generative AI model for speech that generalizes across tasks
#53Re: Voicebox: Generative AI model for speech that generalizes across tasks
#54That little stinger at the end was not as surprising as they thought it was :P It's very cool tech, but it's far from transparent. It has a very obvious "autotune" like sound to it that jumps right out. when they edited that one word it was obvious it had been edited. Again, super cool tech, just not going to replace voice actors or anything.
What stinger at the end?
Re: Voicebox: Generative AI model for speech that generalizes across tasks
#55Re: Voicebox: Generative AI model for speech that generalizes across tasks
#56I am mostly excited for cheaper audiobooks with consistent voices for different characters.
For what its worth, most of the cost of audiobooks doesn't come from paying talent. For intermediate level actors, the going rate is around $50-$100 per finished hour (PFH) and experienced actors it can be around $250-$300. This page does a decent job of laying out pay structures for audiobooks: https://speechify.com/blog/whats-the-meaning-of-per-finished...
An 8 hour audio book might cost the author/producer about $1800-$2k.
Just talking about Audible exclusively, they take about %50 of sales. But it's kinda wishy washy about exactly how much an author will earn in royalties. It's not as much as you might think. Good article from an author here that lays out some sales numbers: https://selfpublishingadvice.org/how-audiobook-authors-are-p...
The other way that a narrator can get paid is called royalty share. That means the author/producer doesn't pay the narrator anything up front and the voice actor then relies on a small percent of each book sale to get paid. Theoretically, if an audiobook ends up really taking off then the narrator potentially could make a lot of money. But that rarely happens. Most audiobooks that you find on Audible have very, very low sales volumes.
To sum it up, it doesn't occur to lost of audiobook fans but voice acting is a very competitive industry. It takes a lot of work to make a name for yourself, and even then the most successful actors probably aren't making much more than a highly paid software engineer. For most wannabe voice actors (including myself), its something you do more for love than necessarily to make a career out of it. Though of course, lots of people do but not the majority.
This is all why I'm personally not a fan of these voice generation models. It's going to eventually make this niche industry non-competitive for real humans except for the talent that is already established. People keep blaming the actors as being too expensive when most are barely making it without secondary jobs.
Re: Voicebox: Generative AI model for speech that generalizes across tasks
#57Earlier quoted context omitted.
Meta products have more users than ever
Doesn’t prove much honestly. How do you even know those users are real ? What value are the products providing ? My Apple products continue to improve my day to day immensely. The meta products are just rubbish on the whole. How has Instagram improved in the last 5 years ?
Re: Voicebox: Generative AI model for speech that generalizes across tasks
#58I think that the "star trek" use case of a live translation is super exciting. I think that this also will force people to have pass phrases that they use to authenticate phone calls with. I normally downplay when people bring up everyone signing everything with a public/private key (impractical for normal users) but clearly there will be a need for authentication protocols as AI proliferates
How "live" can translation ever really be? Properly translating anything from one language to another requires context.
Re: Voicebox: Generative AI model for speech that generalizes across tasks
#59Earlier quoted context omitted.
[flagged]
Do you use an iPhone ? They do pretty amazing things with images now. Even the search for a photo by text is quite amazing. I used it the other day for work for the first time and I found what I needed in my tens of thousands of photos. It almost truly felt like an extension of my memory, it was actually a pretty cool feeling. So sorry I don’t buy the Siri attack as being proof of anything. I have found Siri has impr…
Re: Voicebox: Generative AI model for speech that generalizes across tasks
#60That little stinger at the end was not as surprising as they thought it was :P It's very cool tech, but it's far from transparent. It has a very obvious "autotune" like sound to it that jumps right out. when they edited that one word it was obvious it had been edited. Again, super cool tech, just not going to replace voice actors or anything.
For me its more like I wake up and check if humans have been replaced yet. Oh good, it's another day that I don't have to share one time pads with my mother to ensure that I'm talking to her and not a simulant performing fraud on a massive scale.