What's the best open source text to speech? Eleven Labs and others are interesting but closed source. I want to use them mainly for audiobooks as I have a lot of ePubs and I'm just using the basic Google text to speech voices on my Android, via Moon+ Reader. It works fine but it's still more robotic than state of the art.
Weird A.I. Yankovic: a cursed deep dive into the world of voice cloning
51–60 of 207 posts
Re: Weird A.I. Yankovic: a cursed deep dive into the world of voice cloning
#52I did not know about this: "The center of the A.I. cover songs community is a massive 500,000+ member Discord called A.I. Hub, where members trade new tips, tools, techniques, and links to their original and cover songs."
I poked around there for a while, and my takeaway was "sub-par" all around, which might be the reason for it's relative obscurity? The thing is, I can't tell to what extent it's the tech, and to what extent it's just "very uninteresting source material." Like, there's a whole lot of "classic song done by presently popular rapper," and I'll be the first to insist that there is nearly nothing vocally interesting at all…
Re: Weird A.I. Yankovic: a cursed deep dive into the world of voice cloning
#53This article only covers the musical aspects of AI voice cloning, but there's another dynamic to AI voice cloning that's more complicated: replacing general voice actors in movies/video games/anime (example: https://www.axios.com/2023/07/24/ai-voice-actors-victoria-at... ) Unlike musicians who can't be replaced without significant postprocessing, have enough money to not be impacted by competition, and have legal mus…
Most likey you'd see a lot of people saying that somehow getting rid of voice actors is good for "progress". Whatever that means.
Random aside someone really needs to make a hackernews that focuses more on game development and other arts so blog posts like your talking about would have a proper community to discuss them with.
Re: Weird A.I. Yankovic: a cursed deep dive into the world of voice cloning
#54Earlier quoted context omitted.
> and their voicework is owned by their clients who may not be as incentivised in protecting the VO. The work product produced by their voice for fulfilling the contract is owned. No corp owns someone else's voice.
They don't own the voice , but they own the vocal performance , which ends up being a meaningless legal distinction in practice. It's one reason why VAs rarely take fan requests for a character they voice.
I'm guessing contracts will need to be updated to say that a character's voice made from AI can't be used so a completely different production cannot say they have the actor attached for publicity purposes.
Re: Weird A.I. Yankovic: a cursed deep dive into the world of voice cloning
#55It's kind of wild that these tools just transfer a copy of these models every time they're spun up (whether it's to a Google Colab notebook or a local machine.) This must mean Hugging Face's bandwidth bill must be crazy, or am I missing something (maybe they have a peering agreement? heavily caching things?)
I really wish I could configure this crap to cache somewhere other than my C: drive Or better yet, how about asking me where I want to store my models?
Re: Weird A.I. Yankovic: a cursed deep dive into the world of voice cloning
#56Earlier quoted context omitted.
I always assume 200.to 250 pages per book when someone talks about large quantities of books.
That's fairly short. I read about 100 books a year and it includes thousand page tomes like The Count of Monte Cristo.
100 books/year. That's an impressive feat regardless the number of pages. Are these downloaded ebooks or physical printed copies of books?
Re: Weird A.I. Yankovic: a cursed deep dive into the world of voice cloning
#57Earlier quoted context omitted.
Their Python module caches the downloads, which is checked before downloading them again...but you're probably not wrong on the crazy bandwidth bill. Looks like they have crazy VC money though, considering the current climate.
The Colab notebooks are a fresh and independent session with no caching.
Re: Weird A.I. Yankovic: a cursed deep dive into the world of voice cloning
#58It's kind of wild that these tools just transfer a copy of these models every time they're spun up (whether it's to a Google Colab notebook or a local machine.) This must mean Hugging Face's bandwidth bill must be crazy, or am I missing something (maybe they have a peering agreement? heavily caching things?)
Unmetered 10+ gigabit connections were on the order of $1/mbit/mo wholesale over a decade ago when I priced out a custom CDN so for the cost of 100 TB of data transfer out of AWS you could get a 24/7 sustained 10gbit/s (>3 PB per month at 100% utilization). Bandwidth has always been crazy cheap.
Re: Weird A.I. Yankovic: a cursed deep dive into the world of voice cloning
#59Earlier quoted context omitted.
That's fairly short. I read about 100 books a year and it includes thousand page tomes like The Count of Monte Cristo.
I always assumed that book to be rather short since it just needs to be a number of sandwiches eaten. 100 books/year. That's an impressive feat regardless the number of pages. Are these downloaded ebooks or physical printed copies of books?
Re: Weird A.I. Yankovic: a cursed deep dive into the world of voice cloning
#60The sampled voices sound neither like Michael Jackson nor Weird Al. A good effort, but a professional impersonator could likely do better on either front.
It sounds like Weird Al trying to be Michael Jackson trying to be Weird Al.