Live data from Hacker News

Weird A.I. Yankovic: a cursed deep dive into the world of voice cloning

waxy.org

51–60 of 207 posts

Re: Weird A.I. Yankovic: a cursed deep dive into the world of voice cloning

#51

What's the best open source text to speech? Eleven Labs and others are interesting but closed source. I want to use them mainly for audiobooks as I have a lot of ePubs and I'm just using the basic Google text to speech voices on my Android, via Moon+ Reader. It works fine but it's still more robotic than state of the art.

[deleted]

Re: Weird A.I. Yankovic: a cursed deep dive into the world of voice cloning

#52
post #43
post #2

I did not know about this: "The center of the A.I. cover songs community is a massive 500,000+ member Discord called A.I. Hub, where members trade new tips, tools, techniques, and links to their original and cover songs."

I poked around there for a while, and my takeaway was "sub-par" all around, which might be the reason for it's relative obscurity? The thing is, I can't tell to what extent it's the tech, and to what extent it's just "very uninteresting source material." Like, there's a whole lot of "classic song done by presently popular rapper," and I'll be the first to insist that there is nearly nothing vocally interesting at all…

[deleted]

Re: Weird A.I. Yankovic: a cursed deep dive into the world of voice cloning

#53

This article only covers the musical aspects of AI voice cloning, but there's another dynamic to AI voice cloning that's more complicated: replacing general voice actors in movies/video games/anime (example: https://www.axios.com/2023/07/24/ai-voice-actors-victoria-at... ) Unlike musicians who can't be replaced without significant postprocessing, have enough money to not be impacted by competition, and have legal mus…

I'd be interested.

Most likey you'd see a lot of people saying that somehow getting rid of voice actors is good for "progress". Whatever that means.

Random aside someone really needs to make a hackernews that focuses more on game development and other arts so blog posts like your talking about would have a proper community to discuss them with.

Re: Weird A.I. Yankovic: a cursed deep dive into the world of voice cloning

#54

Earlier quoted context omitted.

> and their voicework is owned by their clients who may not be as incentivised in protecting the VO. The work product produced by their voice for fulfilling the contract is owned. No corp owns someone else's voice.

They don't own the voice , but they own the vocal performance , which ends up being a meaningless legal distinction in practice. It's one reason why VAs rarely take fan requests for a character they voice.

If they are using their real voice, then they kind of screwed themselves. If they are performing a character voice, then at least they only lose out on that kind of work.

I'm guessing contracts will need to be updated to say that a character's voice made from AI can't be used so a completely different production cannot say they have the actor attached for publicity purposes.

Re: Weird A.I. Yankovic: a cursed deep dive into the world of voice cloning

#55
post #25
post #3

It's kind of wild that these tools just transfer a copy of these models every time they're spun up (whether it's to a Google Colab notebook or a local machine.) This must mean Hugging Face's bandwidth bill must be crazy, or am I missing something (maybe they have a peering agreement? heavily caching things?)

I really wish I could configure this crap to cache somewhere other than my C: drive Or better yet, how about asking me where I want to store my models?

I haven’t used windows in a while but I thought it supported some form of cross-volume symlink? Or at least mounting an image stored on another volume to an arbitrary path.

Re: Weird A.I. Yankovic: a cursed deep dive into the world of voice cloning

#56

Earlier quoted context omitted.

I always assume 200.to 250 pages per book when someone talks about large quantities of books.

That's fairly short. I read about 100 books a year and it includes thousand page tomes like The Count of Monte Cristo.

I always assumed that book to be rather short since it just needs to be a number of sandwiches eaten.

100 books/year. That's an impressive feat regardless the number of pages. Are these downloaded ebooks or physical printed copies of books?

Re: Weird A.I. Yankovic: a cursed deep dive into the world of voice cloning

#57
post #6

Earlier quoted context omitted.

Their Python module caches the downloads, which is checked before downloading them again...but you're probably not wrong on the crazy bandwidth bill. Looks like they have crazy VC money though, considering the current climate.

The Colab notebooks are a fresh and independent session with no caching.

Google might cache further up the chain, which could help

Re: Weird A.I. Yankovic: a cursed deep dive into the world of voice cloning

#58
post #3

It's kind of wild that these tools just transfer a copy of these models every time they're spun up (whether it's to a Google Colab notebook or a local machine.) This must mean Hugging Face's bandwidth bill must be crazy, or am I missing something (maybe they have a peering agreement? heavily caching things?)

Unmetered 10+ gigabit connections were on the order of $1/mbit/mo wholesale over a decade ago when I priced out a custom CDN so for the cost of 100 TB of data transfer out of AWS you could get a 24/7 sustained 10gbit/s (>3 PB per month at 100% utilization). Bandwidth has always been crazy cheap.

Not all connections are created equal. Even some big providers clearly have iffy peering agreements upstream that’ll manifest as terrible performance if you have a widely-geographically-distributed bandwidth-heavy load.

Re: Weird A.I. Yankovic: a cursed deep dive into the world of voice cloning

#59

Earlier quoted context omitted.

That's fairly short. I read about 100 books a year and it includes thousand page tomes like The Count of Monte Cristo.

I always assumed that book to be rather short since it just needs to be a number of sandwiches eaten. 100 books/year. That's an impressive feat regardless the number of pages. Are these downloaded ebooks or physical printed copies of books?

It's mostly audiobooks, I have some ePubs that don't have audiobooks anywhere, such as many Japanese light novel fan (or official) translations into English for example. I can get through them as I can understand audio faster than I can read text, as I play back at 3 to 5x speed.

Re: Weird A.I. Yankovic: a cursed deep dive into the world of voice cloning

#60
post #35

The sampled voices sound neither like Michael Jackson nor Weird Al. A good effort, but a professional impersonator could likely do better on either front.

It sounds like Weird Al trying to be Michael Jackson trying to be Weird Al.

As a non native speaker, it does sound a bit like Michael Jackson imo…
Post reply on HN