Live data from Hacker News

Qwen3-TTS family is now open sourced: Voice design, clone, and generation

qwen.ai

161–170 of 229 posts

Re: Qwen3-TTS family is now open sourced: Voice design, clone, and generation

#161

Earlier quoted context omitted.

The demo page only has three samples for Japanese, and one of it pronounces taskete as itsukete (???), so...

Thanks. All modern TTS for Japanese are useless failures.

Just use that green guy, classical TTS ain't broke nanoda.

Re: Qwen3-TTS family is now open sourced: Voice design, clone, and generation

#162
post #133

Earlier quoted context omitted.

Given how easy voice cloning is with this thing I chickened out of sharing the training audio I recorded! That's not really rational considering the internet is full of examples of my voice that anyone could use though. Here's a recent podcast clip: https://www.youtube.com/watch?v=lVDhQMiAbR8&t=3006s

Thanks, so it’s in the [pretty close but still distinguishable] range.

it depends on the medium. A flac will be distinguishable (for now); but put it out over low bandwidth media and you get https://youtube.com/shorts/dpScfg3how8 which, https://www.youtube.com/shorts/zi7BVqVzRx4 is real good and close to what that podcaster's voice sounded like when i made that clone!

i have several other examples from before my repeater ID voice clone. Newer voice models will have to wait till i recover my NAS tomorrow!

this is the newest one i have access to: Dick Powell voice clone off his Richard Diamond Persona: https://soundcloud.com/djoutcold/dick-powell-voice-clone-tes...

i was one-shotting voices years ago that were timbre/tonally identical to the reference voice; however the issue i had was inflection and subtlety. I find that female voices are much easier to clone, or at least it fools my brain into thinking so.

this model, if the results weren't too cherry picked, will be huge improvement!

Re: Qwen3-TTS family is now open sourced: Voice design, clone, and generation

#163
post #33

If you want to try out the voice cloning yourself you can do that an this Hugging Face demo: https://huggingface.co/spaces/Qwen/Qwen3-TTS - switch to the "Voice Clone" tab, paste in some example text and use the microphone option to record yourself reading that text - then paste in other text and have it generate a version of that read using your voice. I shared a recording of audio I generated with that here: https:…

I cloned my voice and had it generate audio for a paragraph from something I wrote. It definitely kind of sounds like me, but I like it much better than listening to my real voice. Some kind of uncanny peak.

Re: Qwen3-TTS family is now open sourced: Voice design, clone, and generation

#164
post #48

Earlier quoted context omitted.

This is terrifying. With this and z-image-turbo, we've crossed a chasm. And a very deep one. We are currently protected by screens, we can, and should assume everything behind a screen is fake unless rigorously (and systematically, i.e. cryptographically) proven otherwise. We're sleepwalking into this, not enough people know about it.

That was my thought too. You’d have “loved ones” calling with their faces and voices asking for money in some emergency. But you’d also have plausible deniability as anything digital can be brushed off as “that’s not evidence, it could be AI generated”.

For now you could ask them to turn away from the camera while keeping their eyes open. If they are a Z-Image they will instantly snap their head to face you.

Re: Qwen3-TTS family is now open sourced: Voice design, clone, and generation

#165
post #33

If you want to try out the voice cloning yourself you can do that an this Hugging Face demo: https://huggingface.co/spaces/Qwen/Qwen3-TTS - switch to the "Voice Clone" tab, paste in some example text and use the microphone option to record yourself reading that text - then paste in other text and have it generate a version of that read using your voice. I shared a recording of audio I generated with that here: https:…

I cloned my voice and had it generate audio for a paragraph from something I wrote. It definitely kind of sounds like me, but I like it much better than listening to my real voice. Some kind of uncanny peak.

They weirdly makes it a canny peak though :)

Re: Qwen3-TTS family is now open sourced: Voice design, clone, and generation

#166

Earlier quoted context omitted.

This feels like one of those tropes that keeps showing up whenever new tech comes out. At the advent of recorded music, im sure buskers and performers were complaing that live music is dead forever. Stage actors were probably complaining that film killed plays. Heck, I bet someome even complained that video itself killed the radio star. Yet here we are, hundreds of years later, live music is still desirable, plays st…

umm, I don't know if you've seen the current state of trying to make a living with music but It's widely accepted as dire. Touring is a loss leader, putting out music for free doesn't pay, stream counts payouts are abysmally low. No one buys songs. All that is before the fact that streaming services are stuffing playlists with AI generated music to further reduce the payouts to artists. > Yet here we are, hundreds of…

Artists were saying this even before streaming, though, much less AI.

I listen pretty exclusively to metal, and a huge chunk of that is bands that are very small. I go to shows where they headliners stick around at the bar and chat with people. Not saying this to be a hipster - I listen to plenty of "mainstream" stuff too - but to show that it's hard to get smaller than this when it comes to people wanting to make a living making music.

None of them made any money off of Spotify or whatever before AI. They probably don't notice a difference, because they never paid attention to the "revenue" there either.

But they do pay attention to Bandcamp. Because Bandcamp has given them more ability to make money off the actual sale of music than they've had in their history - they don't need to rely on a record deal with a big label. They don't need to hope that the small label can somehow get their name out there.

For some genres, some bands, it's more viable than ever before to make a living. For others, yeah, it's getting harder and harder.

Re: Qwen3-TTS family is now open sourced: Voice design, clone, and generation

#167

Earlier quoted context omitted.

This is terrifying. With this and z-image-turbo, we've crossed a chasm. And a very deep one. We are currently protected by screens, we can, and should assume everything behind a screen is fake unless rigorously (and systematically, i.e. cryptographically) proven otherwise. We're sleepwalking into this, not enough people know about it.

> This is terrifying. Far more terrifying is Big Tech having access to a closed version of the same models, in the hands of powerful people with a history of unethical behavior (i.e. Zuckerberg's "Dumb Fucks" comments). In fact it's a miracle and a bit ironic that the Chinese would be the ones to release a plethora of capable open source models, instead of the scraps like we've seen from Google, Meta, OpenAI, etc.

> Far more terrifying is Big Tech having access to a closed version of the same model

Agreed. The only thing worse than everyone having access to this tech is only governments, mega corps and highly-motivated bad actors having access. They've had it a while and there's no putting the genii back in the bottle. The best thing the rest of us can do is use it widely so everyone can adapt to this being the new normal.

Re: Qwen3-TTS family is now open sourced: Voice design, clone, and generation

#168

Earlier quoted context omitted.

I'll pay attention to where the puck is because that is something I can observe, where it is going to be is anybody's guess. Lots of original ideas are coming out of Chinese AI research but there is also lots of junk. I think in the longer term they will have the advantage but right now that simply isn't the case. Your 'cope' accusation has no place here, I have no dog in the race and do not need to cope with anythin…

> Your 'cope' accusation has no place here I will rephrase my statement and continue to stand by it: "Denying the volume of original AI research being done by China - a falsifiable metric - betrays some level of cope." You seem to agree on the fact that China has surpassed the US. As for quality, I'll say expertise is a result of execution. At some point in time during off-shoring, the US had qualitatively better mac…

It may not be cope, could just be ignorance.

Re: Qwen3-TTS family is now open sourced: Voice design, clone, and generation

#169
post #129
post #76

Earlier quoted context omitted.

With a good harness I am getting similar results with GLM 4.7. I am paying for TWO! max accounts and my agents are running 24/7. I still have a small Claude account to do some code reviews. Opus 4.5 does good reviews but at this point GLM 4.7 usually can do the same code reviews. If cost is an issue (for me it is, I pay out of pocket) go with GLM 4.7

Your GitHub profile is... disturbing. 1,354 commits and 464 pull requests in January so far. Regardless of how productive those numbers may seem, that amount of code being published so quickly is concerning, to say the least. It couldn't have possibly been reviewed by a human or properly tested. If this is the future of software development, society is cooked.

It's mostly trying out my orchestration system (https://github.com/mohsen1/claude-code-orchestrator and https://github.com/mohsen1/claude-orchestrator-action) in a repo using GH_PAT.

Stuff like this: https://github.com/mohsen1/claude-code-orchestrator-e2e-test...

Yes, the idea is to really, fully automate software engineering. I don't know if I am going to be successful but I'm on vacation and having fun!

if Opus 4.5/GLM 4.7 can do so much already, I can only imagine what can be done in two years. Might as well adopt to this reality and learn how leverage this advancement

Re: Qwen3-TTS family is now open sourced: Voice design, clone, and generation

#170
post #33

If you want to try out the voice cloning yourself you can do that an this Hugging Face demo: https://huggingface.co/spaces/Qwen/Qwen3-TTS - switch to the "Voice Clone" tab, paste in some example text and use the microphone option to record yourself reading that text - then paste in other text and have it generate a version of that read using your voice. I shared a recording of audio I generated with that here: https:…

This is terrifying. With this and z-image-turbo, we've crossed a chasm. And a very deep one. We are currently protected by screens, we can, and should assume everything behind a screen is fake unless rigorously (and systematically, i.e. cryptographically) proven otherwise. We're sleepwalking into this, not enough people know about it.

I'd be a bit more worried with Z-Image Edit/Base is release. Flux.2 Klein is our and its on par with Zit, and with some fine tuning can just about hit Flux.2. Adding on top of that is Qwen Image Edit 2511 for additional refinement. Anything is possible. Those folks at r/StableDiffusion and falling over the possible release of Z-Image-Omni-Base, a hold me over until actual base is out. I've heard its equal to Flux.2. Crazy time.
Post reply on HN