Live data from Hacker News

Qwen3-TTS family is now open sourced: Voice design, clone, and generation

qwen.ai

211–220 of 229 posts

Re: Qwen3-TTS family is now open sourced: Voice design, clone, and generation

#211

Earlier quoted context omitted.

> This is terrifying. Far more terrifying is Big Tech having access to a closed version of the same models, in the hands of powerful people with a history of unethical behavior (i.e. Zuckerberg's "Dumb Fucks" comments). In fact it's a miracle and a bit ironic that the Chinese would be the ones to release a plethora of capable open source models, instead of the scraps like we've seen from Google, Meta, OpenAI, etc.

> Far more terrifying is Big Tech having access to a closed version of the same model Agreed. The only thing worse than everyone having access to this tech is only governments, mega corps and highly-motivated bad actors having access. They've had it a while and there's no putting the genii back in the bottle. The best thing the rest of us can do is use it widely so everyone can adapt to this being the new normal.

I know genii is the plural of genie, but for a second I thought it was a typo of genai and I kind of like that better.

Re: Qwen3-TTS family is now open sourced: Voice design, clone, and generation

#212
post #193

Earlier quoted context omitted.

If i am ever in the same city as you, i'll buy you dinner. I poked around during my free time today trying to figure out how to run these models, and here is the estimable Simon Willison just presenting it on a platter. hopefully i can make this work on windows (or linux, i guess). thanks so much.

> hopefully i can make this work on windows (or linux, i guess). mlx-audio only works on Apple Silicon

The original script supports CPU inference, nonetheless.

Re: Qwen3-TTS family is now open sourced: Voice design, clone, and generation

#214

Earlier quoted context omitted.

This is terrifying. With this and z-image-turbo, we've crossed a chasm. And a very deep one. We are currently protected by screens, we can, and should assume everything behind a screen is fake unless rigorously (and systematically, i.e. cryptographically) proven otherwise. We're sleepwalking into this, not enough people know about it.

> This is terrifying. Far more terrifying is Big Tech having access to a closed version of the same models, in the hands of powerful people with a history of unethical behavior (i.e. Zuckerberg's "Dumb Fucks" comments). In fact it's a miracle and a bit ironic that the Chinese would be the ones to release a plethora of capable open source models, instead of the scraps like we've seen from Google, Meta, OpenAI, etc.

The really terrifying thing is the next logical step from the instinctual reaction. Eschew miracle, eschew the cognitive bias of feeling warm and fuzzy for the guy who gives you it for free.

Socratic version: how can the Chinese companies afford to make them and give them out for free? Cui bono?

n.b. it's not because they're making money on the API, ex. open openrouter and see how Moonshot or DeepSeek's 1st party inference speed compares to literally any other provider. Note also that this disadvantage can't just be limited to LLMs, due to GPU export rules.

Re: Qwen3-TTS family is now open sourced: Voice design, clone, and generation

#215

Has anyone successfully run this on a Mac? The installation instructions appear to assume an NVIDIA GPU (CUDA, FlashAttention), and I’m not sure whether it works with PyTorch’s Metal/MPS backend.

Yes, using mlx-audio. See https://news.ycombinator.com/item?id=46726440

Re: Qwen3-TTS family is now open sourced: Voice design, clone, and generation

#216
Honestly, this seems like it could be pretty cool for video games. I always liked Oblivion's 'Radiant AI', this could be a natural progression, give characters motivations, relations with the player and other NPCs and have an LLM spit out background dialogue, then have another model generate the audio.

Re: Qwen3-TTS family is now open sourced: Voice design, clone, and generation

#217

Earlier quoted context omitted.

Indeed, I have a future project/goal of "restoring" Have Gun - Will Travel radio episodes to listenable quality using tech like this. There are so many lines where sound effects and tape rot and other "bad recording" things make it very difficult to understand what was sad. Will be amazing, but as with all tech the potential for abuse is very real

hey if you want to collab or trade notes, my email is in my profile. there was java software that did FANTASTIC work cleaning up crappy transfers of audio, like, specifically, it was perfect for "AM Quality Monaural Audio". Observe, original: https://www.youtube.com/watch?v=YiRcOVDAryM my edit (took about an hour, if memory serves, to set up. forgot render time...): https://www.youtube.com/watch?v=xazubVJ0jz4 i say "…

Neat! That's really cool. I'll definitely reach out once I'm ready to move forward on it. Got a few high-priority things sucking up all my free time at the moment :-(

Yeah all my radio plays are from OTRR now. I bought a number of different "collections" from different sources but none of them come even close to the quality and care that the OTRR people have.

Also, always a pleasure to meet someone else who loves old-time radio :-D

What are some of your favorites? Probably my favorite is Abbott & Costello, followed by Have Gun - Will Travel and Gunsmoke. I like the Lone Ranger too but am only a few hours into it so far.

p.s. I am indeed a dude named Ben!

Re: Qwen3-TTS family is now open sourced: Voice design, clone, and generation

#218

Earlier quoted context omitted.

This feels like one of those tropes that keeps showing up whenever new tech comes out. At the advent of recorded music, im sure buskers and performers were complaing that live music is dead forever. Stage actors were probably complaining that film killed plays. Heck, I bet someome even complained that video itself killed the radio star. Yet here we are, hundreds of years later, live music is still desirable, plays st…

umm, I don't know if you've seen the current state of trying to make a living with music but It's widely accepted as dire. Touring is a loss leader, putting out music for free doesn't pay, stream counts payouts are abysmally low. No one buys songs. All that is before the fact that streaming services are stuffing playlists with AI generated music to further reduce the payouts to artists. > Yet here we are, hundreds of…

Is it though? Think about being a musician 200 years ago. In 1826 you needed to essentially be nobility or nobility-adjacent just to be able to touch an instrument let alone make a living from it. 100 years later, 1926 the barrier to entry was still sky high, nobody could make and distribute recordings without extensive investment. Nowadays it's not uncommon for a 17 year old to download some free composer software, sign up for a few accounts and distribute their music to an audience of millions. It's not easy to do, sure, but there is still opportunity that never existed. If you were to take at random a 20 year old from the general population in 1826, 1923, 1943, 1953, 1973, 83, etc, would you REALLY say that any of them have a BETTER opportunity than today?

Re: Qwen3-TTS family is now open sourced: Voice design, clone, and generation

#219

Earlier quoted context omitted.

It's mostly trying out my orchestration system ( https://github.com/mohsen1/claude-code-orchestrator and https://github.com/mohsen1/claude-orchestrator-action ) in a repo using GH_PAT. Stuff like this: https://github.com/mohsen1/claude-code-orchestrator-e2e-test... Yes, the idea is to really, fully automate software engineering. I don't know if I am going to be successful but I'm on vacation and having fun! if Opus 4…

On the contrary, that actually is pretty cool. z.ai subscription is cheap enough that I'm thinking to run it 24/7 too. Curious if you've tried any other AI orchestration tools like Gas Town? What made you decide to build your own, and how is it working for you so far?

I didn't know about Gas Town! Super cool! I will try it once I have a chance. I started with a few dumb Tmux based scripts and eventually I figured I make it into a proper package.

I think using GitHub with issues,PRs and specially leveraging AI code reviewers like Greptile is the way to go Actually. I did an attempt here https://github.com/mohsen1/claude-orchestrator-action but I think it needs a lot more attention to get it right. Ideas in Gas Town are great and I might steal some of those. Running Claude Code in GitHub Action works with GLM 4.7 great.

Microsoft's new Agent SDK is also interesting. Unlocks multi-provider workflows so user can burn out all of their subscriptions or quickly switch providers

Also super interested in collaborating with someone to build something together if you are interested!

Re: Qwen3-TTS family is now open sourced: Voice design, clone, and generation

#220
post #172

Earlier quoted context omitted.

> as anything digital can be brushed off as “that’s not evidence, it could be AI generated”. This won't change anything about Western style courts which have always required an unbroken chain of custody of evidence for evidence to be admissable in court

Court account for a vanishingly small proportion of most people's lives.

So does the presentation of evidence...
Post reply on HN