Live data from Hacker News

Qwen3-TTS family is now open sourced: Voice design, clone, and generation

qwen.ai

221–229 of 229 posts

Re: Qwen3-TTS family is now open sourced: Voice design, clone, and generation

#221

Has anyone successfully run this on a Mac? The installation instructions appear to assume an NVIDIA GPU (CUDA, FlashAttention), and I’m not sure whether it works with PyTorch’s Metal/MPS backend.

Yes, using mlx-audio. See https://news.ycombinator.com/item?id=46726440

Thanks! Simon's example uses the custom voice model (creating a voice from instructions). But that comment led me eventually to this page, which shows how to use mlx-audio for custom voices:

https://huggingface.co/mlx-community/Qwen3-TTS-12Hz-0.6B-Bas...

  uv tool install --force git+https://github.com/Blaizzy/mlx-audio.git --prerelease=allow
    
  python -m mlx_audio.tts.generate --model mlx-community/Qwen3-TTS-12Hz-0.6B-Base-bf16 --text "Hello, this is a test." --ref_audio path_to_audio.wav --ref_text "Transcript of the reference audio." --play

Re: Qwen3-TTS family is now open sourced: Voice design, clone, and generation

#222
post #131

I got this running on macOS using mlx-audio thanks to Prince Canuma: https://x.com/Prince_Canuma/status/2014453857019904423 Here's the script I'm using: https://github.com/simonw/tools/blob/main/python/q3_tts.py You can try it with uv (downloads a 4.5GB model on first run) like this: uv run https://tools.simonwillison.net/python/q3_tts.py \ 'I am a pirate, give me your gold!' \ -i 'gruff voice' -o pirate.wav

If you want to do custom voice cloning, record a sample wav file with a sentence or two, and then try this:

  uv tool install --force git+https://github.com/Blaizzy/mlx-audio.git --prerelease=allow
    
  python -m mlx_audio.tts.generate --model mlx-community/Qwen3-TTS-12Hz-0.6B-Base-bf16 --text "Hello, this is a test." --ref_audio path_to_audio.wav --ref_text "Transcript of the reference audio." --play

Re: Qwen3-TTS family is now open sourced: Voice design, clone, and generation

#224
post #192

Earlier quoted context omitted.

You could pipe the output to an audio file with ffmpeg or pyaudio and save it locally

I don't want to save the audio. I want to save the voice model so I can use it for many different texts, for consistency.

Yes, you can. I was just testing it. I made a "My Custom Voices" tab, and recorded a small sample of my own voice or upload a sample of w/e voice. Then you can use it. I am in the process of training a model of my voice too to see how it handles it using the 1.7b

Works surprisingly good with a 4090. I will also try it on 5090. This is the best one I have seen so far. NGL. 11Labs is cooked lol.

Re: Qwen3-TTS family is now open sourced: Voice design, clone, and generation

#225
post #33

If you want to try out the voice cloning yourself you can do that an this Hugging Face demo: https://huggingface.co/spaces/Qwen/Qwen3-TTS - switch to the "Voice Clone" tab, paste in some example text and use the microphone option to record yourself reading that text - then paste in other text and have it generate a version of that read using your voice. I shared a recording of audio I generated with that here: https:…

well that isnt concerning at all

Re: Qwen3-TTS family is now open sourced: Voice design, clone, and generation

#226
post #179

Earlier quoted context omitted.

Are you using an API proxy to route GLM into the Claude Code CLI? Or do you mean side-by-side usage? Not sure if custom endpoints are supported natively yet.

This works: $ZAI_ANTHROPIC_BASE_URL=xxx $ZAI_ANTHROPIC_AUTH_TOKEN=xxx alias "claude-zai"="ANTHROPIC_BASE_URL=$ZAI_ANTHROPIC_BASE_URL ANTHROPIC_AUTH_TOKEN=$ZAI_ANTHROPIC_AUTH_TOKEN claude" Then you can run `claude`, hit your limit, exit the session and `claude-zai -c` to continue (with context reset, of course). Someone gave me that command a while back.

thats pretty much what I do, I have a bash alias to launch either the normal claude code, or the glm one

Re: Qwen3-TTS family is now open sourced: Voice design, clone, and generation

#227

Earlier quoted context omitted.

> This is terrifying. Far more terrifying is Big Tech having access to a closed version of the same models, in the hands of powerful people with a history of unethical behavior (i.e. Zuckerberg's "Dumb Fucks" comments). In fact it's a miracle and a bit ironic that the Chinese would be the ones to release a plethora of capable open source models, instead of the scraps like we've seen from Google, Meta, OpenAI, etc.

>Far more terrifying is Big Tech having access to a closed version of the same models, in the hands of powerful people with a history of unethical behavior (i.e. Zuckerberg's "Dumb Fucks" comments). Lol what exactly do you think Zuck would do with your voice, drain your bank account??

More likely sell your family ads while using your voice.

Re: Qwen3-TTS family is now open sourced: Voice design, clone, and generation

#228
post #48

Earlier quoted context omitted.

That was my thought too. You’d have “loved ones” calling with their faces and voices asking for money in some emergency. But you’d also have plausible deniability as anything digital can be brushed off as “that’s not evidence, it could be AI generated”.

Only if you focus on the form instead of the content. For a long time my family has had secret words and phrases we use to identify ourselves to each other over secure, but unauthenticated, channels (i.e. the channel is encrypted, but the source is unknown). The military has had to deal with this for some time, and developed various form of IFF that allies could use to identify themselves. E.g. for returning aircraft…

That sounds way too complicated. I get around that by just not having any family any more.

Re: Qwen3-TTS family is now open sourced: Voice design, clone, and generation

#229

Earlier quoted context omitted.

hey if you want to collab or trade notes, my email is in my profile. there was java software that did FANTASTIC work cleaning up crappy transfers of audio, like, specifically, it was perfect for "AM Quality Monaural Audio". Observe, original: https://www.youtube.com/watch?v=YiRcOVDAryM my edit (took about an hour, if memory serves, to set up. forgot render time...): https://www.youtube.com/watch?v=xazubVJ0jz4 i say "…

Neat! That's really cool. I'll definitely reach out once I'm ready to move forward on it. Got a few high-priority things sucking up all my free time at the moment :-( Yeah all my radio plays are from OTRR now. I bought a number of different "collections" from different sources but none of them come even close to the quality and care that the OTRR people have. Also, always a pleasure to meet someone else who loves old…

As far as old time radio: Yours Truly, Johnny Dollar; Richard Diamond (and Rogue's Gallery, same star and crew), and Philip Marlow. I try to get back in to the ones i listened to as a kid 40 odd years ago like Dragnet and Broadway is My Beat, but there's something too slow about them for me now. I like Suspense, as well. I'd have to go digging for some other shows i've enjoyed, as the YT,JD set is hundreds and hundreds of episodes long and keeps me entertained on long car trips...

I also have some newer things, i'm trying to fill my "Coast-to-Coast AM" collection, i've started on Phil Hendrie, and i have most of a show called "Love Line" hosted by Dr Drew Pinsky (and others, usually Adam Carolla). That one i asked permission to clone an archive i found on accident, and the archive owner/host was glad i was doing it. Those are all transcribed, now.

Each is kind of a snapshot of the time they existed, and i don't necessarily want to listen to them all (i don't like Art Bell that much, or talk radio in general!)

And if the reference to "Dude Named Ben" is correct, ITM, and i hope to hear from you when we both have time to fix these old shows!

Post reply on HN