Earlier quoted context omitted.
do you have the RTF for the 1080? I am trying to figure out if the 0.6B model is viable for real-time inference on edge devices.
Yeah, it's not great. I wrote a harness that calculates it as: 3.61s Load Time, 38.78s Gen Time, 18.38s Audio Len, RTF 2.111. The Tao Te Ching audiobook came in at 62 mins in length and it ran for 102 mins, which gives an RTF of 1.645. I do get a warning about flash-attn not being installed, which says that it'll slow down inference. I'm not sure if that feature can be supported on the 1080 and I wasn't up for tinker…
Qwen3-TTS family is now open sourced: Voice design, clone, and generation
151–160 of 229 posts
Re: Qwen3-TTS family is now open sourced: Voice design, clone, and generation
#152Earlier quoted context omitted.
Have you tried specifying the emotion? There's an option to do so and if it's left empty it wouldn't surprise me if it defaulted to rng instead of bland.
For the system prompt I used: > Read this in a calm, clear, and wise audiobook tone. > Do not rush. Allow the meaning to sink in. But maybe I should experiment with something more detailed. Do you have any suggestions?
Character Name: Marcus Cole Voice Profile: A bright, agile male voice with a natural upward lift, delivering lines at a brisk, energetic pace. Pitch leans high with spark, volume projects clearly—near-shouting at peaks—to convey urgency and excitement. Speech flows seamlessly, fluently, each word sharply defined, riding a current of dynamic rhythm. Background: Longtime broadcast booth announcer for national television, specializing in live interstitials and public engagement spots. His voice bridges segments, rallies action, and keeps momentum alive—from voter drives to entertainment news. Presence: Late 50s, neatly groomed, dressed in a crisp shirt under studio lights. Moves with practiced ease, eyes locked on the script, energy coiled and ready. Personality: Energetic, precise, inherently engaging. He doesn’t just read—he propels. Behind the speed is intent: to inform fast, to move people to act. Whether it’s “text VOTE to 5703” or a star-studded tease, he makes it feel immediate, vital.
Re: Qwen3-TTS family is now open sourced: Voice design, clone, and generation
#153I got this running on macOS using mlx-audio thanks to Prince Canuma: https://x.com/Prince_Canuma/status/2014453857019904423 Here's the script I'm using: https://github.com/simonw/tools/blob/main/python/q3_tts.py You can try it with uv (downloads a 4.5GB model on first run) like this: uv run https://tools.simonwillison.net/python/q3_tts.py \ 'I am a pirate, give me your gold!' \ -i 'gruff voice' -o pirate.wav
hopefully i can make this work on windows (or linux, i guess).
thanks so much.
Re: Qwen3-TTS family is now open sourced: Voice design, clone, and generation
#154Earlier quoted context omitted.
This is terrifying. With this and z-image-turbo, we've crossed a chasm. And a very deep one. We are currently protected by screens, we can, and should assume everything behind a screen is fake unless rigorously (and systematically, i.e. cryptographically) proven otherwise. We're sleepwalking into this, not enough people know about it.
That was my thought too. You’d have “loved ones” calling with their faces and voices asking for money in some emergency. But you’d also have plausible deniability as anything digital can be brushed off as “that’s not evidence, it could be AI generated”.
This won't change anything about Western style courts which have always required an unbroken chain of custody of evidence for evidence to be admissable in court
Re: Qwen3-TTS family is now open sourced: Voice design, clone, and generation
#155I got this running on macOS using mlx-audio thanks to Prince Canuma: https://x.com/Prince_Canuma/status/2014453857019904423 Here's the script I'm using: https://github.com/simonw/tools/blob/main/python/q3_tts.py You can try it with uv (downloads a 4.5GB model on first run) like this: uv run https://tools.simonwillison.net/python/q3_tts.py \ 'I am a pirate, give me your gold!' \ -i 'gruff voice' -o pirate.wav
Simon how do you think this would perform on CPU only? Lets say threadripper with 20G ram. (Voice cloning in particular)
anyhow, with faster CPUs and optimizations, you won't be waiting too long. Also 20GB is overkill for an audio model. Only text - LLM - are huge and take infinite memory. SD/FLUX models are under 16GB of ram usage (uh, mine are, at least!), for instance.
Re: Qwen3-TTS family is now open sourced: Voice design, clone, and generation
#156Earlier quoted context omitted.
do you have the RTF for the 1080? I am trying to figure out if the 0.6B model is viable for real-time inference on edge devices.
Yeah, it's not great. I wrote a harness that calculates it as: 3.61s Load Time, 38.78s Gen Time, 18.38s Audio Len, RTF 2.111. The Tao Te Ching audiobook came in at 62 mins in length and it ran for 102 mins, which gives an RTF of 1.645. I do get a warning about flash-attn not being installed, which says that it'll slow down inference. I'm not sure if that feature can be supported on the 1080 and I wasn't up for tinker…
Re: Qwen3-TTS family is now open sourced: Voice design, clone, and generation
#157Earlier quoted context omitted.
Your GitHub profile is... disturbing. 1,354 commits and 464 pull requests in January so far. Regardless of how productive those numbers may seem, that amount of code being published so quickly is concerning, to say the least. It couldn't have possibly been reviewed by a human or properly tested. If this is the future of software development, society is cooked.
You may not like it but this is what a 10x developer looks like. :-)
Re: Qwen3-TTS family is now open sourced: Voice design, clone, and generation
#158Qwen team, please please please, release something to outperform and surpass the coding abilities of Opus 4.5. Although I like the model, I don't like the leadership of that company and how close it is, how divisive they're in terms of politics.
The Chinese labs distill the SOTA models to boost the performance of theirs. They are a trailer hooked up (with a 3-6 month long chain) to the trucks pushing the technology forwards. I've yet to see a trailer overtake it's truck. China would need an architectural breakthrough to leap American labs given the huge compute disparity.
because i've been on youtube and insta, and believe me, no one else even compares, yet.
Re: Qwen3-TTS family is now open sourced: Voice design, clone, and generation
#159Re: Qwen3-TTS family is now open sourced: Voice design, clone, and generation
#160great news, this looks great! is it just me, or do most of the english audio samples sound like anime voices?
1: https://old.reddit.com/r/ZenlessZoneZero/comments/1gqmtl1/th...