Earlier quoted context omitted.
This is terrifying. With this and z-image-turbo, we've crossed a chasm. And a very deep one. We are currently protected by screens, we can, and should assume everything behind a screen is fake unless rigorously (and systematically, i.e. cryptographically) proven otherwise. We're sleepwalking into this, not enough people know about it.
https://www.youtube.com/watch?v=diboERFAjkE pretty much this
Qwen3-TTS family is now open sourced: Voice design, clone, and generation
121–130 of 229 posts
Re: Qwen3-TTS family is now open sourced: Voice design, clone, and generation
#122Haha something that I want to try out. I have started using voice input more and more instead of typing and now I am on my second app and second TTS model, namely Handy and Parakeet V3. Parakeet is pretty good, but there are times it struggles. Would be interesting to see how Qwen compares once Handy has it in.
Re: Qwen3-TTS family is now open sourced: Voice design, clone, and generation
#123Huh. One of the English Voice Clone examples features Obama.
Re: Qwen3-TTS family is now open sourced: Voice design, clone, and generation
#124Re: Qwen3-TTS family is now open sourced: Voice design, clone, and generation
#125If you want to try out the voice cloning yourself you can do that an this Hugging Face demo: https://huggingface.co/spaces/Qwen/Qwen3-TTS - switch to the "Voice Clone" tab, paste in some example text and use the microphone option to record yourself reading that text - then paste in other text and have it generate a version of that read using your voice. I shared a recording of audio I generated with that here: https:…
The HF demo space was overloaded, but I got the demo working locally easily enough. The voice cloning of the 1.7B model captures the tone of the speaker very well, but I found it failed at reproducing the variation in intonation, so it sounds like a monotonous reading of a boring text. I presume this is due to using the base model, and not the one tuned for more expressiveness. edit: Or more likely, the demo not expo…
Re: Qwen3-TTS family is now open sourced: Voice design, clone, and generation
#126If you want to try out the voice cloning yourself you can do that an this Hugging Face demo: https://huggingface.co/spaces/Qwen/Qwen3-TTS - switch to the "Voice Clone" tab, paste in some example text and use the microphone option to record yourself reading that text - then paste in other text and have it generate a version of that read using your voice. I shared a recording of audio I generated with that here: https:…
Re: Qwen3-TTS family is now open sourced: Voice design, clone, and generation
#127Interesting model, I've managed to get the 0.6B param model running on my old 1080 and I can generated 200 character chunks safely without going OOM, so I thought that making an audiobook of the Tao Te Ching would be a good test. Unfortunately each snippet varies drastically in quality: sometimes the speaker is clear and coherent, but other times it bursts out laughing or moaning. In a way it feels a bit like magical…
Re: Qwen3-TTS family is now open sourced: Voice design, clone, and generation
#128Re: Qwen3-TTS family is now open sourced: Voice design, clone, and generation
#129Qwen team, please please please, release something to outperform and surpass the coding abilities of Opus 4.5. Although I like the model, I don't like the leadership of that company and how close it is, how divisive they're in terms of politics.
With a good harness I am getting similar results with GLM 4.7. I am paying for TWO! max accounts and my agents are running 24/7. I still have a small Claude account to do some code reviews. Opus 4.5 does good reviews but at this point GLM 4.7 usually can do the same code reviews. If cost is an issue (for me it is, I pay out of pocket) go with GLM 4.7
Regardless of how productive those numbers may seem, that amount of code being published so quickly is concerning, to say the least. It couldn't have possibly been reviewed by a human or properly tested.
If this is the future of software development, society is cooked.
Re: Qwen3-TTS family is now open sourced: Voice design, clone, and generation
#130Earlier quoted context omitted.
Same issue (I am Danish). Have you tested alternatives? I grabbed Open Code and a Minimax m2.1 subscription, even just the 10usd/mo one to test with. Result? We designed a spec for a slight variation of a tool for which I wrote a spec with Claude - same problem (process supervisor tool), from scratch. Honestly, it worked great, I have played a little further with generating code (this time golang), again, I am happy.…
I've been using GLM 4.7 with Claude Code. best of both worlds. Canceled my Anthropic subscription due to the US politics as well. Already started my "withdrawal" in Jan 2025, Anthropic was one of the few that was left