Qwen3-TTS family is now open sourced: Voice design, clone, and generation
101–110 of 229 posts
Re: Qwen3-TTS family is now open sourced: Voice design, clone, and generation
#102If you want to try out the voice cloning yourself you can do that an this Hugging Face demo: https://huggingface.co/spaces/Qwen/Qwen3-TTS - switch to the "Voice Clone" tab, paste in some example text and use the microphone option to record yourself reading that text - then paste in other text and have it generate a version of that read using your voice. I shared a recording of audio I generated with that here: https:…
Re: Qwen3-TTS family is now open sourced: Voice design, clone, and generation
#103Earlier quoted context omitted.
There are plenty of electronic artists who can't sing. Right now they have to hire someone else to do the singing for them, but I'd wager a lot of them would like to own their music end-to-end. I would. I'm a filmmaker. I've done it photons-on-glass production for fifteen years. Meisner trained, have performed every role from cast to crew. I'm elated that these tools are going to enable me to do more with a smaller b…
Yes, the flipside of this is that we're eroding the last bit of ability for people to make a living through their art. We are capturing the market for people to live off of making illustrations, to making background music, jingles, promotional videos, photographs, graphic design, and funnelling those earnings to NVIDIA. The question I keep asking is whether we care to value as a society for people to make a living th…
Re: Qwen3-TTS family is now open sourced: Voice design, clone, and generation
#104If you want to try out the voice cloning yourself you can do that an this Hugging Face demo: https://huggingface.co/spaces/Qwen/Qwen3-TTS - switch to the "Voice Clone" tab, paste in some example text and use the microphone option to record yourself reading that text - then paste in other text and have it generate a version of that read using your voice. I shared a recording of audio I generated with that here: https:…
The HF demo space was overloaded, but I got the demo working locally easily enough. The voice cloning of the 1.7B model captures the tone of the speaker very well, but I found it failed at reproducing the variation in intonation, so it sounds like a monotonous reading of a boring text. I presume this is due to using the base model, and not the one tuned for more expressiveness. edit: Or more likely, the demo not expo…
Re: Qwen3-TTS family is now open sourced: Voice design, clone, and generation
#105Any recommendations for an iOS app to test models like this? There are a few good ones for text gen, and it’s a great way to try models
Re: Qwen3-TTS family is now open sourced: Voice design, clone, and generation
#106Earlier quoted context omitted.
Yes, the flipside of this is that we're eroding the last bit of ability for people to make a living through their art. We are capturing the market for people to live off of making illustrations, to making background music, jingles, promotional videos, photographs, graphic design, and funnelling those earnings to NVIDIA. The question I keep asking is whether we care to value as a society for people to make a living th…
This feels like one of those tropes that keeps showing up whenever new tech comes out. At the advent of recorded music, im sure buskers and performers were complaing that live music is dead forever. Stage actors were probably complaining that film killed plays. Heck, I bet someome even complained that video itself killed the radio star. Yet here we are, hundreds of years later, live music is still desirable, plays st…
All that is before the fact that streaming services are stuffing playlists with AI generated music to further reduce the payouts to artists.
> Yet here we are, hundreds of years later, live music is still desirable, plays still happen, and faceless voices are still around...
Yes all those things still happen, but it's increasingly untenable to make a living through it.
Re: Qwen3-TTS family is now open sourced: Voice design, clone, and generation
#107If you want to try out the voice cloning yourself you can do that an this Hugging Face demo: https://huggingface.co/spaces/Qwen/Qwen3-TTS - switch to the "Voice Clone" tab, paste in some example text and use the microphone option to record yourself reading that text - then paste in other text and have it generate a version of that read using your voice. I shared a recording of audio I generated with that here: https:…
This is terrifying. With this and z-image-turbo, we've crossed a chasm. And a very deep one. We are currently protected by screens, we can, and should assume everything behind a screen is fake unless rigorously (and systematically, i.e. cryptographically) proven otherwise. We're sleepwalking into this, not enough people know about it.
Re: Qwen3-TTS family is now open sourced: Voice design, clone, and generation
#108Earlier quoted context omitted.
This is terrifying. With this and z-image-turbo, we've crossed a chasm. And a very deep one. We are currently protected by screens, we can, and should assume everything behind a screen is fake unless rigorously (and systematically, i.e. cryptographically) proven otherwise. We're sleepwalking into this, not enough people know about it.
Admittedly I have not dove into it much but, I wonder if we might finally have a usecase for NFTs and web3? We need some sort of way to denote items are persion generated not AI. Would certainly be easier than trying to determine if something is AI generated
Re: Qwen3-TTS family is now open sourced: Voice design, clone, and generation
#109How does the cloning compare to pocket TTS?
Re: Qwen3-TTS family is now open sourced: Voice design, clone, and generation
#110Qwen team, please please please, release something to outperform and surpass the coding abilities of Opus 4.5. Although I like the model, I don't like the leadership of that company and how close it is, how divisive they're in terms of politics.
They were just waiting for someone in the comments to ask!