Live data from Hacker News

Qwen3-Omni: Native Omni AI model for text, image and video

github.com

21–30 of 152 posts

Re: Qwen3-Omni: Native Omni AI model for text, image and video

#21

Earlier quoted context omitted.

Sure but all of these find some way of mapping inputs (any medium) to state space concepts. That's the core of the transformer architecture.

The user you originally replied to specifically mentioned > without going to text first

Yeah, and that's my understanding. Nothing goes video -> text, or audio -> text, or even text -> text without first going through state space. That's where the core of the transformer architecture is.

Re: Qwen3-Omni: Native Omni AI model for text, image and video

#22

The multilingual example in the launch graphic has Qwen3 producing the text: > "Bonjour, pourriez-vous me dire comment se rendreà la place Tian'anmen?" translation: "Hello, could you tell me how to get to Tiananmen Square?" a bold choice!

Westerners only know it from the massacre but it’s actually just like Times Square for them

Re: Qwen3-Omni: Native Omni AI model for text, image and video

#23
post #18

You can try it out on https://chat.qwen.ai/ - sign in with Google or GitHub (signed out users can't use the voice mode) and then click on the voice icon. It has an entertaining selection of different voices, including: *Dylan* - A teenager who grew up in Beijing's hutongs *Peter* - Tianjin crosstalk, professionally supporting others *Cherry* - A sunny, positive, friendly, and natural young lady *Ethan* - A sunny, war…

I only see Omni Flash, is that the one?

Re: Qwen3-Omni: Native Omni AI model for text, image and video

#24
post #19

Speech input + speech output is a big deal. In theory you can talk to it using voice, and it can respond in your language, or translate for someone else, without intermediary technologies. Right now you need wakeword, speech to text, and then text to speech, in addition to your core LLM. A couple can input speech, or output speech, but not both. It looks like they have at least 3 variants in the ~32b range. Depending…

Seems like a big win for language learning, if nothing else. Also seems possible to run locally, especially once the unsloth guys get their hands on it.

Re: Qwen3-Omni: Native Omni AI model for text, image and video

#25
Interesting, the pacing seemed very slow when conversing in english, but when I spoke to it in spanish, it sounded much faster. It's really impressive that these models are going to be able to do real time translation and much more.

The Chinese are going to end up owning the AI market if the American labs don't start competing on open weights. Americans may end up in a situation where they have some $1000-2000 device at home with an open Chinese model running on it, if they care about privacy or owning their data. What a turn of events!

Re: Qwen3-Omni: Native Omni AI model for text, image and video

#26
post #18

You can try it out on https://chat.qwen.ai/ - sign in with Google or GitHub (signed out users can't use the voice mode) and then click on the voice icon. It has an entertaining selection of different voices, including: *Dylan* - A teenager who grew up in Beijing's hutongs *Peter* - Tianjin crosstalk, professionally supporting others *Cherry* - A sunny, positive, friendly, and natural young lady *Ethan* - A sunny, war…

The voices are really fun, thanks for the laughs :)

Re: Qwen3-Omni: Native Omni AI model for text, image and video

#27
post #6

The model weights are 70GB (Hugging Face recently added a file size indicator - see https://huggingface.co/Qwen/Qwen3-Omni-30B-A3B-Instruct/tree... ) so this one is reasonably accessible to run locally. I wonder if we'll see a macOS port soon - currently it very much needs an NVIDIA GPU as far as I can tell.

is there an inference engine for this on macos?

Re: Qwen3-Omni: Native Omni AI model for text, image and video

#28

Interesting, the pacing seemed very slow when conversing in english, but when I spoke to it in spanish, it sounded much faster. It's really impressive that these models are going to be able to do real time translation and much more. The Chinese are going to end up owning the AI market if the American labs don't start competing on open weights. Americans may end up in a situation where they have some $1000-2000 device…

This is exactly what I do. I have two 3090s at home, with Qwen3 on it. This is tied into my Home Assistant install, and I use esp32 devices as voice satellites. It works shockingly well.

Re: Qwen3-Omni: Native Omni AI model for text, image and video

#29

Interesting, the pacing seemed very slow when conversing in english, but when I spoke to it in spanish, it sounded much faster. It's really impressive that these models are going to be able to do real time translation and much more. The Chinese are going to end up owning the AI market if the American labs don't start competing on open weights. Americans may end up in a situation where they have some $1000-2000 device…

Americans may end up in a situation where they have some $1000-2000 device at home with an open Chinese model running on it

Wouldn't worry about that, I'm pretty sure the government is going to ban running Chinese tech in this space sooner or later. And we won't even be able to download it.

Not saying any of the bans will make any kind of sense, but I'm pretty sure they're gonna say this is a "strategic" space. And everything else will follow from there.

Download Chinese models while you can.

Re: Qwen3-Omni: Native Omni AI model for text, image and video

#30
post #29

Interesting, the pacing seemed very slow when conversing in english, but when I spoke to it in spanish, it sounded much faster. It's really impressive that these models are going to be able to do real time translation and much more. The Chinese are going to end up owning the AI market if the American labs don't start competing on open weights. Americans may end up in a situation where they have some $1000-2000 device…

Americans may end up in a situation where they have some $1000-2000 device at home with an open Chinese model running on it Wouldn't worry about that, I'm pretty sure the government is going to ban running Chinese tech in this space sooner or later. And we won't even be able to download it. Not saying any of the bans will make any kind of sense, but I'm pretty sure they're gonna say this is a "strategic" space. And e…

When DeepSeek first hit the news, an American senator proposed adding it to ITAR so they could send people to prison for using it. Didn't pass, thankfully.
Post reply on HN