Qwen3-Omni: Native Omni AI model for text, image and video
1–10 of 152 posts
Re: Qwen3-Omni: Native Omni AI model for text, image and video
#2Re: Qwen3-Omni: Native Omni AI model for text, image and video
#3Re: Qwen3-Omni: Native Omni AI model for text, image and video
#4> "Bonjour, pourriez-vous me dire comment se rendreà la place Tian'anmen?"
translation: "Hello, could you tell me how to get to Tiananmen Square?"
a bold choice!
Re: Qwen3-Omni: Native Omni AI model for text, image and video
#5The qwen thinker/speaker architecture is really fascinating and is more in line with how I imagine human multi modality works - IE, a picture of an apple, the text a p p l e and the sound all map to the same concept without going to text first.
Re: Qwen3-Omni: Native Omni AI model for text, image and video
#6I wonder if we'll see a macOS port soon - currently it very much needs an NVIDIA GPU as far as I can tell.
Re: Qwen3-Omni: Native Omni AI model for text, image and video
#7Re: Qwen3-Omni: Native Omni AI model for text, image and video
#8The qwen thinker/speaker architecture is really fascinating and is more in line with how I imagine human multi modality works - IE, a picture of an apple, the text a p p l e and the sound all map to the same concept without going to text first.
Isn’t that how all LLMs work?
Multi-modal audio models are a lot less common. GPT-4o was meant to be able to do this natively from the start but they ended up shipping separate custom models based on it for their audio features. As far as I can tell GPT-5 doesn't have audio input/output at all - the OpenAI features for that still use GPT-4o-audio.
I don't know if Gemini 2.5 (which is multi-modal for vision and audio) shares the same embedding space for all three, but I expect it probably does.
Re: Qwen3-Omni: Native Omni AI model for text, image and video
#9Re: Qwen3-Omni: Native Omni AI model for text, image and video
#10The model weights are 70GB (Hugging Face recently added a file size indicator - see https://huggingface.co/Qwen/Qwen3-Omni-30B-A3B-Instruct/tree... ) so this one is reasonably accessible to run locally. I wonder if we'll see a macOS port soon - currently it very much needs an NVIDIA GPU as far as I can tell.