The model weights are 70GB (Hugging Face recently added a file size indicator - see https://huggingface.co/Qwen/Qwen3-Omni-30B-A3B-Instruct/tree... ) so this one is reasonably accessible to run locally. I wonder if we'll see a macOS port soon - currently it very much needs an NVIDIA GPU as far as I can tell.
That's at BF16, so it should fit fairly well on 24GB GPUs after quantization to Q4, I'd think. (Much like the other 30B-A3B models in the family.) I'm pretty happy about that - I was worried it'd be another 200B+.
Qwen3-Omni: Native Omni AI model for text, image and video
141–150 of 152 posts
Re: Qwen3-Omni: Native Omni AI model for text, image and video
#142Earlier quoted context omitted.
> Americans may end up in a situation where they have some $1000-2000 device at home with an open Chinese model running on it, if they care about privacy or owning their data. I think HN vastly overestimates the market for something like this. Yes, there are some people who would spend $2,000 to avoid having prompts go to any cloud service. However, most people don’t care. Paying $20 per month for a ChatGPT subscript…
The reason people will pay $2,000 for a private at home AI is porn.
Re: Qwen3-Omni: Native Omni AI model for text, image and video
#143Earlier quoted context omitted.
When has the average American ever been willing to spend a $1,000-2,000 premium for privacy-respecting tech? They already save $20-200 to buy IoT cameras which provide all audio and video from inside their home directly to the government without a warrant (Ring vs Reolink/etc).
Ease of use is a major issue. What percentage of the people that you know are able to install python and dependencies plus the correct open weights models? I'd wager most of your parents can't do it. Most "normies" wouldn't even know what a local model even is, let alone how to install a GPU.
Re: Qwen3-Omni: Native Omni AI model for text, image and video
#144Earlier quoted context omitted.
If there is a walled garden, and you aren't in it, you'll probably push for the walls to come down. No moral basis needed.
Depends on whether you want to be in it. A ladder might be enough to peek over the top and rip it off. Do it better. Which seems to be what is happening.
In the case of technology like RISC, pretty much all the value add is unprotected, so they can sell those products in the US/EU without issue.
Re: Qwen3-Omni: Native Omni AI model for text, image and video
#145Earlier quoted context omitted.
Ooo interesting, I'd love to hear more about the esp32's as voice satellites!
For the physical hardware I use the esp32-s3-box[1]. The esphome[2] suite has firmware you can flash to make the device work with HomeAssistant automatically. I have an esphome profile[3] I use, but I'm considering switching to this[4] profile instead. For the actual AI, I basically set up three docker containers: one for speech to text[5], one for text to speech[6], and then ollama[7] for the actual AI. After that i…
The nails in the video made me laugh
Re: Qwen3-Omni: Native Omni AI model for text, image and video
#146Earlier quoted context omitted.
> Americans may end up in a situation where they have some $1000-2000 device at home with an open Chinese model running on it, if they care about privacy or owning their data. I think HN vastly overestimates the market for something like this. Yes, there are some people who would spend $2,000 to avoid having prompts go to any cloud service. However, most people don’t care. Paying $20 per month for a ChatGPT subscript…
There is going to be a big market for private AI appliances, in my estimation at least. Case in point: I give Gmail OAuth access to nobody. I nearly got burned once and I really don’t want my entire domain nuked. But I want to be able to have an LLM do things only LLMs can do with my email. “Find all emails with ‘autopay’ in the subject from my utility company for the past 12 months, then compare it to the prior year…
My greatest concern for local AI solutions like this is the centrality of email and the obvious security concerns surrounding email auth.
Re: Qwen3-Omni: Native Omni AI model for text, image and video
#147Earlier quoted context omitted.
Could you elaborate on what you mean by "moral basis" in your comment?
It is in their selfish interest to push for open weights. That's not to say they are being selfish, or to judge in any way the morality of their actions. But because of that incentive, you can't logically infer moral agency in their decision to release open-weights, IP-free CPUs, etc.
Re: Qwen3-Omni: Native Omni AI model for text, image and video
#148Re: Qwen3-Omni: Native Omni AI model for text, image and video
#149Earlier quoted context omitted.
I think so, you need to click the big jagged audio icon to start a voice session.
Is the Qwen3-Omni-Flash the same as Qwen3-Omni-30B-A3B, or is the Omni-Flash a different closed-source model?
"... A comprehensive evaluation was performed on a suite of models, including Qwen3-Omni-30B-A3B- Instruct, Qwen3-Omni-30B-A3B-Thinking, and two in-house developed variants, designated Qwen3- Omni-Flash-Instruct and Qwen3-Omni-Flash-Thinking. These “Flash” models were designed to improve both computational efficiency and performance efficacy, integrating new functionalities, notably the support for various dialects. ..."
Re: Qwen3-Omni: Native Omni AI model for text, image and video
#150https://x.com/whowillrickwill/status/1920723985311903767