Live data from Hacker News

Qwen3-Omni: Native Omni AI model for text, image and video

github.com

141–150 of 152 posts

Re: Qwen3-Omni: Native Omni AI model for text, image and video

#141
post #12
post #6

The model weights are 70GB (Hugging Face recently added a file size indicator - see https://huggingface.co/Qwen/Qwen3-Omni-30B-A3B-Instruct/tree... ) so this one is reasonably accessible to run locally. I wonder if we'll see a macOS port soon - currently it very much needs an NVIDIA GPU as far as I can tell.

That's at BF16, so it should fit fairly well on 24GB GPUs after quantization to Q4, I'd think. (Much like the other 30B-A3B models in the family.) I'm pretty happy about that - I was worried it'd be another 200B+.

So like, 1x32GB is all you need for quite a while? Scrolling through the Web makes me feel like I'm out unless I have minimum 128GB of VRAM.

Re: Qwen3-Omni: Native Omni AI model for text, image and video

#142

Earlier quoted context omitted.

> Americans may end up in a situation where they have some $1000-2000 device at home with an open Chinese model running on it, if they care about privacy or owning their data. I think HN vastly overestimates the market for something like this. Yes, there are some people who would spend $2,000 to avoid having prompts go to any cloud service. However, most people don’t care. Paying $20 per month for a ChatGPT subscript…

The reason people will pay $2,000 for a private at home AI is porn.

Given that $2000 might only buy you about 10 date nights with dinner and drinks, the value proposition might actually be pretty good if posterity is not a feature requirement.

Re: Qwen3-Omni: Native Omni AI model for text, image and video

#143
post #63

Earlier quoted context omitted.

When has the average American ever been willing to spend a $1,000-2,000 premium for privacy-respecting tech? They already save $20-200 to buy IoT cameras which provide all audio and video from inside their home directly to the government without a warrant (Ring vs Reolink/etc).

Ease of use is a major issue. What percentage of the people that you know are able to install python and dependencies plus the correct open weights models? I'd wager most of your parents can't do it. Most "normies" wouldn't even know what a local model even is, let alone how to install a GPU.

installing LM Studio is easy and it walks you through choosing a model as well. It is actually well within many people's abilities.

Re: Qwen3-Omni: Native Omni AI model for text, image and video

#144
post #140
post #86

Earlier quoted context omitted.

If there is a walled garden, and you aren't in it, you'll probably push for the walls to come down. No moral basis needed.

Depends on whether you want to be in it. A ladder might be enough to peek over the top and rip it off. Do it better. Which seems to be what is happening.

That only works for China's domestic market. As long as the IP they are "taking inspiration from" is protected in the target markets, they effectively lock themselves out by doing that.

In the case of technology like RISC, pretty much all the value add is unprotected, so they can sell those products in the US/EU without issue.

Re: Qwen3-Omni: Native Omni AI model for text, image and video

#145
post #130

Earlier quoted context omitted.

Ooo interesting, I'd love to hear more about the esp32's as voice satellites!

For the physical hardware I use the esp32-s3-box[1]. The esphome[2] suite has firmware you can flash to make the device work with HomeAssistant automatically. I have an esphome profile[3] I use, but I'm considering switching to this[4] profile instead. For the actual AI, I basically set up three docker containers: one for speech to text[5], one for text to speech[6], and then ollama[7] for the actual AI. After that i…

> 1. https://www.adafruit.com/product/5835

The nails in the video made me laugh

Re: Qwen3-Omni: Native Omni AI model for text, image and video

#146

Earlier quoted context omitted.

> Americans may end up in a situation where they have some $1000-2000 device at home with an open Chinese model running on it, if they care about privacy or owning their data. I think HN vastly overestimates the market for something like this. Yes, there are some people who would spend $2,000 to avoid having prompts go to any cloud service. However, most people don’t care. Paying $20 per month for a ChatGPT subscript…

There is going to be a big market for private AI appliances, in my estimation at least. Case in point: I give Gmail OAuth access to nobody. I nearly got burned once and I really don’t want my entire domain nuked. But I want to be able to have an LLM do things only LLMs can do with my email. “Find all emails with ‘autopay’ in the subject from my utility company for the past 12 months, then compare it to the prior year…

Totally agree on the consumer and SMB play (which is why we're stealthily working on it :). I'm curious what capabilities the next generation of models (and HW) will provide that doesn't exist now. Considering Ryzen 395 / Digits / etc can achieve 40-50+ T/s on capable mid-size models (e.g., OSS120B/Qwen-Next/GLM Air) with some headroom for STT and a lean TTS, I think now is the time to enter but seems to me the 2 key things that are lacking are 1) reliable low-latency multi-modal streaming voice frameworks for STT+STT and 2) reliable fast and secure UI Computer use (without relying on optional accessibility tags/meta).

My greatest concern for local AI solutions like this is the centrality of email and the obvious security concerns surrounding email auth.

Re: Qwen3-Omni: Native Omni AI model for text, image and video

#147
post #99

Earlier quoted context omitted.

Could you elaborate on what you mean by "moral basis" in your comment?

It is in their selfish interest to push for open weights. That's not to say they are being selfish, or to judge in any way the morality of their actions. But because of that incentive, you can't logically infer moral agency in their decision to release open-weights, IP-free CPUs, etc.

By selfish interests you mean the public good?

Re: Qwen3-Omni: Native Omni AI model for text, image and video

#148
post #94
post #59

Earlier quoted context omitted.

I think so, you need to click the big jagged audio icon to start a voice session.

Is the Qwen3-Omni-Flash the same as Qwen3-Omni-30B-A3B, or is the Omni-Flash a different closed-source model?

My question too

Re: Qwen3-Omni: Native Omni AI model for text, image and video

#149
post #94
post #59

Earlier quoted context omitted.

I think so, you need to click the big jagged audio icon to start a voice session.

Is the Qwen3-Omni-Flash the same as Qwen3-Omni-30B-A3B, or is the Omni-Flash a different closed-source model?

In Section 5 of their [technical report](https://arxiv.org/pdf/2509.17765v1) they mention them

"... A comprehensive evaluation was performed on a suite of models, including Qwen3-Omni-30B-A3B- Instruct, Qwen3-Omni-30B-A3B-Thinking, and two in-house developed variants, designated Qwen3- Omni-Flash-Instruct and Qwen3-Omni-Flash-Thinking. These “Flash” models were designed to improve both computational efficiency and performance efficacy, integrating new functionalities, notably the support for various dialects. ..."

Post reply on HN