Live data from Hacker News

Qwen3-Omni: Native Omni AI model for text, image and video

github.com

71–80 of 152 posts

Re: Qwen3-Omni: Native Omni AI model for text, image and video

#71
post #6

The model weights are 70GB (Hugging Face recently added a file size indicator - see https://huggingface.co/Qwen/Qwen3-Omni-30B-A3B-Instruct/tree... ) so this one is reasonably accessible to run locally. I wonder if we'll see a macOS port soon - currently it very much needs an NVIDIA GPU as far as I can tell.

A fun project for somebody who has more time than myself would be to see if they can get it working with the new Mojo stuff from yesterday for Apple. I don't know if the functionality would be fully baked out enough yet to actually do the port successfully, but it would be an interesting try.

New Mojo stuff from Apple?

Re: Qwen3-Omni: Native Omni AI model for text, image and video

#72

Earlier quoted context omitted.

A fun project for somebody who has more time than myself would be to see if they can get it working with the new Mojo stuff from yesterday for Apple. I don't know if the functionality would be fully baked out enough yet to actually do the port successfully, but it would be an interesting try.

New Mojo stuff from Apple?

Nvm found it https://news.ycombinator.com/item?id=45326388

Re: Qwen3-Omni: Native Omni AI model for text, image and video

#74

Does this support realtime speech to speech via API? If so, where is this hosted/documented? I wasn’t able to see any info. I’d love to use this in lieu of OAIs (expensive) real time speech to speech offering.

https://www.alibabacloud.com/help/en/model-studio/realtime?s...

Re: Qwen3-Omni: Native Omni AI model for text, image and video

#75
post #63

Earlier quoted context omitted.

When has the average American ever been willing to spend a $1,000-2,000 premium for privacy-respecting tech? They already save $20-200 to buy IoT cameras which provide all audio and video from inside their home directly to the government without a warrant (Ring vs Reolink/etc).

Ease of use is a major issue. What percentage of the people that you know are able to install python and dependencies plus the correct open weights models? I'd wager most of your parents can't do it. Most "normies" wouldn't even know what a local model even is, let alone how to install a GPU.

That sounds like all software. We can make better software.

Re: Qwen3-Omni: Native Omni AI model for text, image and video

#76
post #63

Earlier quoted context omitted.

When has the average American ever been willing to spend a $1,000-2,000 premium for privacy-respecting tech? They already save $20-200 to buy IoT cameras which provide all audio and video from inside their home directly to the government without a warrant (Ring vs Reolink/etc).

Ease of use is a major issue. What percentage of the people that you know are able to install python and dependencies plus the correct open weights models? I'd wager most of your parents can't do it. Most "normies" wouldn't even know what a local model even is, let alone how to install a GPU.

Wiredpancake got flagged to death but they’re right. MacWhisper provides a great example of good value for dead-simple user-friendly on-device processing.

Re: Qwen3-Omni: Native Omni AI model for text, image and video

#77

Interesting, the pacing seemed very slow when conversing in english, but when I spoke to it in spanish, it sounded much faster. It's really impressive that these models are going to be able to do real time translation and much more. The Chinese are going to end up owning the AI market if the American labs don't start competing on open weights. Americans may end up in a situation where they have some $1000-2000 device…

> Americans may end up in a situation where they have some $1000-2000 device at home with an open Chinese model running on it, if they care about privacy or owning their data.

I think HN vastly overestimates the market for something like this. Yes, there are some people who would spend $2,000 to avoid having prompts go to any cloud service.

However, most people don’t care. Paying $20 per month for a ChatGPT subscription is a bargain and they automatically get access to new versions as they come.

I think the at-home self hosting hobby is interesting, but it’s never going to be a mainstream thing.

Re: Qwen3-Omni: Native Omni AI model for text, image and video

#78

Interesting, the pacing seemed very slow when conversing in english, but when I spoke to it in spanish, it sounded much faster. It's really impressive that these models are going to be able to do real time translation and much more. The Chinese are going to end up owning the AI market if the American labs don't start competing on open weights. Americans may end up in a situation where they have some $1000-2000 device…

sitting here in the US, reading that China is strongly urging the adoption of Linux and pushing for open CPU architectures like RISC-V and also self-hosted open models

are we the baddies??

Re: Qwen3-Omni: Native Omni AI model for text, image and video

#79
post #29

Earlier quoted context omitted.

Americans may end up in a situation where they have some $1000-2000 device at home with an open Chinese model running on it Wouldn't worry about that, I'm pretty sure the government is going to ban running Chinese tech in this space sooner or later. And we won't even be able to download it. Not saying any of the bans will make any kind of sense, but I'm pretty sure they're gonna say this is a "strategic" space. And e…

government hardly has the capacity to ban foreign weights

The danger is that lawmakers, confused about the difference between foreign weights and foreign APIs, accidentally ban both.

Re: Qwen3-Omni: Native Omni AI model for text, image and video

#80

Interesting, the pacing seemed very slow when conversing in english, but when I spoke to it in spanish, it sounded much faster. It's really impressive that these models are going to be able to do real time translation and much more. The Chinese are going to end up owning the AI market if the American labs don't start competing on open weights. Americans may end up in a situation where they have some $1000-2000 device…

> Americans may end up in a situation where they have some $1000-2000 device at home with an open Chinese model running on it, if they care about privacy or owning their data. I think HN vastly overestimates the market for something like this. Yes, there are some people who would spend $2,000 to avoid having prompts go to any cloud service. However, most people don’t care. Paying $20 per month for a ChatGPT subscript…

I agree the market is niche atm, but I can't help but disagree with your outlook long term. Self hosted models don't have the problems ChatGPT subscribers are facing with models seemingly performing worse over time, they don't need to worry about usage quotas, they don't need to worry about getting locked out of their services, etc.

All of these things have a dark side, though; but it's likely unnecessary for me to elaborate on that.

Post reply on HN