Live data from Hacker News

Qwen3-Omni: Native Omni AI model for text, image and video

github.com

51–60 of 152 posts

Re: Qwen3-Omni: Native Omni AI model for text, image and video

#51
post #16
post #13

Earlier quoted context omitted.

Not really, it's a significant place which is why the protest (and hence massacre) was there, so especially for Chinese people (I expect) merely referencing it doesn't so immediately refer to the massacre, they have plenty of other connotations for it. e.g. if something similar happened in Trafalgar Square, I expect it would still be primarily a major square in London to me, not oh my god they must be referring to th…

Not to mention, Tiananmen Square is one of the major tourist destinations in Beijing (similar to National Mall in Washington DC), for both domestic and foreign visitors.

This is true. I also think they've put some real effort into steering the model away from certain topics. If you ask too closely you'll get a response like:

"As an AI assistant, I must remind you that your statements may involve false and potentially illegal information. Please observe the relevant laws and regulations and ask questions in a civilized manner when you speak."

Re: Qwen3-Omni: Native Omni AI model for text, image and video

#52

Earlier quoted context omitted.

If it does in the future, do we just hope it won’t be retroactive? Is this water boiling yet?

ex post facto law is explicitly banned in the US Constitution

Dogs can't play basketball, either, but we've sure been getting dunked on a lot lately.

Re: Qwen3-Omni: Native Omni AI model for text, image and video

#54
The real point of leverage for here is performance/size. Getting traction in the open weights space kinda forces that the models need to innovate on efficiency. This means the open weight models may get leverage that the closed weight ones don't think about.

If we had some aggregated cluster reasoning mechanisms, When would 8x 30B models running on an h100 server out perform in terms of accuracy 1 240B model on the same server.

Re: Qwen3-Omni: Native Omni AI model for text, image and video

#55
post #6

The model weights are 70GB (Hugging Face recently added a file size indicator - see https://huggingface.co/Qwen/Qwen3-Omni-30B-A3B-Instruct/tree... ) so this one is reasonably accessible to run locally. I wonder if we'll see a macOS port soon - currently it very much needs an NVIDIA GPU as far as I can tell.

Would it run on 5090? Or is it possible to link multiple GPUs or has NVIDIA locked it down?

Re: Qwen3-Omni: Native Omni AI model for text, image and video

#56

Interesting, the pacing seemed very slow when conversing in english, but when I spoke to it in spanish, it sounded much faster. It's really impressive that these models are going to be able to do real time translation and much more. The Chinese are going to end up owning the AI market if the American labs don't start competing on open weights. Americans may end up in a situation where they have some $1000-2000 device…

[deleted]

Re: Qwen3-Omni: Native Omni AI model for text, image and video

#57

Interesting, the pacing seemed very slow when conversing in english, but when I spoke to it in spanish, it sounded much faster. It's really impressive that these models are going to be able to do real time translation and much more. The Chinese are going to end up owning the AI market if the American labs don't start competing on open weights. Americans may end up in a situation where they have some $1000-2000 device…

When has the average American ever been willing to spend a $1,000-2,000 premium for privacy-respecting tech? They already save $20-200 to buy IoT cameras which provide all audio and video from inside their home directly to the government without a warrant (Ring vs Reolink/etc).

To be fair, it isn't $1000-2000 extra, it's the new laptop/pc you just bought that is powerful enough (now, or in the near future) to run these open weight models.

Re: Qwen3-Omni: Native Omni AI model for text, image and video

#58
post #6

The model weights are 70GB (Hugging Face recently added a file size indicator - see https://huggingface.co/Qwen/Qwen3-Omni-30B-A3B-Instruct/tree... ) so this one is reasonably accessible to run locally. I wonder if we'll see a macOS port soon - currently it very much needs an NVIDIA GPU as far as I can tell.

is there an inference engine for this on macos?

Not yet as far as I can tell - might take a while for someone to pull that together given the complexity involved in handling audio and image and text and video at once.

Re: Qwen3-Omni: Native Omni AI model for text, image and video

#60

Interesting, the pacing seemed very slow when conversing in english, but when I spoke to it in spanish, it sounded much faster. It's really impressive that these models are going to be able to do real time translation and much more. The Chinese are going to end up owning the AI market if the American labs don't start competing on open weights. Americans may end up in a situation where they have some $1000-2000 device…

There is some irony about buying Chinese hardware to run American software on it for the past decade(s), and now the exact reverse.
Post reply on HN