Live data from Hacker News

Qwen3-Omni: Native Omni AI model for text, image and video

github.com

111–120 of 152 posts

Re: Qwen3-Omni: Native Omni AI model for text, image and video

#111
post #68

I usually ask these models to tell me a short story, and most times the prose is stiff and the story reads like a mass market straight to KDP kids book. But wow, first shot generated something, light, mildly funny, and chill. Quite a surprise. Pasted here for your own judgement: *Title: The Last Lightbulb* The power had been out for three days. Rain drummed against the windows of the old cabin, and the only light cam…

That's not bad actually.

Re: Qwen3-Omni: Native Omni AI model for text, image and video

#114

Earlier quoted context omitted.

I run Home Assistant on an RPi4 and have an ESP32-based Core2 with mic ( https://shop.m5stack.com/products/m5stack-core2-esp32-iot-de... ), along with a 16GB 4070 Ti Super in an always-on Windows system I only use for occasional gaming and serving media. I'd love to set up something like you have. Can you recommended a starting place, or ideally, a step-by-step tutorial? I've never set up any AI system. Would you say…

I mean - the AI itself will help you get all that setup. Claude code is your friend. I run proxmox on an old Dell R710 in my closet that hosts my homeassistant (amongst others) VM and then I've setup my "gaming" PC (which hasn't done any gaming in quite some time) to dual boot (Windows or Deb/Proxmox) and just keep it booted into Deb as another proxmox node. That PC also has a 4070 Super that I have setup to passthru…

[dead]

Re: Qwen3-Omni: Native Omni AI model for text, image and video

#115

Interesting, the pacing seemed very slow when conversing in english, but when I spoke to it in spanish, it sounded much faster. It's really impressive that these models are going to be able to do real time translation and much more. The Chinese are going to end up owning the AI market if the American labs don't start competing on open weights. Americans may end up in a situation where they have some $1000-2000 device…

When has the average American ever been willing to spend a $1,000-2,000 premium for privacy-respecting tech? They already save $20-200 to buy IoT cameras which provide all audio and video from inside their home directly to the government without a warrant (Ring vs Reolink/etc).

You mean like a home with a yard large enough to keep the neighbors out of sight?

Granted, based on how annoyingly chill we are with advertisements and government surveillance, I suppose this desire for privacy never extended beyond the neighbors.

Re: Qwen3-Omni: Native Omni AI model for text, image and video

#116
has anyone figured out how to ask a question with text and have it speak the answer in the app ? i can generate text or talk but not jump between

i was lead to believe that it was possibly by the first image here: https://qwen.ai/blog?id=1f04779964b26eacd0025e68698258faacc7... that shows a voice output (top left) next to the written out detail of the thinking mode

Re: Qwen3-Omni: Native Omni AI model for text, image and video

#117

Interesting, the pacing seemed very slow when conversing in english, but when I spoke to it in spanish, it sounded much faster. It's really impressive that these models are going to be able to do real time translation and much more. The Chinese are going to end up owning the AI market if the American labs don't start competing on open weights. Americans may end up in a situation where they have some $1000-2000 device…

Is there a AI market for open weights? Companies like Alibaba, Tencent, Meta or Microsoft makes a lot sense. They can build on open weights, and not losing values, potentially beneficial for share prices. The only winner is application and cloud providers, I don't see how they can make money from the weights itself to be honest.

Re: Qwen3-Omni: Native Omni AI model for text, image and video

#118

Earlier quoted context omitted.

I run Home Assistant on an RPi4 and have an ESP32-based Core2 with mic ( https://shop.m5stack.com/products/m5stack-core2-esp32-iot-de... ), along with a 16GB 4070 Ti Super in an always-on Windows system I only use for occasional gaming and serving media. I'd love to set up something like you have. Can you recommended a starting place, or ideally, a step-by-step tutorial? I've never set up any AI system. Would you say…

I mean - the AI itself will help you get all that setup. Claude code is your friend. I run proxmox on an old Dell R710 in my closet that hosts my homeassistant (amongst others) VM and then I've setup my "gaming" PC (which hasn't done any gaming in quite some time) to dual boot (Windows or Deb/Proxmox) and just keep it booted into Deb as another proxmox node. That PC also has a 4070 Super that I have setup to passthru…

>use opus (get the max plan)

I dont have max plan, but on the Pro i tried for a month, i was able to blow trough my 5 hour limit by a single prompt (with 70k context codebase attached). The idea of paying so much money to get few questions per "workday" seems insane to me

Re: Qwen3-Omni: Native Omni AI model for text, image and video

#119

Interesting, the pacing seemed very slow when conversing in english, but when I spoke to it in spanish, it sounded much faster. It's really impressive that these models are going to be able to do real time translation and much more. The Chinese are going to end up owning the AI market if the American labs don't start competing on open weights. Americans may end up in a situation where they have some $1000-2000 device…

Is there a AI market for open weights? Companies like Alibaba, Tencent, Meta or Microsoft makes a lot sense. They can build on open weights, and not losing values, potentially beneficial for share prices. The only winner is application and cloud providers, I don't see how they can make money from the weights itself to be honest.

It promotes an open research environment where external researchers have the opportunity to learn, improve and build. And it keeps the big companies in check, they can't become monopolies or duopolies and increase API prices (as is usually the playbook) if you can get the same quality responses from a smaller provider on OpenRouter

Re: Qwen3-Omni: Native Omni AI model for text, image and video

#120
post #99

Earlier quoted context omitted.

Could you elaborate on what you mean by "moral basis" in your comment?

I mean China's push for open weights/source/architecture probably has more to do with them wanting legal access to markets than it does with those things being morally superior.

If by being selfish they end up doing morally superior thing, then, I much prefer to go with the Chinese.

Even more so now that Trump is in command.

Post reply on HN