Live data from Hacker News

Qwen3-Omni: Native Omni AI model for text, image and video

github.com

101–110 of 152 posts

Re: Qwen3-Omni: Native Omni AI model for text, image and video

#101
post #100

Earlier quoted context omitted.

https://www.alibabacloud.com/help/en/model-studio/realtime?s...

Thank you. I'm looking for this as well. The realtime model is a closed-source model and it's different than the open Qwen3-Omni-30B-A3B, right? I wonder how hard is it to turn the open-source model to be a realtime model.

Why do you say they are different models? I've been looking at this today and haven't seen anything explicitly state that.

Re: Qwen3-Omni: Native Omni AI model for text, image and video

#102
post #99
post #86

Earlier quoted context omitted.

If there is a walled garden, and you aren't in it, you'll probably push for the walls to come down. No moral basis needed.

Could you elaborate on what you mean by "moral basis" in your comment?

I mean China's push for open weights/source/architecture probably has more to do with them wanting legal access to markets than it does with those things being morally superior.

Re: Qwen3-Omni: Native Omni AI model for text, image and video

#104

Earlier quoted context omitted.

> Americans may end up in a situation where they have some $1000-2000 device at home with an open Chinese model running on it, if they care about privacy or owning their data. I think HN vastly overestimates the market for something like this. Yes, there are some people who would spend $2,000 to avoid having prompts go to any cloud service. However, most people don’t care. Paying $20 per month for a ChatGPT subscript…

There is going to be a big market for private AI appliances, in my estimation at least. Case in point: I give Gmail OAuth access to nobody. I nearly got burned once and I really don’t want my entire domain nuked. But I want to be able to have an LLM do things only LLMs can do with my email. “Find all emails with ‘autopay’ in the subject from my utility company for the past 12 months, then compare it to the prior year…

How would using oauth through Google nuke ur domain?

Re: Qwen3-Omni: Native Omni AI model for text, image and video

#105
post #100

Earlier quoted context omitted.

Thank you. I'm looking for this as well. The realtime model is a closed-source model and it's different than the open Qwen3-Omni-30B-A3B, right? I wonder how hard is it to turn the open-source model to be a realtime model.

Why do you say they are different models? I've been looking at this today and haven't seen anything explicitly state that.

This is just my assumption given that they listed a lot of different models here: https://modelstudio.console.alibabacloud.com/?spm=a3c0i.2876...

This is an older link, but they listed two different sections here, commercial and open source models: https://modelstudio.console.alibabacloud.com/?tab=doc#/doc/?...

For the realtime multimodal, I'm not seeing the open source models tab: https://modelstudio.console.alibabacloud.com/?tab=doc#/doc/?...

Re: Qwen3-Omni: Native Omni AI model for text, image and video

#107

Interesting, the pacing seemed very slow when conversing in english, but when I spoke to it in spanish, it sounded much faster. It's really impressive that these models are going to be able to do real time translation and much more. The Chinese are going to end up owning the AI market if the American labs don't start competing on open weights. Americans may end up in a situation where they have some $1000-2000 device…

>Interesting, the pacing seemed very slow when conversing in english, but when I spoke to it in spanish, it sounded much faster

So did you run the model offline on your own computer and get realtime audio?

Can you tell me the GPU or specifications you used?

I inquired with ChatGPT:

https://chatgpt.com/share/68d23c2c-2928-800b-bdde-040d8cb40b...

It seems it needs around a $2,500 GPU, do you have one?

I tried Qwen online via its website interface a few months ago, and found it to be very good.

I've run some offline models including Deepseek-R1 70B on CPU (pretty slow, my server has 128 GB of RAM but no GPU) and I'm looking into what kind of setup I would need to run an offline model on GPU myself.

Re: Qwen3-Omni: Native Omni AI model for text, image and video

#109

Earlier quoted context omitted.

There is going to be a big market for private AI appliances, in my estimation at least. Case in point: I give Gmail OAuth access to nobody. I nearly got burned once and I really don’t want my entire domain nuked. But I want to be able to have an LLM do things only LLMs can do with my email. “Find all emails with ‘autopay’ in the subject from my utility company for the past 12 months, then compare it to the prior year…

How would using oauth through Google nuke ur domain?

Depends on the setup, but programmatic access to a Gmail account that's used for admin purposes would allow for hijacking via key/password exfiltration of anything in the mailbox, sending unattended approvals, and autonomous conversations with third parties that aren't on the lookout for impersonation. In the average case, the address book would probably get scraped and the account would be used to blast spam to the rest of the internet.

Moving further, if the OAuth Token confers access to the rest of a user's Google suite, any information in Drive can be compromised. If the token has broader access to a Google Workspace account, there's room for inspecting, modifying, and destroying important information belonging to multiple users. If it's got admin privileges, a third party can start making changes to the org's configuration at large, sending spam from the domain to tank its reputation while earning a quick buck, or engage in phishing on internal users.

The next step would be racking up bills in Google's Cloud, but that's hopefully locked behind a different token. All the same, a bit of lateral movement goes a long way ;)

Re: Qwen3-Omni: Native Omni AI model for text, image and video

#110
post #99
post #86

Earlier quoted context omitted.

If there is a walled garden, and you aren't in it, you'll probably push for the walls to come down. No moral basis needed.

Could you elaborate on what you mean by "moral basis" in your comment?

It is in their selfish interest to push for open weights.

That's not to say they are being selfish, or to judge in any way the morality of their actions. But because of that incentive, you can't logically infer moral agency in their decision to release open-weights, IP-free CPUs, etc.

Post reply on HN