Live data from Hacker News

Qwen3-Omni: Native Omni AI model for text, image and video

github.com

41–50 of 152 posts

Re: Qwen3-Omni: Native Omni AI model for text, image and video

#41

I recently needed to scan hundreds of low quality invoices and run them through OCR for invoice numbers and dates. I really took for granted how seamless this is in some applications, and was shocked how much work went into producing decent results. I was obviously really naive. Either way, it gets me excited any time I see progress with OCR. I should give this a try against my (small) dataset.

I don't think I understand your comment. What were your results for Qwen? Or is that what you meant for how much work was needed?

Re: Qwen3-Omni: Native Omni AI model for text, image and video

#42

Interesting, the pacing seemed very slow when conversing in english, but when I spoke to it in spanish, it sounded much faster. It's really impressive that these models are going to be able to do real time translation and much more. The Chinese are going to end up owning the AI market if the American labs don't start competing on open weights. Americans may end up in a situation where they have some $1000-2000 device…

[deleted]

Re: Qwen3-Omni: Native Omni AI model for text, image and video

#43
post #28

Interesting, the pacing seemed very slow when conversing in english, but when I spoke to it in spanish, it sounded much faster. It's really impressive that these models are going to be able to do real time translation and much more. The Chinese are going to end up owning the AI market if the American labs don't start competing on open weights. Americans may end up in a situation where they have some $1000-2000 device…

This is exactly what I do. I have two 3090s at home, with Qwen3 on it. This is tied into my Home Assistant install, and I use esp32 devices as voice satellites. It works shockingly well.

omg, this is something I've had in mind for quite some time, I even bought some i2s devices to test it out. Do you have some pointers on how to do it?

Re: Qwen3-Omni: Native Omni AI model for text, image and video

#44

I recently needed to scan hundreds of low quality invoices and run them through OCR for invoice numbers and dates. I really took for granted how seamless this is in some applications, and was shocked how much work went into producing decent results. I was obviously really naive. Either way, it gets me excited any time I see progress with OCR. I should give this a try against my (small) dataset.

What is your point about Qwen? Or is it just a general statement regarding LLM?

Re: Qwen3-Omni: Native Omni AI model for text, image and video

#45
post #30
post #29

Earlier quoted context omitted.

Americans may end up in a situation where they have some $1000-2000 device at home with an open Chinese model running on it Wouldn't worry about that, I'm pretty sure the government is going to ban running Chinese tech in this space sooner or later. And we won't even be able to download it. Not saying any of the bans will make any kind of sense, but I'm pretty sure they're gonna say this is a "strategic" space. And e…

When DeepSeek first hit the news, an American senator proposed adding it to ITAR so they could send people to prison for using it. Didn't pass, thankfully.

If it does in the future, do we just hope it won’t be retroactive? Is this water boiling yet?

Re: Qwen3-Omni: Native Omni AI model for text, image and video

#46

Interesting, the pacing seemed very slow when conversing in english, but when I spoke to it in spanish, it sounded much faster. It's really impressive that these models are going to be able to do real time translation and much more. The Chinese are going to end up owning the AI market if the American labs don't start competing on open weights. Americans may end up in a situation where they have some $1000-2000 device…

When has the average American ever been willing to spend a $1,000-2,000 premium for privacy-respecting tech? They already save $20-200 to buy IoT cameras which provide all audio and video from inside their home directly to the government without a warrant (Ring vs Reolink/etc).

Re: Qwen3-Omni: Native Omni AI model for text, image and video

#47
post #18

You can try it out on https://chat.qwen.ai/ - sign in with Google or GitHub (signed out users can't use the voice mode) and then click on the voice icon. It has an entertaining selection of different voices, including: *Dylan* - A teenager who grew up in Beijing's hutongs *Peter* - Tianjin crosstalk, professionally supporting others *Cherry* - A sunny, positive, friendly, and natural young lady *Ethan* - A sunny, war…

I only see Omni Flash, is that the one?

same, did you figure it out?

Re: Qwen3-Omni: Native Omni AI model for text, image and video

#48
post #30

Earlier quoted context omitted.

When DeepSeek first hit the news, an American senator proposed adding it to ITAR so they could send people to prison for using it. Didn't pass, thankfully.

If it does in the future, do we just hope it won’t be retroactive? Is this water boiling yet?

ex post facto law is explicitly banned in the US Constitution

Re: Qwen3-Omni: Native Omni AI model for text, image and video

#50
post #28

Interesting, the pacing seemed very slow when conversing in english, but when I spoke to it in spanish, it sounded much faster. It's really impressive that these models are going to be able to do real time translation and much more. The Chinese are going to end up owning the AI market if the American labs don't start competing on open weights. Americans may end up in a situation where they have some $1000-2000 device…

This is exactly what I do. I have two 3090s at home, with Qwen3 on it. This is tied into my Home Assistant install, and I use esp32 devices as voice satellites. It works shockingly well.

Can you tell me about these voice satellites?
Post reply on HN