The model weights are 70GB (Hugging Face recently added a file size indicator - see https://huggingface.co/Qwen/Qwen3-Omni-30B-A3B-Instruct/tree... ) so this one is reasonably accessible to run locally. I wonder if we'll see a macOS port soon - currently it very much needs an NVIDIA GPU as far as I can tell.
That's at BF16, so it should fit fairly well on 24GB GPUs after quantization to Q4, I'd think. (Much like the other 30B-A3B models in the family.) I'm pretty happy about that - I was worried it'd be another 200B+.
Qwen3-Omni: Native Omni AI model for text, image and video
61–70 of 152 posts
Re: Qwen3-Omni: Native Omni AI model for text, image and video
#62Earlier quoted context omitted.
If it does in the future, do we just hope it won’t be retroactive? Is this water boiling yet?
ex post facto law is explicitly banned in the US Constitution
Re: Qwen3-Omni: Native Omni AI model for text, image and video
#63Interesting, the pacing seemed very slow when conversing in english, but when I spoke to it in spanish, it sounded much faster. It's really impressive that these models are going to be able to do real time translation and much more. The Chinese are going to end up owning the AI market if the American labs don't start competing on open weights. Americans may end up in a situation where they have some $1000-2000 device…
When has the average American ever been willing to spend a $1,000-2,000 premium for privacy-respecting tech? They already save $20-200 to buy IoT cameras which provide all audio and video from inside their home directly to the government without a warrant (Ring vs Reolink/etc).
What percentage of the people that you know are able to install python and dependencies plus the correct open weights models?
I'd wager most of your parents can't do it.
Most "normies" wouldn't even know what a local model even is, let alone how to install a GPU.
Re: Qwen3-Omni: Native Omni AI model for text, image and video
#64The model weights are 70GB (Hugging Face recently added a file size indicator - see https://huggingface.co/Qwen/Qwen3-Omni-30B-A3B-Instruct/tree... ) so this one is reasonably accessible to run locally. I wonder if we'll see a macOS port soon - currently it very much needs an NVIDIA GPU as far as I can tell.
Would it run on 5090? Or is it possible to link multiple GPUs or has NVIDIA locked it down?
Re: Qwen3-Omni: Native Omni AI model for text, image and video
#65Earlier quoted context omitted.
That's at BF16, so it should fit fairly well on 24GB GPUs after quantization to Q4, I'd think. (Much like the other 30B-A3B models in the family.) I'm pretty happy about that - I was worried it'd be another 200B+.
are there any that would run on 16GB Apple M1?
Re: Qwen3-Omni: Native Omni AI model for text, image and video
#66Earlier quoted context omitted.
This is exactly what I do. I have two 3090s at home, with Qwen3 on it. This is tied into my Home Assistant install, and I use esp32 devices as voice satellites. It works shockingly well.
I run Home Assistant on an RPi4 and have an ESP32-based Core2 with mic ( https://shop.m5stack.com/products/m5stack-core2-esp32-iot-de... ), along with a 16GB 4070 Ti Super in an always-on Windows system I only use for occasional gaming and serving media. I'd love to set up something like you have. Can you recommended a starting place, or ideally, a step-by-step tutorial? I've never set up any AI system. Would you say…
Claude code is your friend.
I run proxmox on an old Dell R710 in my closet that hosts my homeassistant (amongst others) VM and then I've setup my "gaming" PC (which hasn't done any gaming in quite some time) to dual boot (Windows or Deb/Proxmox) and just keep it booted into Deb as another proxmox node. That PC also has a 4070 Super that I have setup to passthru to a VM and on that VM I've got various services utilizing the GPU. This includes some that are utilized by my hetzner bare metal servers for things like image/text embeddings as well as local LLM use (though, rather minimal due to VRAM constraints) and some image/video object detection stuff with my security cameras (slowly working on a remote water gun turret to keep the racoons from trying to eat the kittens that stray cats keep having in my driveway/workshop).
Install claude code (or, opencode, it's also good) - use Opus (get the max plan) and give it a directory that it can use as it's working directory (don't open it in ~/Documents and just start doing things) and prompt it with something as simple as this:
"I have an existing home assistant setup at home and I'd like to determine what sort of self-hosted AI I could setup and integrate with that home assistant install - can you help me get started? Please also maintain some notes in .md files in this working directory with those note files named and organized as you see appropriate so that we can share relevant context and information with future sessions. (example: Hardware information, local urls, network layout, etc) If you're unsure of something, ask me questions. Do not perform any destructive actions without first confirming with me."
Plan mode. _ALWAYS_ use plan mode to get the task setup, if there's something about the plan you don't like, say no and give it notes - it will return with a new plan. Eventually agree to the plan when it's right - then work through that plan not in plan mode, but if it gets off the plan, get back in plan mode to get the/a plan set and then again let it go and just steer it in regular mode.
Re: Qwen3-Omni: Native Omni AI model for text, image and video
#67Re: Qwen3-Omni: Native Omni AI model for text, image and video
#68Pasted here for your own judgement:
*Title: The Last Lightbulb*
The power had been out for three days. Rain drummed against the windows of the old cabin, and the only light came from a flickering candle on the kitchen table.
Maggie, wrapped in a wool blanket, squinted at the last working flashlight. “We’ve got one bulb left, Jack. One.”
Jack, hunched over a board game he’d dug out of the closet, didn’t look up. “Then don’t turn it on unless you’re reading Shakespeare or delivering a baby.”
She rolled her eyes. “I need it to find the can opener. I’m not eating cold beans with my fingers again.”
Jack finally glanced up, grinning. “You did that yesterday and called it ‘rustic dining.’”
“Desperate times,” she muttered, clicking the flashlight on. The beam cut through the gloom—and immediately began to dim.
“No—!” Jack lunged, but too late. The light sputtered… then died.
Silence. Then Maggie sighed. “Well. There goes civilization.”
Jack leaned back, chuckling. “Guess we’re officially cavemen now.”
“Cavewoman,” she corrected, fumbling in the dark. “And I’m going to bed. Wake me when the grid remembers we exist.”
As she shuffled off, Jack called after her, “Hey—if you find the can opener in the dark, you’re officially magic.”
A pause. Then, from down the hall: “I found socks that match. That’s basically witchcraft.”
Jack smiled into the dark. “Goodnight, witch.”
“Goodnight, caveman.”
Outside, the rain kept falling. Inside, the dark didn’t feel so heavy anymore.
— The End —
Re: Qwen3-Omni: Native Omni AI model for text, image and video
#69I recently needed to scan hundreds of low quality invoices and run them through OCR for invoice numbers and dates. I really took for granted how seamless this is in some applications, and was shocked how much work went into producing decent results. I was obviously really naive. Either way, it gets me excited any time I see progress with OCR. I should give this a try against my (small) dataset.
What is your point about Qwen? Or is it just a general statement regarding LLM?
All I'm saying is I'm excited to try Qwen to see if it out performs my gnarly algorithm.
Re: Qwen3-Omni: Native Omni AI model for text, image and video
#70You can try it out on https://chat.qwen.ai/ - sign in with Google or GitHub (signed out users can't use the voice mode) and then click on the voice icon. It has an entertaining selection of different voices, including: *Dylan* - A teenager who grew up in Beijing's hutongs *Peter* - Tianjin crosstalk, professionally supporting others *Cherry* - A sunny, positive, friendly, and natural young lady *Ethan* - A sunny, war…
In Russian, Ryan sounds like a westerner who started reading Russian words a month ago.
Dylan sounds somewhat authentic, while everyone else is a different degree of heavy-asian-accented Russian.