Live data from Hacker News

Moondream 3 Preview: Frontier-level reasoning at a blazing speed

moondream.ai

41–46 of 46 posts

Re: Moondream 3 Preview: Frontier-level reasoning at a blazing speed

#41
post #11
post #7

Moondream 2 has been very useful for me: I've been using it to automatically label object detection datasets for novel classes and distill an orders of magnitude smaller but similarly accurate CNN. One oddity is that I haven't seen the claimed improvements beyond the 2025-01-09 tag - subsequent releases improve recall but degrade precision pretty significantly. It'd be amazing if object detection VLMs like this repor…

Thanks! If you could shoot me a note at vik@m87.ai with any examples of the precision/recall issues you saw I'd appreciate it a ton.

are you planning to release a GGUF?

Re: Moondream 3 Preview: Frontier-level reasoning at a blazing speed

#42

Spent 5 minutes trying to get basic pricing info for Moondream cloud. Seems it simply does not exist (or at least not until you've actually signed up?). There's 5,000 free requests but I need to sense-check the pricing as viable as step 0 of evaluating - long before hooking it up to an app.

We are looking to launch our cloud very soon. We are still optimizing our inference to get you the best pricing we can offer. Follow @moondreamai on X if you want your ear to the ground for our launch!

Will you add this to OpenRouter too?

Re: Moondream 3 Preview: Frontier-level reasoning at a blazing speed

#43

Really impressive performance from the Moondream model, but looking at the results from the big 3 labs, it's absolutely wild how poorly Claude and OpenAI perform. Gemini isn't as good as Moondream, but it's clearly the only one that's even half way decent at these vision tasks. I didn't realize how big a performance gap there was.

I'm not sure why they haven't been acquired yet by any of the big ones, since clearly Moondream is pretty good! Definitely seems like something Anthropic/OpenAI/whoever would want to fold into their platforms and such. Everyone involved in creating it should probably be swimming in money and visual use cases for LLMs should become far less useless with the reach of the big orgs.

Re: Moondream 3 Preview: Frontier-level reasoning at a blazing speed

#45

Can anyone suggest what's the cheapest hardware to run this model locally with a reasonable performance?

Since there's no quantized version available at the moment, you'll need ~20 GB of memory for the weights plus some extra for the KV cache. CPU with 32 GB RAM will be the cheapest and still reasonably fast given the relatively small number of activated parameters.

Thank you!

I don't even know what a "quantized version" is, but I was expecting answers about NVIDIA graphics cards and their memory. My computer has 24GB of memory, but I'll go for 64GB to run this locally on a new computer.

Post reply on HN