Live data from Hacker News

Moondream 3 Preview: Frontier-level reasoning at a blazing speed

moondream.ai

31–40 of 46 posts

Re: Moondream 3 Preview: Frontier-level reasoning at a blazing speed

#31
post #11

Earlier quoted context omitted.

Thanks! If you could shoot me a note at vik@m87.ai with any examples of the precision/recall issues you saw I'd appreciate it a ton.

Will do!

Wonderful to see "at the coalface" collaboration happen on this stuff at HN. More than just a newsfeed!

Re: Moondream 3 Preview: Frontier-level reasoning at a blazing speed

#32
This looks amazing. I'm a big fan of Gemini for bounding box operations, the idea that a 9B model could outperform it is incredibly exciting!

I noticed that Moondream 2 was Apache 2 licensed but the 3 preview is currently BSL ("You can’t (without a deal): offer the model’s functionality to anyone outside your organization—e.g., an external API, or managed hosting for customers") - is that a permanent change to your licensing policies?

Re: Moondream 3 Preview: Frontier-level reasoning at a blazing speed

#34
post #32

This looks amazing. I'm a big fan of Gemini for bounding box operations, the idea that a 9B model could outperform it is incredibly exciting! I noticed that Moondream 2 was Apache 2 licensed but the 3 preview is currently BSL ("You can’t (without a deal): offer the model’s functionality to anyone outside your organization—e.g., an external API, or managed hosting for customers") - is that a permanent change to your l…

I just noticed in https://huggingface.co/moondream/moondream3-preview/blob/mai... that the license is set to change to Apache 2 after two years.

Re: Moondream 3 Preview: Frontier-level reasoning at a blazing speed

#35
Really impressive performance from the Moondream model, but looking at the results from the big 3 labs, it's absolutely wild how poorly Claude and OpenAI perform. Gemini isn't as good as Moondream, but it's clearly the only one that's even half way decent at these vision tasks. I didn't realize how big a performance gap there was.

Re: Moondream 3 Preview: Frontier-level reasoning at a blazing speed

#36

Really impressive performance from the Moondream model, but looking at the results from the big 3 labs, it's absolutely wild how poorly Claude and OpenAI perform. Gemini isn't as good as Moondream, but it's clearly the only one that's even half way decent at these vision tasks. I didn't realize how big a performance gap there was.

Gemini is really fantastic at anything that's OCR-adjacent, and it promptly falls over on most other image-related tasks.

Re: Moondream 3 Preview: Frontier-level reasoning at a blazing speed

#37

Spent 5 minutes trying to get basic pricing info for Moondream cloud. Seems it simply does not exist (or at least not until you've actually signed up?). There's 5,000 free requests but I need to sense-check the pricing as viable as step 0 of evaluating - long before hooking it up to an app.

We are looking to launch our cloud very soon. We are still optimizing our inference to get you the best pricing we can offer. Follow @moondreamai on X if you want your ear to the ground for our launch!

Re: Moondream 3 Preview: Frontier-level reasoning at a blazing speed

#39

Can anyone suggest what's the cheapest hardware to run this model locally with a reasonable performance?

Since there's no quantized version available at the moment, you'll need ~20 GB of memory for the weights plus some extra for the KV cache. CPU with 32 GB RAM will be the cheapest and still reasonably fast given the relatively small number of activated parameters.

Re: Moondream 3 Preview: Frontier-level reasoning at a blazing speed

#40

Really impressive performance from the Moondream model, but looking at the results from the big 3 labs, it's absolutely wild how poorly Claude and OpenAI perform. Gemini isn't as good as Moondream, but it's clearly the only one that's even half way decent at these vision tasks. I didn't realize how big a performance gap there was.

Funnily enough, Gemini is also the only one able to read a D20. ChatGPT consistently gets it wrong, and Claude mostly argues it can't read the face of the die that's facing up because it's obstructed (it's not lol).
Post reply on HN