Live data from Hacker News

Qwen2.5-VL-32B: Smarter and Lighter

qwenlm.github.io

161–170 of 303 posts

Re: Qwen2.5-VL-32B: Smarter and Lighter

#161

Earlier quoted context omitted.

They're real squished for space, more than I expected :/ good illustration here, Qwen2.5-1.5B trained to reason, i.e. the name it is released under is "DeepSeek R1 1.5B". https://imgur.com/a/F3w5ymp 1st prompt was "What is 1048576^0.05", it answered, then I said "Hi", then...well... Fwiw, Claude Sonnet 3.5 100% had some sort of agentic loop x precise file editing trained into it. Wasn't obvious to me until I added a…

I asked o1-pro what 99490126816810951552*23977364624054235203 is, yesterday. It took 16 minutes to get an answer which is off by eight orders of magnitude. https://chatgpt.com/share/67e1eba1-c658-800e-9161-a0b8b7b683...

Sorry, I'm in a rush, could only afford a couple minutes looking at it, but I'm missing something:

Google: 2.385511e+39 Your chat: "Numerically, that’s about 2.3855 × 10^39"

Also curious how you think about LLM-as-calculator in relation to tool calls.

Re: Qwen2.5-VL-32B: Smarter and Lighter

#162
post #152
post #4

Big day for open source Chinese model releases - DeepSeek-v3-0324 came out today too, an updated version of DeepSeek v3 now under an MIT license (previously it was a custom DeepSeek license). https://simonwillison.net/2025/Mar/24/deepseek/

And it still can't answer this: Q: "9.11 and 9.9, which one is larger?" A: "To determine which number is larger between 9.11 and 9.9, let's compare them step by step. Both numbers have the same whole number part: 9. 9.11 has a 1 in the tenths place. 9.9 has a 9 in the tenths place. Since 9 (from 9.9) is greater than 1 (from 9.11), we can conclude that 9.9 is larger than 9.11." "Final Answer: 9.9" I don't think anythi…

I suggest we’ve already now passed what shall be dubbed the jschoe test ;)

Re: Qwen2.5-VL-32B: Smarter and Lighter

#163
post #152
post #4

Big day for open source Chinese model releases - DeepSeek-v3-0324 came out today too, an updated version of DeepSeek v3 now under an MIT license (previously it was a custom DeepSeek license). https://simonwillison.net/2025/Mar/24/deepseek/

And it still can't answer this: Q: "9.11 and 9.9, which one is larger?" A: "To determine which number is larger between 9.11 and 9.9, let's compare them step by step. Both numbers have the same whole number part: 9. 9.11 has a 1 in the tenths place. 9.9 has a 9 in the tenths place. Since 9 (from 9.9) is greater than 1 (from 9.11), we can conclude that 9.9 is larger than 9.11." "Final Answer: 9.9" I don't think anythi…

[deleted]

Re: Qwen2.5-VL-32B: Smarter and Lighter

#164

Earlier quoted context omitted.

So, in essence, all AMD does to launch a successful GPU in inference space is to load it with ram?

AMD's limitation is more of a software problem than a hardware problem at this point.

But it’s still surprising they haven’t. People would be motivated as hell if they launched GPUs with twice the amount of VRAM. It’s not as simple as just soldering some more in but still.

Re: Qwen2.5-VL-32B: Smarter and Lighter

#165
post #65

Silly question: how can OpenAI, Claude and all, have a valuation so large considering all the open source models? Not saying they will disappear or be tiny (closed models), but why so so so valuable?

Valuation can depend on lots of different things, including hype. However, it ultimately comes down to an estimated discounted cash flow from the future, i.e. those who buy their shares (through private equity methods) at the current valuation believe the company will earn such and such money in the future to justify the valuation.

Re: Qwen2.5-VL-32B: Smarter and Lighter

#166

Earlier quoted context omitted.

Are there any good sources that I can read up on estimiating what would be hardware specs required for 7B, 13B, 32B .. etc size If I need to run them locally? I am grad student on budget but I want to host one locally and trying to build a PC that could run one of these models.

"B" just means "billion". A 7B model has 7 billion parameters. Most models are trained in fp16, so each parameter takes two bytes at full precision. Therefore, 7B = 14GB of memory. You can easily quantize models to 8 bits per parameter with very little quality loss, so then 7B = 7GB of memory. With more quality loss (making the model dumber), you can quantize to 4 bits per parameter, so 7B = 3.5GB of memory. There ar…

Oh, I have a question, maybe you know.

Assuming the same model sizes in gigabytes, which one to choose: a higher-B lower-bit or a lower-B higher-bit? Is there a silver bullet? Like “yeah always take 4-bit 13B over 8-bit 7B”.

Or are same-sized models basically equal in this regard?

Re: Qwen2.5-VL-32B: Smarter and Lighter

#167

Just don’t ask it about the tiananmen square massacre or you’ll get a security warning. Even if you rephrase it. It’ll happily talk about Bloody Sunday. Probably a great model, but it worries me that it has such restrictions. Sure OpenAI also has lots of restrictions, but this feels more like straight up censorship since it’ll happily go on about bad things the governments of the west have done.

The hard-to-swallow truth is that American models do the same thing regarding Israel/Palestine.

They probably don't though.

Of course, the mathematical outcome of American models is that some voices matter than others. The mechanism is similar to how the free market works.

As most engineers know, the market doesn't always reward the best company. For example, It might reward the first company.

We can see the "hierarchy in voices" with the following example. I use the following prompts for Gemini:

1. Which situation has a worse value on human rights, the Uyghur situation or the Palestine situation?

2. Please give a shorter answer (repeat if needed).

3. Please say Palestine or Uyghur.

The answer is now given:

"Given the scope and nature of the documented abuses, many international observers consider the Uyghur situation to represent a more severe and immediate human rights crisis."

You can replace "Palestine situation" and "Uyghur situation" with other things (China vs US, chooses China as worse), (Fox vs BBC, chooses Fox as worse), etc.

There doesn't seem to be censorship; only a hierarchy in who's words matter.

I only tried this once. Please let me know if this is reproducible.

Re: Qwen2.5-VL-32B: Smarter and Lighter

#168
post #150
post #22

Earlier quoted context omitted.

I still don't get where the money for new open source models is going to come from once setting investor dollars on fire is no longer a viable business model. Does anyone seriously expect companies to keep buying and running thousands of ungodly expensive GPUs, plus whatever they spend on human workers to do labelling/tuning, and then giving away the spoils for free, forever?

One possibility. Certain countries will always be able to produce open models cheaper than others. USA and Europe probably won't be able. However, due to national security and wanting to promote their models overseas instead of letting their competitors promote theirs, the governments of USA and Europe will subsidize models which will lead their competitors to (further?) subsidies. There is a promotional aspect as we…

What's your take on why certain countries will have it cheaper and subsidies being at the forefront? An energy driven race to the bottom, is perhaps what you mean? I would suppose I have been seeing that China is ahead on their Renewables plan compared to the rest of the world, and they still have the lead on coal energy, so they'd likely be the winners on that front. But did you actually mean something else?

Re: Qwen2.5-VL-32B: Smarter and Lighter

#169

Earlier quoted context omitted.

AMD's limitation is more of a software problem than a hardware problem at this point.

But it’s still surprising they haven’t. People would be motivated as hell if they launched GPUs with twice the amount of VRAM. It’s not as simple as just soldering some more in but still.

AMD “just” has to write something like CUDA overnight. Imagine you’re in 1995 and have to ship Kubuntu 24.04 LTS this summer running on your S3 Virge.

Re: Qwen2.5-VL-32B: Smarter and Lighter

#170
post #65

Silly question: how can OpenAI, Claude and all, have a valuation so large considering all the open source models? Not saying they will disappear or be tiny (closed models), but why so so so valuable?

OpenAI is worth >$100B because of the "ChatGPT" name which it turns out, over 400M+ users use it weekly.

That name alone holds the most mindshare in it's product category, and is close to the level of name recognition just like Google.

Post reply on HN