Live data from Hacker News

Qwen2.5-VL-32B: Smarter and Lighter

qwenlm.github.io

181–190 of 303 posts

Re: Qwen2.5-VL-32B: Smarter and Lighter

#181
post #135

Earlier quoted context omitted.

And it’s quite easy to set up a Cloudflare tunnel to make your open-webui instance accessible online too just you

... or a TailScale network. I've been leaving open-webui running on my laptop on my desk and then going out into the word and accessing it from my phone via TailScale, works great.

I would use tail scale. But I specifically want to use open web-ui from a place I can’t install a Tailscale client

Re: Qwen2.5-VL-32B: Smarter and Lighter

#182
post #152
post #4

Big day for open source Chinese model releases - DeepSeek-v3-0324 came out today too, an updated version of DeepSeek v3 now under an MIT license (previously it was a custom DeepSeek license). https://simonwillison.net/2025/Mar/24/deepseek/

And it still can't answer this: Q: "9.11 and 9.9, which one is larger?" A: "To determine which number is larger between 9.11 and 9.9, let's compare them step by step. Both numbers have the same whole number part: 9. 9.11 has a 1 in the tenths place. 9.9 has a 9 in the tenths place. Since 9 (from 9.9) is greater than 1 (from 9.11), we can conclude that 9.9 is larger than 9.11." "Final Answer: 9.9" I don't think anythi…

But that’s correct. 9.9 = 9.90 > 9.11. Seems that it answered the question absolutely correctly.

Re: Qwen2.5-VL-32B: Smarter and Lighter

#185

Earlier quoted context omitted.

Nah, it's great for things that Western models are censored on. The True Hacker will keep an Eastern and Western model available, depending on what they need information on.

I tried to ask it about Java exploits that would allow me to gain RCE, but it refused just as most western models do. That was the only thing I could think to ask really. Do you have a better example maybe?

Adult content and things like making biological/chemical/nuclear weapons are the other main topics that usually get censored. I don’t think the Chinese models tend to be less censored than western models in these dimensions. You can sometimes find “uncensored“ models on HuggingFace where people basically finetune sensitive topics back in. There is a finetuned version of R1 called 1776 that will correctly answer Chinese-censored questions, for example.

Re: Qwen2.5-VL-32B: Smarter and Lighter

#187
post #152

Earlier quoted context omitted.

And it still can't answer this: Q: "9.11 and 9.9, which one is larger?" A: "To determine which number is larger between 9.11 and 9.9, let's compare them step by step. Both numbers have the same whole number part: 9. 9.11 has a 1 in the tenths place. 9.9 has a 9 in the tenths place. Since 9 (from 9.9) is greater than 1 (from 9.11), we can conclude that 9.9 is larger than 9.11." "Final Answer: 9.9" I don't think anythi…

This is hilarious, especially if it's unintentional.

Poe's law in effect.

Re: Qwen2.5-VL-32B: Smarter and Lighter

#189

Earlier quoted context omitted.

I asked o1-pro what 99490126816810951552*23977364624054235203 is, yesterday. It took 16 minutes to get an answer which is off by eight orders of magnitude. https://chatgpt.com/share/67e1eba1-c658-800e-9161-a0b8b7b683...

Sorry, I'm in a rush, could only afford a couple minutes looking at it, but I'm missing something: Google: 2.385511e+39 Your chat: "Numerically, that’s about 2.3855 × 10^39" Also curious how you think about LLM-as-calculator in relation to tool calls.

If you look at the precise answer, it's got 8 too many digits, despite it getting the right number of digits in the estimate you looked at.

> Also curious how you think about LLM-as-calculator in relation to tool calls.

I just tried this because I heard all existing models are bad at this kind of problem, and wanted to try it with the most powerful one I have access to. I think it shows that you really want an AI to be able to use computational tools in appropriate circumstances.

Re: Qwen2.5-VL-32B: Smarter and Lighter

#190
post #152
post #4

Big day for open source Chinese model releases - DeepSeek-v3-0324 came out today too, an updated version of DeepSeek v3 now under an MIT license (previously it was a custom DeepSeek license). https://simonwillison.net/2025/Mar/24/deepseek/

And it still can't answer this: Q: "9.11 and 9.9, which one is larger?" A: "To determine which number is larger between 9.11 and 9.9, let's compare them step by step. Both numbers have the same whole number part: 9. 9.11 has a 1 in the tenths place. 9.9 has a 9 in the tenths place. Since 9 (from 9.9) is greater than 1 (from 9.11), we can conclude that 9.9 is larger than 9.11." "Final Answer: 9.9" I don't think anythi…

> I don't think anything else needs to be said here.

Will this humbling moment change your opinion?

Post reply on HN