Big day for open source Chinese model releases - DeepSeek-v3-0324 came out today too, an updated version of DeepSeek v3 now under an MIT license (previously it was a custom DeepSeek license). https://simonwillison.net/2025/Mar/24/deepseek/
And it still can't answer this: Q: "9.11 and 9.9, which one is larger?" A: "To determine which number is larger between 9.11 and 9.9, let's compare them step by step. Both numbers have the same whole number part: 9. 9.11 has a 1 in the tenths place. 9.9 has a 9 in the tenths place. Since 9 (from 9.9) is greater than 1 (from 9.11), we can conclude that 9.9 is larger than 9.11." "Final Answer: 9.9" I don't think anythi…
Qwen2.5-VL-32B: Smarter and Lighter
191–200 of 303 posts
Re: Qwen2.5-VL-32B: Smarter and Lighter
#192Re: Qwen2.5-VL-32B: Smarter and Lighter
#193Silly question: how can OpenAI, Claude and all, have a valuation so large considering all the open source models? Not saying they will disappear or be tiny (closed models), but why so so so valuable?
The average user won't self-host a model.
Re: Qwen2.5-VL-32B: Smarter and Lighter
#194Earlier quoted context omitted.
And it still can't answer this: Q: "9.11 and 9.9, which one is larger?" A: "To determine which number is larger between 9.11 and 9.9, let's compare them step by step. Both numbers have the same whole number part: 9. 9.11 has a 1 in the tenths place. 9.9 has a 9 in the tenths place. Since 9 (from 9.9) is greater than 1 (from 9.11), we can conclude that 9.9 is larger than 9.11." "Final Answer: 9.9" I don't think anythi…
But that’s correct. 9.9 = 9.90 > 9.11. Seems that it answered the question absolutely correctly.
Re: Qwen2.5-VL-32B: Smarter and Lighter
#195Re: Qwen2.5-VL-32B: Smarter and Lighter
#196Deepseek has proved that fp8 is more cost-effectiveness than fp16, isn't it valid for dozens-B model?
Re: Qwen2.5-VL-32B: Smarter and Lighter
#19732B is one of my favourite model sizes at this point - large enough to be extremely capable (generally equivalent to GPT-4 March 2023 level performance, which is when LLMs first got really useful) but small enough you can run them on a single GPU or a reasonably well specced Mac laptop (32GB or more).
Re: Qwen2.5-VL-32B: Smarter and Lighter
#198Earlier quoted context omitted.
And it still can't answer this: Q: "9.11 and 9.9, which one is larger?" A: "To determine which number is larger between 9.11 and 9.9, let's compare them step by step. Both numbers have the same whole number part: 9. 9.11 has a 1 in the tenths place. 9.9 has a 9 in the tenths place. Since 9 (from 9.9) is greater than 1 (from 9.11), we can conclude that 9.9 is larger than 9.11." "Final Answer: 9.9" I don't think anythi…
I suggest we’ve already now passed what shall be dubbed the jschoe test ;)
It's interesting to think that maybe one of the most realistic consequences of reaching artificial superintelligence will be when its answers start wildly diverging from human expectations and we think it's being "increasingly wrong".
Re: Qwen2.5-VL-32B: Smarter and Lighter
#199So today is Qwen. Tomorrow a new SOTA model from Google apparently, R2 next week. We haven't hit the wall yet.
Google's announcements are mostly vaporware anyway. Btw, where is Gemini Ultra 1 ? how about Gemini Ultra 2?
Re: Qwen2.5-VL-32B: Smarter and Lighter
#200Silly question: how can OpenAI, Claude and all, have a valuation so large considering all the open source models? Not saying they will disappear or be tiny (closed models), but why so so so valuable?
That being said, they have a user base and integrations. As long as they stay close or a bit ahead of the Chinese models they'll be fine. If the Chinese models significantly jumps ahead of them, well, then they are pretty much dead. Add open source to the mix and they become history.