We haven't hit the wall yet.
Qwen2.5-VL-32B: Smarter and Lighter
11–20 of 303 posts
Re: Qwen2.5-VL-32B: Smarter and Lighter
#12Big day for open source Chinese model releases - DeepSeek-v3-0324 came out today too, an updated version of DeepSeek v3 now under an MIT license (previously it was a custom DeepSeek license). https://simonwillison.net/2025/Mar/24/deepseek/
it seems that this free version "may use your prompts and completions to train new models" https://openrouter.ai/deepseek/deepseek-chat-v3-0324:free do you think this needs attention?
Re: Qwen2.5-VL-32B: Smarter and Lighter
#1332B is one of my favourite model sizes at this point - large enough to be extremely capable (generally equivalent to GPT-4 March 2023 level performance, which is when LLMs first got really useful) but small enough you can run them on a single GPU or a reasonably well specced Mac laptop (32GB or more).
Re: Qwen2.5-VL-32B: Smarter and Lighter
#14I've seen some people claim it should make the models better at text, but I find that a little difficult to believe without data.
Re: Qwen2.5-VL-32B: Smarter and Lighter
#15(It's not mentioned anywhere in the blog post.)
Re: Qwen2.5-VL-32B: Smarter and Lighter
#16Big day for open source Chinese model releases - DeepSeek-v3-0324 came out today too, an updated version of DeepSeek v3 now under an MIT license (previously it was a custom DeepSeek license). https://simonwillison.net/2025/Mar/24/deepseek/
it seems that this free version "may use your prompts and completions to train new models" https://openrouter.ai/deepseek/deepseek-chat-v3-0324:free do you think this needs attention?
Re: Qwen2.5-VL-32B: Smarter and Lighter
#17Big day for open source Chinese model releases - DeepSeek-v3-0324 came out today too, an updated version of DeepSeek v3 now under an MIT license (previously it was a custom DeepSeek license). https://simonwillison.net/2025/Mar/24/deepseek/
The foundation model companies are screwed. Only shovel makers (Nvidia, infra companies) and product companies are going to win.
Re: Qwen2.5-VL-32B: Smarter and Lighter
#18Big day for open source Chinese model releases - DeepSeek-v3-0324 came out today too, an updated version of DeepSeek v3 now under an MIT license (previously it was a custom DeepSeek license). https://simonwillison.net/2025/Mar/24/deepseek/
it seems that this free version "may use your prompts and completions to train new models" https://openrouter.ai/deepseek/deepseek-chat-v3-0324:free do you think this needs attention?
Re: Qwen2.5-VL-32B: Smarter and Lighter
#19Re: Qwen2.5-VL-32B: Smarter and Lighter
#2032B is one of my favourite model sizes at this point - large enough to be extremely capable (generally equivalent to GPT-4 March 2023 level performance, which is when LLMs first got really useful) but small enough you can run them on a single GPU or a reasonably well specced Mac laptop (32GB or more).
I don't think these models are GPT-4 level. Yes they seem to be on benchmarks, but it has been known that models increasingly use A/B testing in dataset curation and synthesis(using GPT 4 level models) to optimize not just the benchmarks but things which could be benchmarked like academics.
To pick just the most popular one, https://lmarena.ai/?leaderboard= has GPT-4-0314 ranked 83rd now.