Silly question: how can OpenAI, Claude and all, have a valuation so large considering all the open source models? Not saying they will disappear or be tiny (closed models), but why so so so valuable?
People cannot normally invest in their competitors. It's not unlikely that chinese products may be banned / tarriff'd
Qwen2.5-VL-32B: Smarter and Lighter
91–100 of 303 posts
Re: Qwen2.5-VL-32B: Smarter and Lighter
#92So today is Qwen. Tomorrow a new SOTA model from Google apparently, R2 next week. We haven't hit the wall yet.
Re: Qwen2.5-VL-32B: Smarter and Lighter
#93Earlier quoted context omitted.
Would a 40GB A6000 fully accommodate a 32B model? I assume an fp16 quantization is still necessary?
At FP16 you‘d need 64GB just for the weights, and it‘d be 2x as slow as a Q8 version, likely with little improvement. You‘ll also need space for attention and context etc, so 80-100GB (or even more) VRAM would be better. Many people „just“ use 4x consumer GPUs like the 3090 (24GB each) which scales well. They’d probably buy a mining rig, EPYC CPU, Mainboard with sufficient PCIe lanes, PCIe risers, 1600W PSU (might ne…
Re: Qwen2.5-VL-32B: Smarter and Lighter
#94Earlier quoted context omitted.
Since we are on HN here, I can highly recommend open-webui with some OpenAI-compatible provider. I'm running with Deep Infra for more than a year now and am very happy. New models are usually available within one or two days after release. Also have some friends who use the service almost daily.
I'm using open-webui at home with a couple of different models. gemma2-9b fits in VRAM on a NV 3060 card + performs nicely.
Re: Qwen2.5-VL-32B: Smarter and Lighter
#95Earlier quoted context omitted.
I've been waiting since November for 1, just 1*, model other than Claude than can reliably do agentic tool call loops. As long as the Chinese open models are chasing reasoning and benchmark maxxing vs. mid-2024 US private models, I'm very comfortable with somewhat ignoring these models. (this isn't idle prognostication hinging on my personal hobby horse. I got skin in the game, I'm virtually certain I have the only A…
is there some reason you cant train a 1b model to just do agentic stuff?
Re: Qwen2.5-VL-32B: Smarter and Lighter
#96Earlier quoted context omitted.
Pretty soon I won't be using any American models. It'll be a 100% Chinese open source stack. The foundation model companies are screwed. Only shovel makers (Nvidia, infra companies) and product companies are going to win.
I still don't get where the money for new open source models is going to come from once setting investor dollars on fire is no longer a viable business model. Does anyone seriously expect companies to keep buying and running thousands of ungodly expensive GPUs, plus whatever they spend on human workers to do labelling/tuning, and then giving away the spoils for free, forever?
This has clearly been part of a viable business model for a long time. Why should LLM models be any different?
Re: Qwen2.5-VL-32B: Smarter and Lighter
#97Big day for open source Chinese model releases - DeepSeek-v3-0324 came out today too, an updated version of DeepSeek v3 now under an MIT license (previously it was a custom DeepSeek license). https://simonwillison.net/2025/Mar/24/deepseek/
Why do you keep promoting your blog on every LLM post?
Re: Qwen2.5-VL-32B: Smarter and Lighter
#98Earlier quoted context omitted.
it seems that this free version "may use your prompts and completions to train new models" https://openrouter.ai/deepseek/deepseek-chat-v3-0324:free do you think this needs attention?
good grief! people are okay with it when OpenAI and Google do it, but as soon as open source providers do it, people get defensive about it...
Big companies have so much data they won't be having a human look at mine specifically. Some small place probably has the engineer looking at my logs as user #4.
Also, big companies have security teams whose job is securing the data, and it won't be going over some unencrypted link to cloudflare because OP was too lazy to set up Https certs.
Re: Qwen2.5-VL-32B: Smarter and Lighter
#99Big day for open source Chinese model releases - DeepSeek-v3-0324 came out today too, an updated version of DeepSeek v3 now under an MIT license (previously it was a custom DeepSeek license). https://simonwillison.net/2025/Mar/24/deepseek/
Why do you keep promoting your blog on every LLM post?
Re: Qwen2.5-VL-32B: Smarter and Lighter
#100tbh I’d settle for just lighter