Live data from Hacker News

Qwen2.5-VL-32B: Smarter and Lighter

qwenlm.github.io

91–100 of 303 posts

Re: Qwen2.5-VL-32B: Smarter and Lighter

#91
post #87
post #65

Silly question: how can OpenAI, Claude and all, have a valuation so large considering all the open source models? Not saying they will disappear or be tiny (closed models), but why so so so valuable?

People cannot normally invest in their competitors. It's not unlikely that chinese products may be banned / tarriff'd

There are non-Chinese open LLMs (Mistral, LLama, etc), so I don't think that explains it.

Re: Qwen2.5-VL-32B: Smarter and Lighter

#93
post #72
post #55

Earlier quoted context omitted.

Would a 40GB A6000 fully accommodate a 32B model? I assume an fp16 quantization is still necessary?

At FP16 you‘d need 64GB just for the weights, and it‘d be 2x as slow as a Q8 version, likely with little improvement. You‘ll also need space for attention and context etc, so 80-100GB (or even more) VRAM would be better. Many people „just“ use 4x consumer GPUs like the 3090 (24GB each) which scales well. They’d probably buy a mining rig, EPYC CPU, Mainboard with sufficient PCIe lanes, PCIe risers, 1600W PSU (might ne…

would it be better for energy efficiency and overall performance to use workstation cards like A5000 or A4000? Those can be found on eBay.

Re: Qwen2.5-VL-32B: Smarter and Lighter

#94
post #18

Earlier quoted context omitted.

Since we are on HN here, I can highly recommend open-webui with some OpenAI-compatible provider. I'm running with Deep Infra for more than a year now and am very happy. New models are usually available within one or two days after release. Also have some friends who use the service almost daily.

I'm using open-webui at home with a couple of different models. gemma2-9b fits in VRAM on a NV 3060 card + performs nicely.

What is the memory of your NV3060? 8GB?

Re: Qwen2.5-VL-32B: Smarter and Lighter

#95

Earlier quoted context omitted.

I've been waiting since November for 1, just 1*, model other than Claude than can reliably do agentic tool call loops. As long as the Chinese open models are chasing reasoning and benchmark maxxing vs. mid-2024 US private models, I'm very comfortable with somewhat ignoring these models. (this isn't idle prognostication hinging on my personal hobby horse. I got skin in the game, I'm virtually certain I have the only A…

is there some reason you cant train a 1b model to just do agentic stuff?

The Berkeley Function Calling Leaderboard [1] might be of interest to you. As of now, it looks like Hammer2.1-3b is the strongest model under 7 billion parameters. Its overall score is ~82% of GPT-4o's. There is also Hammer2.1-1.5b at 1.5 billion parameters that is ~76% of GPT-4o.

[1] https://gorilla.cs.berkeley.edu/leaderboard.html

Re: Qwen2.5-VL-32B: Smarter and Lighter

#96
post #22
post #17

Earlier quoted context omitted.

Pretty soon I won't be using any American models. It'll be a 100% Chinese open source stack. The foundation model companies are screwed. Only shovel makers (Nvidia, infra companies) and product companies are going to win.

I still don't get where the money for new open source models is going to come from once setting investor dollars on fire is no longer a viable business model. Does anyone seriously expect companies to keep buying and running thousands of ungodly expensive GPUs, plus whatever they spend on human workers to do labelling/tuning, and then giving away the spoils for free, forever?

There are lots of open-source projects that took many millions of dollars to create. Kubernetes, React, Postgres, Chromium, etc. etc.

This has clearly been part of a viable business model for a long time. Why should LLM models be any different?

Re: Qwen2.5-VL-32B: Smarter and Lighter

#97
post #85
post #4

Big day for open source Chinese model releases - DeepSeek-v3-0324 came out today too, an updated version of DeepSeek v3 now under an MIT license (previously it was a custom DeepSeek license). https://simonwillison.net/2025/Mar/24/deepseek/

Why do you keep promoting your blog on every LLM post?

I think they didn’t want to rewrite their post. It’s more substantial and researched than any comment here, and all their posts are full of information. I think they should get a pass, and calling it self-promotion is a stretch.

Re: Qwen2.5-VL-32B: Smarter and Lighter

#98

Earlier quoted context omitted.

it seems that this free version "may use your prompts and completions to train new models" https://openrouter.ai/deepseek/deepseek-chat-v3-0324:free do you think this needs attention?

good grief! people are okay with it when OpenAI and Google do it, but as soon as open source providers do it, people get defensive about it...

I trust big companies far more with my data than small ones.

Big companies have so much data they won't be having a human look at mine specifically. Some small place probably has the engineer looking at my logs as user #4.

Also, big companies have security teams whose job is securing the data, and it won't be going over some unencrypted link to cloudflare because OP was too lazy to set up Https certs.

Re: Qwen2.5-VL-32B: Smarter and Lighter

#99
post #85
post #4

Big day for open source Chinese model releases - DeepSeek-v3-0324 came out today too, an updated version of DeepSeek v3 now under an MIT license (previously it was a custom DeepSeek license). https://simonwillison.net/2025/Mar/24/deepseek/

Why do you keep promoting your blog on every LLM post?

Because I want people to read it. I only promote it if I think it's useful and relevant.
Post reply on HN