Live data from Hacker News

Qwen2.5-VL-32B: Smarter and Lighter

qwenlm.github.io

51–60 of 303 posts

Re: Qwen2.5-VL-32B: Smarter and Lighter

#51

What is the recommended model to process a RAG of PDF text documents? I've seen some recommendations for Mistral:7b. Looking to run on a consumer pedestrian home PC (ollama) with a Nvidia 4060ti and Ryzen 5700x.

Apparently there are two versions of the 4060Ti, with 8GB and 16GB of VRAM respectively. I've got an 8GB 3060 that runs gemma2:9b nicely, and that will parse PDF files; gemma3:4b also seems to analyze PDFs decently.

Re: Qwen2.5-VL-32B: Smarter and Lighter

#52
post #31

Earlier quoted context omitted.

If nationalist propaganda counts as ads, that might already be supporting Chinese models. Ask them about Tiananmen Square. Any kind of media with zero or near zero copying/distribution costs becomes a deflationary race to the bottom. Someone will eventually release something that's free, and at that point nothing can compete with free unless it's some kind of very specialized offering. Then you run into a the problem…

These ads can also have ads blockers though. Perplexity released the deepseek r1 1331? ( I am not sure I forgot) It basically removes chinese censorships / yes you can ask it about the tiananmen square. I think the next iteration of these ai model ads would be sneaky which might be hard to remove Though it's funny you comment about chinese censorship yet american censorship is fine lol

XAI to the rescue!!1!

... (no, not the unintelligible one - the xplainable one)

Re: Qwen2.5-VL-32B: Smarter and Lighter

#53
post #17

Earlier quoted context omitted.

Pretty soon I won't be using any American models. It'll be a 100% Chinese open source stack. The foundation model companies are screwed. Only shovel makers (Nvidia, infra companies) and product companies are going to win.

I've been waiting since November for 1, just 1*, model other than Claude than can reliably do agentic tool call loops. As long as the Chinese open models are chasing reasoning and benchmark maxxing vs. mid-2024 US private models, I'm very comfortable with somewhat ignoring these models. (this isn't idle prognostication hinging on my personal hobby horse. I got skin in the game, I'm virtually certain I have the only A…

is there some reason you cant train a 1b model to just do agentic stuff?

Re: Qwen2.5-VL-32B: Smarter and Lighter

#54

Just don’t ask it about the tiananmen square massacre or you’ll get a security warning. Even if you rephrase it. It’ll happily talk about Bloody Sunday. Probably a great model, but it worries me that it has such restrictions. Sure OpenAI also has lots of restrictions, but this feels more like straight up censorship since it’ll happily go on about bad things the governments of the west have done.

Daily reminder that all commerical LLMs are going to align with the governments their corporations exist under.

https://imgur.com/a/censorship-much-CBxXOgt

It's not even nefarious: they don't want the model spewing out content that will get them in trouble in the most general sense. It just so happens most governments have things that will get you in trouble.

The US is very obsessed with voter manipulation these days, so OpenAI and Anthropic's models are extra sensitive if the wording implies they're being used for that.

China doesn't like talking about past or ongoing human rights violations, so their models will be extra sensitive about that.

Re: Qwen2.5-VL-32B: Smarter and Lighter

#55
post #6

32B is one of my favourite model sizes at this point - large enough to be extremely capable (generally equivalent to GPT-4 March 2023 level performance, which is when LLMs first got really useful) but small enough you can run them on a single GPU or a reasonably well specced Mac laptop (32GB or more).

32B don't fully fit 16GB of VRAM. Still fine for higher quality answers, worth the extra wait in some cases.

Would a 40GB A6000 fully accommodate a 32B model? I assume an fp16 quantization is still necessary?

Re: Qwen2.5-VL-32B: Smarter and Lighter

#56

Wish I knew better how to estimate what sized video card one needs. HuggingFace link says this is bfloat16, so at least 64GB? I guess the -7B might run on my 16GB AMD card?

https://huggingface.co/spaces/NyxKrage/LLM-Model-VRAM-Calcul...

That will help you quickly calculate the model VRAM usage as well as the VRAM usage of the context length you want to use. You can put "Qwen/Qwen2.5-VL-32B-Instruct" in the "Model (unquantized)" field. Funnily enough the calculator lacks the option to see without quantizing the model, usually because nobody worried about VRAM bothers running >8 bit quants.

Re: Qwen2.5-VL-32B: Smarter and Lighter

#57

Earlier quoted context omitted.

I've only recently started looking into running these models locally on my system. I have limited knowledge regarding LLMs and even more limited when it comes to building my own PC. Are there any good sources that I can read up on estimiating what would be hardware specs required for 7B, 13B, 32B .. etc size If I need to run them locally?

VRAM Required = Number of Parameters (in billions) × Number of Bytes per Parameter × Overhead[0]. [0]: https://twm.me/posts/calculate-vram-requirements-local-llms/

Thats neat! thanks

Re: Qwen2.5-VL-32B: Smarter and Lighter

#58
post #30

Earlier quoted context omitted.

Money from the Chinese defense budget? Everyone using these models undercuts US companies. Eventually China wins.

And wez the end user get open source models. Also china doesn't have access to that many gpus because of the chips act. And i hate it , i hate it when america sounds more communist than china who open sources their stuff because free markets. I actually think that more countries need to invest into AI and not companies wanting profit. This could be the decision that can impact the next century.

If only you knew how many terawatt hours were burned on biasing models to prevent them from becoming racist

Re: Qwen2.5-VL-32B: Smarter and Lighter

#59

Earlier quoted context omitted.

Nah, it's great for things that Western models are censored on. The True Hacker will keep an Eastern and Western model available, depending on what they need information on.

Wouldn’t they just run R1 locally and not have any censorship at all? The model isn’t censored at its core, it’s censored through the system prompt. Perplexity and Huggingface have their own versions of R1 that is not censored.

I tried R1 through Kagi and it’s similarly censored. Even the distill of llama running on Groq is censored.

Re: Qwen2.5-VL-32B: Smarter and Lighter

#60
post #4

Big day for open source Chinese model releases - DeepSeek-v3-0324 came out today too, an updated version of DeepSeek v3 now under an MIT license (previously it was a custom DeepSeek license). https://simonwillison.net/2025/Mar/24/deepseek/

it seems that this free version "may use your prompts and completions to train new models" https://openrouter.ai/deepseek/deepseek-chat-v3-0324:free do you think this needs attention?

good grief! people are okay with it when OpenAI and Google do it, but as soon as open source providers do it, people get defensive about it...
Post reply on HN