What is the recommended model to process a RAG of PDF text documents? I've seen some recommendations for Mistral:7b. Looking to run on a consumer pedestrian home PC (ollama) with a Nvidia 4060ti and Ryzen 5700x.
Qwen2.5-VL-32B: Smarter and Lighter
51–60 of 303 posts
Re: Qwen2.5-VL-32B: Smarter and Lighter
#52Earlier quoted context omitted.
If nationalist propaganda counts as ads, that might already be supporting Chinese models. Ask them about Tiananmen Square. Any kind of media with zero or near zero copying/distribution costs becomes a deflationary race to the bottom. Someone will eventually release something that's free, and at that point nothing can compete with free unless it's some kind of very specialized offering. Then you run into a the problem…
These ads can also have ads blockers though. Perplexity released the deepseek r1 1331? ( I am not sure I forgot) It basically removes chinese censorships / yes you can ask it about the tiananmen square. I think the next iteration of these ai model ads would be sneaky which might be hard to remove Though it's funny you comment about chinese censorship yet american censorship is fine lol
... (no, not the unintelligible one - the xplainable one)
Re: Qwen2.5-VL-32B: Smarter and Lighter
#53Earlier quoted context omitted.
Pretty soon I won't be using any American models. It'll be a 100% Chinese open source stack. The foundation model companies are screwed. Only shovel makers (Nvidia, infra companies) and product companies are going to win.
I've been waiting since November for 1, just 1*, model other than Claude than can reliably do agentic tool call loops. As long as the Chinese open models are chasing reasoning and benchmark maxxing vs. mid-2024 US private models, I'm very comfortable with somewhat ignoring these models. (this isn't idle prognostication hinging on my personal hobby horse. I got skin in the game, I'm virtually certain I have the only A…
Re: Qwen2.5-VL-32B: Smarter and Lighter
#54Just don’t ask it about the tiananmen square massacre or you’ll get a security warning. Even if you rephrase it. It’ll happily talk about Bloody Sunday. Probably a great model, but it worries me that it has such restrictions. Sure OpenAI also has lots of restrictions, but this feels more like straight up censorship since it’ll happily go on about bad things the governments of the west have done.
https://imgur.com/a/censorship-much-CBxXOgt
It's not even nefarious: they don't want the model spewing out content that will get them in trouble in the most general sense. It just so happens most governments have things that will get you in trouble.
The US is very obsessed with voter manipulation these days, so OpenAI and Anthropic's models are extra sensitive if the wording implies they're being used for that.
China doesn't like talking about past or ongoing human rights violations, so their models will be extra sensitive about that.
Re: Qwen2.5-VL-32B: Smarter and Lighter
#5532B is one of my favourite model sizes at this point - large enough to be extremely capable (generally equivalent to GPT-4 March 2023 level performance, which is when LLMs first got really useful) but small enough you can run them on a single GPU or a reasonably well specced Mac laptop (32GB or more).
32B don't fully fit 16GB of VRAM. Still fine for higher quality answers, worth the extra wait in some cases.
Re: Qwen2.5-VL-32B: Smarter and Lighter
#56Wish I knew better how to estimate what sized video card one needs. HuggingFace link says this is bfloat16, so at least 64GB? I guess the -7B might run on my 16GB AMD card?
That will help you quickly calculate the model VRAM usage as well as the VRAM usage of the context length you want to use. You can put "Qwen/Qwen2.5-VL-32B-Instruct" in the "Model (unquantized)" field. Funnily enough the calculator lacks the option to see without quantizing the model, usually because nobody worried about VRAM bothers running >8 bit quants.
Re: Qwen2.5-VL-32B: Smarter and Lighter
#57Earlier quoted context omitted.
I've only recently started looking into running these models locally on my system. I have limited knowledge regarding LLMs and even more limited when it comes to building my own PC. Are there any good sources that I can read up on estimiating what would be hardware specs required for 7B, 13B, 32B .. etc size If I need to run them locally?
VRAM Required = Number of Parameters (in billions) × Number of Bytes per Parameter × Overhead[0]. [0]: https://twm.me/posts/calculate-vram-requirements-local-llms/
Re: Qwen2.5-VL-32B: Smarter and Lighter
#58Earlier quoted context omitted.
Money from the Chinese defense budget? Everyone using these models undercuts US companies. Eventually China wins.
And wez the end user get open source models. Also china doesn't have access to that many gpus because of the chips act. And i hate it , i hate it when america sounds more communist than china who open sources their stuff because free markets. I actually think that more countries need to invest into AI and not companies wanting profit. This could be the decision that can impact the next century.
Re: Qwen2.5-VL-32B: Smarter and Lighter
#59Earlier quoted context omitted.
Nah, it's great for things that Western models are censored on. The True Hacker will keep an Eastern and Western model available, depending on what they need information on.
Wouldn’t they just run R1 locally and not have any censorship at all? The model isn’t censored at its core, it’s censored through the system prompt. Perplexity and Huggingface have their own versions of R1 that is not censored.
Re: Qwen2.5-VL-32B: Smarter and Lighter
#60Big day for open source Chinese model releases - DeepSeek-v3-0324 came out today too, an updated version of DeepSeek v3 now under an MIT license (previously it was a custom DeepSeek license). https://simonwillison.net/2025/Mar/24/deepseek/
it seems that this free version "may use your prompts and completions to train new models" https://openrouter.ai/deepseek/deepseek-chat-v3-0324:free do you think this needs attention?