Qwen2.5-VL-32B: Smarter and Lighter
211–220 of 303 posts
Re: Qwen2.5-VL-32B: Smarter and Lighter
#212Earlier quoted context omitted.
One possibility. Certain countries will always be able to produce open models cheaper than others. USA and Europe probably won't be able. However, due to national security and wanting to promote their models overseas instead of letting their competitors promote theirs, the governments of USA and Europe will subsidize models which will lead their competitors to (further?) subsidies. There is a promotional aspect as we…
What's your take on why certain countries will have it cheaper and subsidies being at the forefront? An energy driven race to the bottom, is perhaps what you mean? I would suppose I have been seeing that China is ahead on their Renewables plan compared to the rest of the world, and they still have the lead on coal energy, so they'd likely be the winners on that front. But did you actually mean something else?
Re: Qwen2.5-VL-32B: Smarter and Lighter
#213Earlier quoted context omitted.
"B" just means "billion". A 7B model has 7 billion parameters. Most models are trained in fp16, so each parameter takes two bytes at full precision. Therefore, 7B = 14GB of memory. You can easily quantize models to 8 bits per parameter with very little quality loss, so then 7B = 7GB of memory. With more quality loss (making the model dumber), you can quantize to 4 bits per parameter, so 7B = 3.5GB of memory. There ar…
Oh, I have a question, maybe you know. Assuming the same model sizes in gigabytes, which one to choose: a higher-B lower-bit or a lower-B higher-bit? Is there a silver bullet? Like “yeah always take 4-bit 13B over 8-bit 7B”. Or are same-sized models basically equal in this regard?
Re: Qwen2.5-VL-32B: Smarter and Lighter
#214Earlier quoted context omitted.
The average user won't self-host a model.
...yet
For small guys and everyone else.. it'll probably be cost neutral to keep paying OpenAi, Google etc directly rather than paying some cloud provider to host an at best on-par model at equivalent prices.
Re: Qwen2.5-VL-32B: Smarter and Lighter
#215Earlier quoted context omitted.
Since we are on HN here, I can highly recommend open-webui with some OpenAI-compatible provider. I'm running with Deep Infra for more than a year now and am very happy. New models are usually available within one or two days after release. Also have some friends who use the service almost daily.
I too run openweb-ui locally and use deepinfra.com as my backend. It has been working very well, and I am quite happy with deepinfra's pricing and privacy policy. I have set up the same thing at work for my colleagues, and they find it better than openai for their tasks.
I've tried LibreChat before, but the app is terrible at generating titles for chats instead of leaving it as "New Chat". Also it lacks a working Code Interpreter.
Re: Qwen2.5-VL-32B: Smarter and Lighter
#216Big day for open source Chinese model releases - DeepSeek-v3-0324 came out today too, an updated version of DeepSeek v3 now under an MIT license (previously it was a custom DeepSeek license). https://simonwillison.net/2025/Mar/24/deepseek/
it seems that this free version "may use your prompts and completions to train new models" https://openrouter.ai/deepseek/deepseek-chat-v3-0324:free do you think this needs attention?
Re: Qwen2.5-VL-32B: Smarter and Lighter
#217Earlier quoted context omitted.
The hard-to-swallow truth is that American models do the same thing regarding Israel/Palestine.
They probably don't though. Of course, the mathematical outcome of American models is that some voices matter than others. The mechanism is similar to how the free market works. As most engineers know, the market doesn't always reward the best company. For example, It might reward the first company. We can see the "hierarchy in voices" with the following example. I use the following prompts for Gemini: 1. Which situa…
Re: Qwen2.5-VL-32B: Smarter and Lighter
#218Earlier quoted context omitted.
I also have the 8GB 4060ti variant. Want to upgrade to a 4070 super, but prices on them are still ridiculous. Could be had for $599 a handful of months ago, now on ebay going for $750 plus. Thanks for the recommendations. I'll give gemma3:12b a try and if needed go down to gemma:4b.
May I ask why you don’t get a used 3090 with 24GB VRAM?
Re: Qwen2.5-VL-32B: Smarter and Lighter
#219Earlier quoted context omitted.
... or a TailScale network. I've been leaving open-webui running on my laptop on my desk and then going out into the word and accessing it from my phone via TailScale, works great.
I would use tail scale. But I specifically want to use open web-ui from a place I can’t install a Tailscale client
Re: Qwen2.5-VL-32B: Smarter and Lighter
#220Earlier quoted context omitted.
And wez the end user get open source models. Also china doesn't have access to that many gpus because of the chips act. And i hate it , i hate it when america sounds more communist than china who open sources their stuff because free markets. I actually think that more countries need to invest into AI and not companies wanting profit. This could be the decision that can impact the next century.
China has allowed quite a bit of market liberalism, so it isn’t that surprising if their AI stuff is responding to the market. But, I don’t really see the connection on the flip side. Why should proprietary AI be associated with communism? If anything I guess a communist handling of AI would also be to share the model.
For example , Chatgpt etc. self hosts them on their own gpu and they can generate 10tk/s or something.
Now there exists groq , cerebras who can do token generation of 4000 tk/s but they kind of require a open source model.
So that is why I feel its not really abiding by the true capitalist philosophy