Live data from Hacker News

Qwen2.5-VL-32B: Smarter and Lighter

qwenlm.github.io

71–80 of 303 posts

Re: Qwen2.5-VL-32B: Smarter and Lighter

#71

Earlier quoted context omitted.

And wez the end user get open source models. Also china doesn't have access to that many gpus because of the chips act. And i hate it , i hate it when america sounds more communist than china who open sources their stuff because free markets. I actually think that more countries need to invest into AI and not companies wanting profit. This could be the decision that can impact the next century.

If only you knew how many terawatt hours were burned on biasing models to prevent them from becoming racist

To be honest, maybe I am going off topic but I wish for the level of innovation in the ai industry in the energy industry.

It feels as an outsider that very little progress is made on the energy issue. I genuinely think that ai can be accelerated so so much more if energy could be more cheap / green

Re: Qwen2.5-VL-32B: Smarter and Lighter

#72
post #55

Earlier quoted context omitted.

32B don't fully fit 16GB of VRAM. Still fine for higher quality answers, worth the extra wait in some cases.

Would a 40GB A6000 fully accommodate a 32B model? I assume an fp16 quantization is still necessary?

At FP16 you‘d need 64GB just for the weights, and it‘d be 2x as slow as a Q8 version, likely with little improvement. You‘ll also need space for attention and context etc, so 80-100GB (or even more) VRAM would be better.

Many people „just“ use 4x consumer GPUs like the 3090 (24GB each) which scales well. They’d probably buy a mining rig, EPYC CPU, Mainboard with sufficient PCIe lanes, PCIe risers, 1600W PSU (might need to limit the GPUs to 300W), and 128GB RAM. Depending what you pay for the GPUs that‘ll be 3.5-4.5k

Re: Qwen2.5-VL-32B: Smarter and Lighter

#73
post #22
post #17

Earlier quoted context omitted.

Pretty soon I won't be using any American models. It'll be a 100% Chinese open source stack. The foundation model companies are screwed. Only shovel makers (Nvidia, infra companies) and product companies are going to win.

I still don't get where the money for new open source models is going to come from once setting investor dollars on fire is no longer a viable business model. Does anyone seriously expect companies to keep buying and running thousands of ungodly expensive GPUs, plus whatever they spend on human workers to do labelling/tuning, and then giving away the spoils for free, forever?

I think the only people who will ever make money are the shovel makers, the models will always be free because you’ll just get open source models chasing the paid ones and never being all that far behind, especially when this S curve growth phase slows down.

Re: Qwen2.5-VL-32B: Smarter and Lighter

#74
post #6

32B is one of my favourite model sizes at this point - large enough to be extremely capable (generally equivalent to GPT-4 March 2023 level performance, which is when LLMs first got really useful) but small enough you can run them on a single GPU or a reasonably well specced Mac laptop (32GB or more).

I just started self hosting as well on my local machine, been using https://lmstudio.ai/ Locally for now. I think the 32b models are actually good enough that I might stop paying for ChatGPT plus and Claude. I get around 20 tok/second on my m3 and I can get 100 tok/second on smaller models or quantized. 80-100 tok/second is the best for interactive usage if you go above that you basically can’t read as fast as it gen…

what spec is your local mac?

Re: Qwen2.5-VL-32B: Smarter and Lighter

#75

So today is Qwen. Tomorrow a new SOTA model from Google apparently, R2 next week. We haven't hit the wall yet.

Google's announcements are mostly vaporware anyway. Btw, where is Gemini Ultra 1 ? how about Gemini Ultra 2?

It is already on the LLM arena right, codename nebula? But you are right they can fuck up their releases royally.

Re: Qwen2.5-VL-32B: Smarter and Lighter

#76
post #18

Earlier quoted context omitted.

it seems that this free version "may use your prompts and completions to train new models" https://openrouter.ai/deepseek/deepseek-chat-v3-0324:free do you think this needs attention?

Since we are on HN here, I can highly recommend open-webui with some OpenAI-compatible provider. I'm running with Deep Infra for more than a year now and am very happy. New models are usually available within one or two days after release. Also have some friends who use the service almost daily.

I too run openweb-ui locally and use deepinfra.com as my backend. It has been working very well, and I am quite happy with deepinfra's pricing and privacy policy.

I have set up the same thing at work for my colleagues, and they find it better than openai for their tasks.

Re: Qwen2.5-VL-32B: Smarter and Lighter

#77
post #30

Earlier quoted context omitted.

Money from the Chinese defense budget? Everyone using these models undercuts US companies. Eventually China wins.

And wez the end user get open source models. Also china doesn't have access to that many gpus because of the chips act. And i hate it , i hate it when america sounds more communist than china who open sources their stuff because free markets. I actually think that more countries need to invest into AI and not companies wanting profit. This could be the decision that can impact the next century.

China has allowed quite a bit of market liberalism, so it isn’t that surprising if their AI stuff is responding to the market.

But, I don’t really see the connection on the flip side. Why should proprietary AI be associated with communism? If anything I guess a communist handling of AI would also be to share the model.

Re: Qwen2.5-VL-32B: Smarter and Lighter

#78
post #65

Silly question: how can OpenAI, Claude and all, have a valuation so large considering all the open source models? Not saying they will disappear or be tiny (closed models), but why so so so valuable?

It's user base and brand. Just like with Pepsi and Coca Cola. There's a reason OpenAI ran a Super Bowl ad.

Re: Qwen2.5-VL-32B: Smarter and Lighter

#79
post #22
post #17

Earlier quoted context omitted.

Pretty soon I won't be using any American models. It'll be a 100% Chinese open source stack. The foundation model companies are screwed. Only shovel makers (Nvidia, infra companies) and product companies are going to win.

I still don't get where the money for new open source models is going to come from once setting investor dollars on fire is no longer a viable business model. Does anyone seriously expect companies to keep buying and running thousands of ungodly expensive GPUs, plus whatever they spend on human workers to do labelling/tuning, and then giving away the spoils for free, forever?

Many sources, Chinese government could be one.

Re: Qwen2.5-VL-32B: Smarter and Lighter

#80

So today is Qwen. Tomorrow a new SOTA model from Google apparently, R2 next week. We haven't hit the wall yet.

> We haven't hit the wall yet.

The models are iterative improvements, but I haven't seen night and day differences since GPT3 and 3.5

Post reply on HN