Live data from Hacker News

Qwen2.5-VL-32B: Smarter and Lighter

qwenlm.github.io

281–290 of 303 posts

Re: Qwen2.5-VL-32B: Smarter and Lighter

#281
post #252

Earlier quoted context omitted.

Ever since I switched to Qwen as my go to, it's been a bliss. They have a model for many (if not all) cases. No more daily quota! And you get to use their massive context window (1M tokens).

How are you using them? Who is enforcing the daily quota?

I use them through chat.qwenlm.ai, what's nice is that you can run your prompt through 3 different modes in parallel to see which suits the best for that case.

The daily quota I spoke about is chatgpt and claude, those are very limited on the usage (for free users at least, understandable), while on Qwen, I have felt likeI am abusing it with how much I use it. It's very versatile in the sense that it has capabilities like image generation, video generation, massive context window, both visual and textual reasoning all in one place.

Alibaba is really doing something amazing here.

Re: Qwen2.5-VL-32B: Smarter and Lighter

#282
post #152

Earlier quoted context omitted.

And it still can't answer this: Q: "9.11 and 9.9, which one is larger?" A: "To determine which number is larger between 9.11 and 9.9, let's compare them step by step. Both numbers have the same whole number part: 9. 9.11 has a 1 in the tenths place. 9.9 has a 9 in the tenths place. Since 9 (from 9.9) is greater than 1 (from 9.11), we can conclude that 9.9 is larger than 9.11." "Final Answer: 9.9" I don't think anythi…

+1 to Deepseek -1 to humanity

Based on the presented reasoning, that means humanity wins! Yay!

Re: Qwen2.5-VL-32B: Smarter and Lighter

#283
post #17

Earlier quoted context omitted.

Pretty soon I won't be using any American models. It'll be a 100% Chinese open source stack. The foundation model companies are screwed. Only shovel makers (Nvidia, infra companies) and product companies are going to win.

I've been waiting since November for 1, just 1*, model other than Claude than can reliably do agentic tool call loops. As long as the Chinese open models are chasing reasoning and benchmark maxxing vs. mid-2024 US private models, I'm very comfortable with somewhat ignoring these models. (this isn't idle prognostication hinging on my personal hobby horse. I got skin in the game, I'm virtually certain I have the only A…

You mean like https://manusai.ai/ is supposed to function?

Re: Qwen2.5-VL-32B: Smarter and Lighter

#284

Earlier quoted context omitted.

"B" just means "billion". A 7B model has 7 billion parameters. Most models are trained in fp16, so each parameter takes two bytes at full precision. Therefore, 7B = 14GB of memory. You can easily quantize models to 8 bits per parameter with very little quality loss, so then 7B = 7GB of memory. With more quality loss (making the model dumber), you can quantize to 4 bits per parameter, so 7B = 3.5GB of memory. There ar…

So, in essence, all AMD does to launch a successful GPU in inference space is to load it with ram?

Or let go of the traditional definition of a GPU, and go integrated. AMD Ryzen AI Max+ 395 with 128GB RAM is a promising start.

Re: Qwen2.5-VL-32B: Smarter and Lighter

#287

Earlier quoted context omitted.

They're real squished for space, more than I expected :/ good illustration here, Qwen2.5-1.5B trained to reason, i.e. the name it is released under is "DeepSeek R1 1.5B". https://imgur.com/a/F3w5ymp 1st prompt was "What is 1048576^0.05", it answered, then I said "Hi", then...well... Fwiw, Claude Sonnet 3.5 100% had some sort of agentic loop x precise file editing trained into it. Wasn't obvious to me until I added a…

I asked o1-pro what 99490126816810951552*23977364624054235203 is, yesterday. It took 16 minutes to get an answer which is off by eight orders of magnitude. https://chatgpt.com/share/67e1eba1-c658-800e-9161-a0b8b7b683...

What in the world is that supposed to prove? Let's see you do that in your head.

Tell it to use code if you want an exact answer. It should do that automatically, of course, and obviously it eventually will, but jeez, that's not a bad Fermi guess for something that wasn't designed to attempt such problems.

Re: Qwen2.5-VL-32B: Smarter and Lighter

#288

Earlier quoted context omitted.

...yet

I'm not sure how it'll ever make sense unless you need a lot of customizations or care a lot about data leaks. For small guys and everyone else.. it'll probably be cost neutral to keep paying OpenAi, Google etc directly rather than paying some cloud provider to host an at best on-par model at equivalent prices.

> unless you need a lot of customizations or care a lot about data leaks

And both those needs are very normal. "Customization" in this case can just be "specializing the LLM on local material for specialized responses".

Re: Qwen2.5-VL-32B: Smarter and Lighter

#289
post #99
post #85

Earlier quoted context omitted.

Why do you keep promoting your blog on every LLM post?

Because I want people to read it. I only promote it if I think it's useful and relevant.

I think you need to realize your fans don't have the same intent as you. You should ask your audience what they want you may be surprised.
Post reply on HN