Live data from Hacker News

Qwen2.5-VL-32B: Smarter and Lighter

qwenlm.github.io

261–270 of 303 posts

Re: Qwen2.5-VL-32B: Smarter and Lighter

#261
post #47

Earlier quoted context omitted.

Maybe from NVIDIA? "Commoditize your product's complement". https://www.joelonsoftware.com/2002/06/12/strategy-letter-v/

This is the reason IMO. Fundamentally China right now is better at manufacturing (e.g. robotics). AI is the complement to this - AI increases the demand for tech manufactured goods. Whereas America is in the opposite position w.r.t which side is their advantage (i.e. the software). AI for China is an enabler into a potentially bigger market which is robots/manufacturing/etc. Commoditizing the AI/intelligence part mea…

This may make sense if there is a centralized force to dictate how much these Chinese foundational model companies charge for their models. I know in the west people just blanketly believes that the state controls everything in China. However it can't be further from the truth. Most of the Chinese foundational model companies like moonshot, 01.ai, minimax, etc used to try to make money on those models. The VC money raised by those companies are in them to make money, not to voluntarily advance state competativeness. Deepseek is just an outlier backed by a billionaire. This billionaire has long been given money to various charities by hundered of millions per year before deepseek. Open-source SOTA models are not out-of-character move for him given his track record.

The thing is, model is in effect a piece of software that has almost 0 marginal cost. You just need a few, maybe even one company to release SOTA models consistently to really crash the valuation of every model companies because every one can acquire that single piece of software without cost to leave other model companies by themselves. The foundational model scene is basically in an extremely unstable state readily to return to a stable state of the model cost goes to 0. You really don't need the state competition assumption to explain the current state of affairs.

Re: Qwen2.5-VL-32B: Smarter and Lighter

#262
post #260

Earlier quoted context omitted.

Does quantised MLX support vision though? Is UV the best way to run it?

uv is just a Python package manager. No idea why they thought it was relevant to mention that

Because that one-liner will result in the model instantly running on your machine, which is much more useful than trying to figure out all the dependencies, invariably failing, and deciding that technology is horrible and that all you ever wanted was to be a carpenter.

Re: Qwen2.5-VL-32B: Smarter and Lighter

#263

Earlier quoted context omitted.

Does Grok still deny that trans women are women?

That’s a red herring. The original point was about censorship on Middle East topics.

It sounded like you were saying that grok is doing none of the censoring that anything else is doing. And I don't think it's a red herring, since it was definitely doing that in the past

Re: Qwen2.5-VL-32B: Smarter and Lighter

#264
post #241
post #212

Earlier quoted context omitted.

The problem with china is, they will have to figure out latency. Right now DeepSeek models hosted in china are having very high latency. It could because of DDoS and not strong enough infrastructure but probably also because of Great Firewall, runtime censoring prompt and servers physical location (big ping to US and EU countries).

> Right now DeepSeek models hosted in china are having very high latency. If you are talking about DeepSeek's own hosted API service. It's because they deliberately decided to run the service in heavily overloaded conditions and have very aggressive batching policy to extract more out of their (limited) H800s. Yes, for some reason (the reason I heard is "our boss don't want to run such a business" which sounds absurd…

> the reason I heard is "our boss don't want to run such a business" which sounds absurd

Liang gave up the No.1 Chinese hedge fund position to create AGI, he has very good chance to short the entire US share market and pocket some stupid amount of $ when R2 is released, he has pretty much unlimited support from local and central Chinese government. Trying to make some pennies from hosting models is not going to sustain what he enjoys now.

Re: Qwen2.5-VL-32B: Smarter and Lighter

#265

It's still BF16 model. Deepseek has proved that fp8 is more cost-effectiveness than fp16, isn't it valid for dozens-B model?

I don't understand what's your point? The optimiser is not fp8 anyway, is it? It's just the weights. I think the extent of fp8 "effectiveness" is greatly exaggerated. Yes, DeepSeek did use fp8, and they did implement it nicely, but it doesn't mean everybody is now got to be using fp8 all of the sudden.

Re: Qwen2.5-VL-32B: Smarter and Lighter

#266
post #212

Earlier quoted context omitted.

What's your take on why certain countries will have it cheaper and subsidies being at the forefront? An energy driven race to the bottom, is perhaps what you mean? I would suppose I have been seeing that China is ahead on their Renewables plan compared to the rest of the world, and they still have the lead on coal energy, so they'd likely be the winners on that front. But did you actually mean something else?

The problem with china is, they will have to figure out latency. Right now DeepSeek models hosted in china are having very high latency. It could because of DDoS and not strong enough infrastructure but probably also because of Great Firewall, runtime censoring prompt and servers physical location (big ping to US and EU countries).

Surely ping time is basically irrelevant dealing with LLMs? It has to be dwarfed by inference time.

Re: Qwen2.5-VL-32B: Smarter and Lighter

#267

Earlier quoted context omitted.

AMD's limitation is more of a software problem than a hardware problem at this point.

But it’s still surprising they haven’t. People would be motivated as hell if they launched GPUs with twice the amount of VRAM. It’s not as simple as just soldering some more in but still.

They sort of have. I'm using a 7900xtx, which has 24gb of vram. The next competitor would be a 4090, which would cost more than double today; granted, that would be much faster.

Technically there is also the 3090, which is more comparable price wise. I don't know about performance, though.

VRAM is supply limited enough that going bigger isn't as easy as it sounds. AMD can probably sell as much as they get their hands on, so they may as well still more GPUs, too.

Re: Qwen2.5-VL-32B: Smarter and Lighter

#268

Earlier quoted context omitted.

These ads can also have ads blockers though. Perplexity released the deepseek r1 1331? ( I am not sure I forgot) It basically removes chinese censorships / yes you can ask it about the tiananmen square. I think the next iteration of these ai model ads would be sneaky which might be hard to remove Though it's funny you comment about chinese censorship yet american censorship is fine lol

There are lots of "alliterated" versions of models too, which is where people will essentially remove the models ability to reject responding to a prompt. The huihui r1 14b alliterated had some trouble telling me about tiananmen square, basically dodging the question by telling me about itself, but after some coaxing I was able to get the info out of it. I say this because I think that the Perplexity model is tuned o…

Abliterated? Alliterated LLMs might be fun though…

Re: Qwen2.5-VL-32B: Smarter and Lighter

#269

Earlier quoted context omitted.

China has allowed quite a bit of market liberalism, so it isn’t that surprising if their AI stuff is responding to the market. But, I don’t really see the connection on the flip side. Why should proprietary AI be associated with communism? If anything I guess a communist handling of AI would also be to share the model.

My reasoning for proprietary AI to be associated with communism is that they aren't competing in a free market way where everyone does one thing and do its best. They are simultaneously trying to do all things internally. For example , Chatgpt etc. self hosts them on their own gpu and they can generate 10tk/s or something. Now there exists groq , cerebras who can do token generation of 4000 tk/s but they kind of requ…

It seems to me like they are acting like true capitalists; they seem very happy with the idea that capital (rather than labor) gives them the right to profit. But, they don’t seem to be too attached to free-market-ism.

Re: Qwen2.5-VL-32B: Smarter and Lighter

#270

Earlier quoted context omitted.

That’s a red herring. The original point was about censorship on Middle East topics.

It sounded like you were saying that grok is doing none of the censoring that anything else is doing. And I don't think it's a red herring, since it was definitely doing that in the past

I did not argue for Grok, but was showing that Qwen--to my surprise--is censoring topics that are rather censored in the West, not in China. And Grok--on the contrary--did not censor said topic.
Post reply on HN