Live data from Hacker News

Qwen2.5-VL-32B: Smarter and Lighter

qwenlm.github.io

301–303 of 303 posts

Re: Qwen2.5-VL-32B: Smarter and Lighter

#301
post #47

Earlier quoted context omitted.

This is the reason IMO. Fundamentally China right now is better at manufacturing (e.g. robotics). AI is the complement to this - AI increases the demand for tech manufactured goods. Whereas America is in the opposite position w.r.t which side is their advantage (i.e. the software). AI for China is an enabler into a potentially bigger market which is robots/manufacturing/etc. Commoditizing the AI/intelligence part mea…

This may make sense if there is a centralized force to dictate how much these Chinese foundational model companies charge for their models. I know in the west people just blanketly believes that the state controls everything in China. However it can't be further from the truth. Most of the Chinese foundational model companies like moonshot, 01.ai, minimax, etc used to try to make money on those models. The VC money r…

I'm not saying there is a centralised force - I didn't say the government per se. Its enough to say many of the models coming out of China - the AI portion isn't their main income source especially for the major models that people are hyping up (Qwen, DeepSeek, etc). This model (Qwen) from Alibaba is a side model more likely complimenting their main business and cloud offerings. DeepSeek started as a way to use AI for trading models firstly; then spun up on the side. I'm more speaking about China's general position - for them AI seems to be more of a compliment than the main business as compared say to the major AI labs in America (ex Google). My opinion is that robotics in particular just extends that going forward.

Given as you say the long term cost of AI models is marginally zero, I don't think this is a bad position to be in.

Re: Qwen2.5-VL-32B: Smarter and Lighter

#302
post #299

Earlier quoted context omitted.

Yes, exactly, and no trivially: Manus is Sonnet with tools

Right. Apparently they also claim it's more than that: https://xcancel.com/peakji/status/1898997311646437487

No, they don't, that's just a bunch of other stuff (ex. Something something we don't differ from academic papers on agents (???))

Re: Qwen2.5-VL-32B: Smarter and Lighter

#303
post #14

Does anyone know how making the models multimodal impacts their text capabilities? The article is claiming this achieves good performance on pure text as well, but I'm curious if there is any analysis on how much impact it usually has. I've seen some people claim it should make the models better at text, but I find that a little difficult to believe without data.

I am having a hard time finding controlled testing, but the premise is straightforward: different modalities encourage different skills and understandings. Text builds up more formal idea tokenization and strengthens logic/reasoning while images require it learns a more robust geometric intuition. Since these learnings are applied to the same latent space, the strengths can be cross-applied. The same applies to human…

That comparison actually makes human reasoning abilities more impressive.

Helen Keller still learned robust generalizations.

Post reply on HN