Live data from Hacker News

Qwen3-VL

qwen.ai

91–100 of 166 posts

Re: Qwen3-VL

#91
post #14

China is winning the hearts of developers in this race so far. At least, they won mine already.

so.. why do you think they are trying this hard to win your heart?

Open source is communism after all? In any case, maybe everyone realized what Zuckerberg was also saying from the start and that is that models will be more of a utility, rather than advantage.

Re: Qwen3-VL

#92

As I mentioned yesterday - I recently needed to process hundreds of low quality images of invoices (for a construction project). I had a script that had used pil/opencv, pytesseract, and open ai as a fallback. It still has a staggering number of failures. Today I tried a handful of the really poor quality invoices and Qwen spat out all the information I needed without an issue. What's crazier is it gave me the boundi…

I would recommend taking a look at this service: https://learn.microsoft.com/en-us/rest/api/computervision/re...

Re: Qwen3-VL

#93
post #69

Imagine the demand for a 128GB/256GB/512GB unified memory stuffed hardware linux box shipping with Qwen models already up and running. Although I´m agAInst steps towards AGI, it feels safer to have these things running locally and disconnected from each other, than some giant GW cloud agentic data centers connected to everyone and everything.

I bought an GMKtec evo 2 that is a 128 GB unified memory system. Strong recommend.

That's AMD Ryzen AI Max+ 395, right? Lots of those boxes popping up recently, but isn't that dog slow? And I can't believe I'm saying this - but maybe RAM filled-up mac might be a better option?

Re: Qwen3-VL

#94

Extremely impressive, but can one really run these >200B param models on prem in any cost effective way? Even if you get your hands on cards with 80GB ram, you still need to tie them together in a low-latency high-BW manner. It seems to me that small/medium sized players would still need a third party to get inference going on these frontier-quality models, and we're not in a fully self-owned self-hosted place yet. I…

A Framework Desktop exposes 96GB of RAM for inference and costs a few thou USD.

You need memory on the GPU, not in the system itself (unless you have unified memory such as the M-architecture). So we're talking about cards like the H200 that have 141GB of memory and cost between 25 to 40k.

Re: Qwen3-VL

#95
Qwen has some really great models. I recently used qwen/qwen3-next-80b-a3b-thinking as a drop-in replacement for GPT-4.1-mini in an agent workflow. Cost 4 times less for input tokens and half for output, instant cost savings. As far as I can measure, system output has kept the same quality.

Re: Qwen3-VL

#96

Earlier quoted context omitted.

A Framework Desktop exposes 96GB of RAM for inference and costs a few thou USD.

You need memory on the GPU, not in the system itself (unless you have unified memory such as the M-architecture). So we're talking about cards like the H200 that have 141GB of memory and cost between 25 to 40k.

Did you casually glance at how the hardware in the Framework Desktop (Strix Halo) works before commenting?

Re: Qwen3-VL

#97
post #60

Earlier quoted context omitted.

Arguably they’ve already won. Check the names at the top the next time you see a paper from an American company, a lot of them are Chinese.

you can’t tell if someone is American or Chinese by looking at their name I actually claim something even stronger, which is it’s what’s in your heart that really determines if you’re American :)

The PRC espionage system doesn't care what passport you have or even where you are born. They have a broader and more ethnic-focus definition.

Re: Qwen3-VL

#98
post #90
post #21

Earlier quoted context omitted.

They still suck at explaining which model they serve is which, though. They also released today Qwen3-VL Plus [1] today alongside Qwen3-VL 235B [2] and they don't tell us which one is better. Note that Qwen3-VL-Plus is a very different model compared to Qwen-VL-Plus. Also, qwen-plus-2025-09-11 [3] vs qwen3-235b-a22b-instruct-2507 [4]. What's the difference? Which one is better? Who knows. You know it's bad when OpenA…

> They still suck at explaining which model they serve is which, though. "they" in this sentence probably applies to all "AI" companies. Even the naming/versioning of OpenAI models is ridiculous, and then you can never find out which is actually better for your needs. Every AI company writes several paragraphs of fluffy text with lots of hand waving, saying how this model is better for complex tasks while this other…

Both Deepseek and Claude are exceptions. Simple versions and Sonnet is overall worse but faster than Opus for the same version.

Re: Qwen3-VL

#99

Earlier quoted context omitted.

You need memory on the GPU, not in the system itself (unless you have unified memory such as the M-architecture). So we're talking about cards like the H200 that have 141GB of memory and cost between 25 to 40k.

Did you casually glance at how the hardware in the Framework Desktop (Strix Halo) works before commenting?

I didn't glace at it, I read it :-) The architecture is a 'unified memory bus', so yes the GPU has access to that memory.

My comment was a bit unfortunate as it implied I didn't agree with yours, sorry for that. I simply want to clarify that there's a difference between 'GPU memory' and 'system memory'.

The Frame.work desktop is a nice deal. I wouldn't buy the Ryzen AI+ myself, from what I read it maxes out at about 60 tokens / sec which is low for my use cases.

Re: Qwen3-VL

#100
post #35

The Chinese are doing what they have been doing to the manufacturing industry as well. Take the core technology and just optimize, optimize, optimize for 10x the cost/efficiency. As simple as that. Super impressive. These models might be bechmaxxed but as another comment said, i see so many that it might as well be the most impressive benchmaxxing today, if not just a genuinely SOTA open source model. They even relea…

> Take the core technology and just optimize, optimize, optimize for 10x the cost/efficiency. As simple as that. Super impressive. This "just" is incorrect. The Qwen team invented things like DeepStack https://arxiv.org/abs/2406.04334 (Also I hate this "The Chinese" thing. Do we say "The British" if it came from a DeepMind team in the UK? Or what if there are Chinese born US citizens working in Paris for Mistral? Giv…

The naming makes some sense here. It's backed by the very Chinese Alibaba and the government directly as well. It's almost a national project.
Post reply on HN