China is winning the hearts of developers in this race so far. At least, they won mine already.
so.. why do you think they are trying this hard to win your heart?
Qwen3-VL
91–100 of 166 posts
Re: Qwen3-VL
#92As I mentioned yesterday - I recently needed to process hundreds of low quality images of invoices (for a construction project). I had a script that had used pil/opencv, pytesseract, and open ai as a fallback. It still has a staggering number of failures. Today I tried a handful of the really poor quality invoices and Qwen spat out all the information I needed without an issue. What's crazier is it gave me the boundi…
Re: Qwen3-VL
#93Imagine the demand for a 128GB/256GB/512GB unified memory stuffed hardware linux box shipping with Qwen models already up and running. Although I´m agAInst steps towards AGI, it feels safer to have these things running locally and disconnected from each other, than some giant GW cloud agentic data centers connected to everyone and everything.
I bought an GMKtec evo 2 that is a 128 GB unified memory system. Strong recommend.
Re: Qwen3-VL
#94Extremely impressive, but can one really run these >200B param models on prem in any cost effective way? Even if you get your hands on cards with 80GB ram, you still need to tie them together in a low-latency high-BW manner. It seems to me that small/medium sized players would still need a third party to get inference going on these frontier-quality models, and we're not in a fully self-owned self-hosted place yet. I…
A Framework Desktop exposes 96GB of RAM for inference and costs a few thou USD.
Re: Qwen3-VL
#95Re: Qwen3-VL
#96Earlier quoted context omitted.
A Framework Desktop exposes 96GB of RAM for inference and costs a few thou USD.
You need memory on the GPU, not in the system itself (unless you have unified memory such as the M-architecture). So we're talking about cards like the H200 that have 141GB of memory and cost between 25 to 40k.
Re: Qwen3-VL
#97Earlier quoted context omitted.
Arguably they’ve already won. Check the names at the top the next time you see a paper from an American company, a lot of them are Chinese.
you can’t tell if someone is American or Chinese by looking at their name I actually claim something even stronger, which is it’s what’s in your heart that really determines if you’re American :)
Re: Qwen3-VL
#98Earlier quoted context omitted.
They still suck at explaining which model they serve is which, though. They also released today Qwen3-VL Plus [1] today alongside Qwen3-VL 235B [2] and they don't tell us which one is better. Note that Qwen3-VL-Plus is a very different model compared to Qwen-VL-Plus. Also, qwen-plus-2025-09-11 [3] vs qwen3-235b-a22b-instruct-2507 [4]. What's the difference? Which one is better? Who knows. You know it's bad when OpenA…
> They still suck at explaining which model they serve is which, though. "they" in this sentence probably applies to all "AI" companies. Even the naming/versioning of OpenAI models is ridiculous, and then you can never find out which is actually better for your needs. Every AI company writes several paragraphs of fluffy text with lots of hand waving, saying how this model is better for complex tasks while this other…
Re: Qwen3-VL
#99Earlier quoted context omitted.
You need memory on the GPU, not in the system itself (unless you have unified memory such as the M-architecture). So we're talking about cards like the H200 that have 141GB of memory and cost between 25 to 40k.
Did you casually glance at how the hardware in the Framework Desktop (Strix Halo) works before commenting?
My comment was a bit unfortunate as it implied I didn't agree with yours, sorry for that. I simply want to clarify that there's a difference between 'GPU memory' and 'system memory'.
The Frame.work desktop is a nice deal. I wouldn't buy the Ryzen AI+ myself, from what I read it maxes out at about 60 tokens / sec which is low for my use cases.
Re: Qwen3-VL
#100The Chinese are doing what they have been doing to the manufacturing industry as well. Take the core technology and just optimize, optimize, optimize for 10x the cost/efficiency. As simple as that. Super impressive. These models might be bechmaxxed but as another comment said, i see so many that it might as well be the most impressive benchmaxxing today, if not just a genuinely SOTA open source model. They even relea…
> Take the core technology and just optimize, optimize, optimize for 10x the cost/efficiency. As simple as that. Super impressive. This "just" is incorrect. The Qwen team invented things like DeepStack https://arxiv.org/abs/2406.04334 (Also I hate this "The Chinese" thing. Do we say "The British" if it came from a DeepMind team in the UK? Or what if there are Chinese born US citizens working in Paris for Mistral? Giv…