Live data from Hacker News

Qwen3-VL

qwen.ai

141–150 of 166 posts

Re: Qwen3-VL

#141

Earlier quoted context omitted.

It did not, unfortunately. When CV failed gpt-4o failed as well. I even had a list of valid invoice numbers & dates to help the models. Still, most failed. Construction invoices are not great.

Did you try few-shotting examples when you hit problem cases? In my ziploc case, the model was failing if red sharpie was used vs black. A few shot hint fixed that.

Tbh, I had run the images through a few filters. The images that went through to AI were high contrast, black and white, with noise such as highlighters removed. I had tried 1 shot and few shot.

I think it was largely a formatting issue. Like some of these invoices have nonsense layouts. Perhaps Qwen works well because it doesn't assume left to right, top to bottom? Just speculating though

Re: Qwen3-VL

#142

Earlier quoted context omitted.

Of course they do, eventually. Also, it seems like they're not burning nearly as much money as some of their US competitors.

AI is part of China's 5 year plan and been given special blessing by Xi Jinping directly. That pretty much translates to blank checks from the party without much oversight or expected ROI. The approach is basically brute forcing something into existence rather than organically letting it grow. China is notorious for this approach, ghost cities, high speed rail to nowhere, solar panel production in the face of a huge…

[deleted]

Re: Qwen3-VL

#143

As I mentioned yesterday - I recently needed to process hundreds of low quality images of invoices (for a construction project). I had a script that had used pil/opencv, pytesseract, and open ai as a fallback. It still has a staggering number of failures. Today I tried a handful of the really poor quality invoices and Qwen spat out all the information I needed without an issue. What's crazier is it gave me the boundi…

i had success with tabula. you may not need ai. but fine if it works too.

Re: Qwen3-VL

#144

Earlier quoted context omitted.

With Qwen I went as stupid as I could: please provide the bounding box metadata for pytesseract for the above image. And it spat it out.

It’s funny that many of us say please. I don’t think it impacts the output, but it also feels wrong without it sometimes.

It's a good habit to build now in case AGI actually happens out of the blue.

Re: Qwen3-VL

#145

Earlier quoted context omitted.

So where did you load up Qwen and how did you supply the pdf or photo files? I don't know how to use these models, but want to learn

You can use their models here chat.qwenlm.ai, its their official website

I wouldn't recommend using anything that can transmit data back to the CCP. The model itself is fine since it's open source (and you can run it firewalled if you're really paranoid), but directly using Alibaba's AI chat website should be discouraged.

Re: Qwen3-VL

#146

Earlier quoted context omitted.

So where did you load up Qwen and how did you supply the pdf or photo files? I don't know how to use these models, but want to learn

LM Studio[0] is the best "i'm new here and what is this!?" tool for dipping your toes in the water. If the model supports "vision" or "sound", that tool makes it relatively painless to take your input file + text and feed it to the model. [0]: https://lmstudio.ai/

Jumping from this for visibility - LM Studio really is the best option out there. Ollama is another runtime that I've used, but I've found it makes too many assumptions about what a computer is capable of and it's almost impossible to override those settings. It often overloads weaker computers and underestimates stronger ones.

LM Studio isn't as "set it and forget it" as Ollama is, and it does have a bit of a learning curve. But if you're doing any kind of AI development and you don't want to mess around with writing llama-cpp scripts all the time, it really can't be beat (for now).

Re: Qwen3-VL

#147

The Chinese are doing what they have been doing to the manufacturing industry as well. Take the core technology and just optimize, optimize, optimize for 10x the cost/efficiency. As simple as that. Super impressive. These models might be bechmaxxed but as another comment said, i see so many that it might as well be the most impressive benchmaxxing today, if not just a genuinely SOTA open source model. They even relea…

> Take the core technology and just optimize, optimize, optimize for 10x the cost/efficiency.

This is what really grinds my gears about American AI and American technology in general lately, as an American myself. We used to do that! But over the last 10-15 years, it seems like all this country can do is try to throw more and more resources at something instead of optimizing what we already have.

Download more ram for this progressive web app.

Buy a Threadripper CPU to run this game that looks worse than the ones you played on the Nintendo Gamecube in the early 2000s.

Generate more electricity (hello Elon Musk).

Y'all remember your algorithms classes from college, right? Why not apply that here? Because China is doing just that, and frankly making us look stupid by comparison.

Re: Qwen3-VL

#148
post #13
post #3

That has got to be the most benchmarks I've ever seen posted with an announcement. Kudos for not just cherrypicking a favorable set.

We should stop reporting saturated benchmarks.

Yeah, especially since many of those are just poor targets after being out so long / contaminating too much.

Re: Qwen3-VL

#149
post #69

Earlier quoted context omitted.

I bought an GMKtec evo 2 that is a 128 GB unified memory system. Strong recommend.

That's AMD Ryzen AI Max+ 395, right? Lots of those boxes popping up recently, but isn't that dog slow? And I can't believe I'm saying this - but maybe RAM filled-up mac might be a better option?

A GMKtec or a Framework desktop with a Strix Halo/AI Max CPU is about the cheapest way to run a model that needs to fit into about 120GB of memory. Macs have twice the memory bandwidth of these units, so will run significantly faster, but they're also much more expensive. Technically, you could run these models on any desktop PC with 128GB of RAM, but that's a whole different level of "dog slow." It really depends on how much you're prepared to pay to run these bigger models locally.

Re: Qwen3-VL

#150
post #120

Earlier quoted context omitted.

Note that the small print on the GMKtec site says that prices do not include customs and VAT. Which seem to amount to 19% in the EU. So, almost 2,4K€.

There seems to be a EU shop as well, but I can't see it's without VAT, not even on checkout page. There's a 50 EUR discount code though.

Loooking closely, the shop does not seem to be located within the EU. And the 50€ discount does not apply to the 128GB config. Also, if you are interested, it might help to have a look into the user forum: https://de.gmktec.com/community/xenforum
Post reply on HN