Live data from Hacker News

Qwen3-VL

qwen.ai

41–50 of 166 posts

Re: Qwen3-VL

#42

Earlier quoted context omitted.

With Qwen I went as stupid as I could: please provide the bounding box metadata for pytesseract for the above image. And it spat it out.

It’s funny that many of us say please. I don’t think it impacts the output, but it also feels wrong without it sometimes.

Depends on the model, but e.g. [1] found many models perform better if you are more polite. Though interestingly being rude can also sometimes improve performance at the cost of higher bias

Intuitively it makes sense. The best sources tend to be either of moderately high politeness (professional language) or 4chan-like (rude, biased but honest)

1: https://arxiv.org/pdf/2402.14531

Re: Qwen3-VL

#43
This model is literally amazing. Everyone should try to get their hands on a H100 and just call it a day.

Re: Qwen3-VL

#44

Earlier quoted context omitted.

It’s funny that many of us say please. I don’t think it impacts the output, but it also feels wrong without it sometimes.

Depends on the model, but e.g. [1] found many models perform better if you are more polite. Though interestingly being rude can also sometimes improve performance at the cost of higher bias Intuitively it makes sense. The best sources tend to be either of moderately high politeness (professional language) or 4chan-like (rude, biased but honest) 1: https://arxiv.org/pdf/2402.14531

When I want an LLM to be be brief, I will say things like "be brief", "don't ramble", etc.

When that fails, "shut the fuck up" always seems to do the trick.

Re: Qwen3-VL

#46
post #35

The Chinese are doing what they have been doing to the manufacturing industry as well. Take the core technology and just optimize, optimize, optimize for 10x the cost/efficiency. As simple as that. Super impressive. These models might be bechmaxxed but as another comment said, i see so many that it might as well be the most impressive benchmaxxing today, if not just a genuinely SOTA open source model. They even relea…

> Take the core technology and just optimize, optimize, optimize for 10x the cost/efficiency. As simple as that. Super impressive. This "just" is incorrect. The Qwen team invented things like DeepStack https://arxiv.org/abs/2406.04334 (Also I hate this "The Chinese" thing. Do we say "The British" if it came from a DeepMind team in the UK? Or what if there are Chinese born US citizens working in Paris for Mistral? Giv…

The Americans do that all the time. :P

Re: Qwen3-VL

#47

As I mentioned yesterday - I recently needed to process hundreds of low quality images of invoices (for a construction project). I had a script that had used pil/opencv, pytesseract, and open ai as a fallback. It still has a staggering number of failures. Today I tried a handful of the really poor quality invoices and Qwen spat out all the information I needed without an issue. What's crazier is it gave me the boundi…

So where did you load up Qwen and how did you supply the pdf or photo files? I don't know how to use these models, but want to learn

Re: Qwen3-VL

#48

As I mentioned yesterday - I recently needed to process hundreds of low quality images of invoices (for a construction project). I had a script that had used pil/opencv, pytesseract, and open ai as a fallback. It still has a staggering number of failures. Today I tried a handful of the really poor quality invoices and Qwen spat out all the information I needed without an issue. What's crazier is it gave me the boundi…

So where did you load up Qwen and how did you supply the pdf or photo files? I don't know how to use these models, but want to learn

LM Studio[0] is the best "i'm new here and what is this!?" tool for dipping your toes in the water.

If the model supports "vision" or "sound", that tool makes it relatively painless to take your input file + text and feed it to the model.

[0]: https://lmstudio.ai/

Re: Qwen3-VL

#49
post #27

Earlier quoted context omitted.

Interesting. I have in the past tried to get bounding boxes of property boundaries on satellite maps estimated by VLLM models but had no success. Do you have any tips on how to improve the results?

Do you have some example images and the prompt you tried?

also documented stack setup if could.

Re: Qwen3-VL

#50

Earlier quoted context omitted.

It’s funny that many of us say please. I don’t think it impacts the output, but it also feels wrong without it sometimes.

Depends on the model, but e.g. [1] found many models perform better if you are more polite. Though interestingly being rude can also sometimes improve performance at the cost of higher bias Intuitively it makes sense. The best sources tend to be either of moderately high politeness (professional language) or 4chan-like (rude, biased but honest) 1: https://arxiv.org/pdf/2402.14531

Bevore GPT5 was released I already had the feeling like the webui response was declining and I started to try to get more out of the responses and dissing it and saying how useless their response was did actually improve the output (I think).
Post reply on HN