Cool! Pity they are not releasing a smaller A3B MoE model
Qwen3-VL
41–50 of 166 posts
Re: Qwen3-VL
#42Earlier quoted context omitted.
With Qwen I went as stupid as I could: please provide the bounding box metadata for pytesseract for the above image. And it spat it out.
It’s funny that many of us say please. I don’t think it impacts the output, but it also feels wrong without it sometimes.
Intuitively it makes sense. The best sources tend to be either of moderately high politeness (professional language) or 4chan-like (rude, biased but honest)
Re: Qwen3-VL
#43Re: Qwen3-VL
#44Earlier quoted context omitted.
It’s funny that many of us say please. I don’t think it impacts the output, but it also feels wrong without it sometimes.
Depends on the model, but e.g. [1] found many models perform better if you are more polite. Though interestingly being rude can also sometimes improve performance at the cost of higher bias Intuitively it makes sense. The best sources tend to be either of moderately high politeness (professional language) or 4chan-like (rude, biased but honest) 1: https://arxiv.org/pdf/2402.14531
When that fails, "shut the fuck up" always seems to do the trick.
Re: Qwen3-VL
#45Re: Qwen3-VL
#46The Chinese are doing what they have been doing to the manufacturing industry as well. Take the core technology and just optimize, optimize, optimize for 10x the cost/efficiency. As simple as that. Super impressive. These models might be bechmaxxed but as another comment said, i see so many that it might as well be the most impressive benchmaxxing today, if not just a genuinely SOTA open source model. They even relea…
> Take the core technology and just optimize, optimize, optimize for 10x the cost/efficiency. As simple as that. Super impressive. This "just" is incorrect. The Qwen team invented things like DeepStack https://arxiv.org/abs/2406.04334 (Also I hate this "The Chinese" thing. Do we say "The British" if it came from a DeepMind team in the UK? Or what if there are Chinese born US citizens working in Paris for Mistral? Giv…
Re: Qwen3-VL
#47As I mentioned yesterday - I recently needed to process hundreds of low quality images of invoices (for a construction project). I had a script that had used pil/opencv, pytesseract, and open ai as a fallback. It still has a staggering number of failures. Today I tried a handful of the really poor quality invoices and Qwen spat out all the information I needed without an issue. What's crazier is it gave me the boundi…
Re: Qwen3-VL
#48As I mentioned yesterday - I recently needed to process hundreds of low quality images of invoices (for a construction project). I had a script that had used pil/opencv, pytesseract, and open ai as a fallback. It still has a staggering number of failures. Today I tried a handful of the really poor quality invoices and Qwen spat out all the information I needed without an issue. What's crazier is it gave me the boundi…
So where did you load up Qwen and how did you supply the pdf or photo files? I don't know how to use these models, but want to learn
If the model supports "vision" or "sound", that tool makes it relatively painless to take your input file + text and feed it to the model.
[0]: https://lmstudio.ai/
Re: Qwen3-VL
#49Earlier quoted context omitted.
Interesting. I have in the past tried to get bounding boxes of property boundaries on satellite maps estimated by VLLM models but had no success. Do you have any tips on how to improve the results?
Do you have some example images and the prompt you tried?
Re: Qwen3-VL
#50Earlier quoted context omitted.
It’s funny that many of us say please. I don’t think it impacts the output, but it also feels wrong without it sometimes.
Depends on the model, but e.g. [1] found many models perform better if you are more polite. Though interestingly being rude can also sometimes improve performance at the cost of higher bias Intuitively it makes sense. The best sources tend to be either of moderately high politeness (professional language) or 4chan-like (rude, biased but honest) 1: https://arxiv.org/pdf/2402.14531