Anecdotal, opinion: Gpt is really good in vision stuff, or at least their MoE seems to be really cohesive. From my experience Claude models can be really good at language but the moment they need to look at a picture and decide why the design is not good what parts need improvement it degrades a lot. My easiest benchmark is giving them a screenshot of a feature in my app and tell it "identify non-normative UI blocks…
GPT 5.6 Sol is the best "vision" model OpenAI ever released
101–110 of 194 posts
Re: GPT 5.6 Sol is the best "vision" model OpenAI ever released
#102Overall I've been hooked on using agents from different companies for what they are best at (Thanks to Theo). Fable is expensive, but unmatched for planning and top level organization of other agents. Sol is fast, will persistantly go after goals (sometimes to its detriment), and does well with computer use.
Re: GPT 5.6 Sol is the best "vision" model OpenAI ever released
#103It is funny to me seeing Sol used for what a "traditional" AI model can do already (counting pills). We have vision models for our pharmacy and I could never imagine taking the latency hit to use a Sol in our robotics, it would be likely 25-50x slower.
I’m evaluating these VLMs to figure out which ones are good enough to auto-annotate my data, so I can fine-tune my detector.
I wrote a bit more about this here: https://x.com/skalskip92/status/2080334344061694429?s=20
Re: GPT 5.6 Sol is the best "vision" model OpenAI ever released
#104Re: GPT 5.6 Sol is the best "vision" model OpenAI ever released
#105I didn't expect Gemini 3.5 Flash to top basically every metric in this article.
Re: GPT 5.6 Sol is the best "vision" model OpenAI ever released
#106The summary "There are still clear limits. Gemini 3.5 Flash remains a better practical choice [than GPT 5.6 Sol] for high-volume detection and counting in our benchmark, especially at its price." seems rather understated ! GPT 5.6 Sol was outperformed on all benchmarks by Gemini 3.5 Flash, apart from a single exception (OCR) where Fable was the winner. Gemini 3.5 Flash not only outperformed GPT 5.6 Sol, but did so at…
Yeah I was thinking about giving Luna a go with my PDF data extraction, but I think I‘ll stay on Gemini. It does a very good job.
For example, 3.7 Flash is #1 on MMLU Pro and AA’s agentic spreadsheets/docs benchmark, etc. Yes, beating Fable.
Agentic coding is only one dimension.
Re: GPT 5.6 Sol is the best "vision" model OpenAI ever released
#107Re: GPT 5.6 Sol is the best "vision" model OpenAI ever released
#108The summary "There are still clear limits. Gemini 3.5 Flash remains a better practical choice [than GPT 5.6 Sol] for high-volume detection and counting in our benchmark, especially at its price." seems rather understated ! GPT 5.6 Sol was outperformed on all benchmarks by Gemini 3.5 Flash, apart from a single exception (OCR) where Fable was the winner. Gemini 3.5 Flash not only outperformed GPT 5.6 Sol, but did so at…
Hi, I’m the author of this blog post. I wrote it about 4 weeks ago, and the VLM world is moving so fast that it’s already kinda outdated. I think Gemini 3.7 Flash might be a better choice now, especially when you factor in the price. Here’s a comparison of the best low-cost models I put together last week. What’s crazy is that Gemini 3.7 Flash is now 50% off on OpenRouter, and this chart doesn’t even account for that…
Re: GPT 5.6 Sol is the best "vision" model OpenAI ever released
#109It is funny to me seeing Sol used for what a "traditional" AI model can do already (counting pills). We have vision models for our pharmacy and I could never imagine taking the latency hit to use a Sol in our robotics, it would be likely 25-50x slower.
Hi! I’m the author of this blog. I’m evaluating these VLMs to figure out which ones are good enough to auto-annotate my data, so I can fine-tune my detector. I wrote a bit more about this here: https://x.com/skalskip92/status/2080334344061694429?s=20
Re: GPT 5.6 Sol is the best "vision" model OpenAI ever released
#110The summary "There are still clear limits. Gemini 3.5 Flash remains a better practical choice [than GPT 5.6 Sol] for high-volume detection and counting in our benchmark, especially at its price." seems rather understated ! GPT 5.6 Sol was outperformed on all benchmarks by Gemini 3.5 Flash, apart from a single exception (OCR) where Fable was the winner. Gemini 3.5 Flash not only outperformed GPT 5.6 Sol, but did so at…