Live data from Hacker News

GPT 5.6 Sol is the best "vision" model OpenAI ever released

blog.roboflow.com

151–160 of 194 posts

Re: GPT 5.6 Sol is the best "vision" model OpenAI ever released

#151

Earlier quoted context omitted.

Curious why you didn't try Gemini 3 pro? That is the model I've been using for OCR entry of handwritten datasheets (JPGS of datasheets, structured JSON output). At my scale, the cost of 3 pro is basically not an issue, but if there are improvements in quality, I'd definitely be willing to explore other models

The “pro” moniker means nothing these models aren’t successors and barely have a common ancestor, they are independently baked in the training oven and assigned a semantic version randomly by someone trying to show initiative but not trying to do on the toes of the last guy who got promoted first So 3 pro is outdated and will likely never exit preview The “flash” and “lite” models are the real “pro” in colloquial ide…

They are smaller models, and you can tell. Small models make dumb common-sense mistakes that big models never do. This is the "smell" many talk about.

Re: GPT 5.6 Sol is the best "vision" model OpenAI ever released

#152

Earlier quoted context omitted.

The “pro” moniker means nothing these models aren’t successors and barely have a common ancestor, they are independently baked in the training oven and assigned a semantic version randomly by someone trying to show initiative but not trying to do on the toes of the last guy who got promoted first So 3 pro is outdated and will likely never exit preview The “flash” and “lite” models are the real “pro” in colloquial ide…

They are smaller models, and you can tell. Small models make dumb common-sense mistakes that big models never do. This is the "smell" many talk about.

Do you have cases where you still see 3.1 pro outperforming 3.7 flash?

Re: GPT 5.6 Sol is the best "vision" model OpenAI ever released

#153

Earlier quoted context omitted.

The “pro” moniker means nothing these models aren’t successors and barely have a common ancestor, they are independently baked in the training oven and assigned a semantic version randomly by someone trying to show initiative but not trying to do on the toes of the last guy who got promoted first So 3 pro is outdated and will likely never exit preview The “flash” and “lite” models are the real “pro” in colloquial ide…

They are smaller models, and you can tell. Small models make dumb common-sense mistakes that big models never do. This is the "smell" many talk about.

hasn't been an issue since 3.5 for me, what have you seen, say, in the last two months

Re: GPT 5.6 Sol is the best "vision" model OpenAI ever released

#156

Earlier quoted context omitted.

I think that is because people perceive OpenCV as 'hard to use' and LLMs as easy to use.

OpenCV is no longer hard to use, it just takes longer. Still, a little more complicated than asking LLM to count. To use an LLM, you just prompt it with an image + text saying "count the pills in this image". To use OpenCV, ... you just prompt an LLM with an image + text saying "count the pills in this image, using OpenCV instead of eyeballing it". (I like to throw in "produce intermediary artifacts so I can see the…

I no longer use it but never felt it was particularly complicated, but since the days of resnet there are much faster ways to the goal.

Re: GPT 5.6 Sol is the best "vision" model OpenAI ever released

#157

Gemini 3 Flash should really be included in this comparison. Or at least 3.7. In most of my testing, 3.5 and 3.6 were both a downgrade in terms of vision capabilities, relative to 3, and at a much higher cost. 3.7 is slightly better than 3, finally.

but 3.7 flash is expensive for img inputs no ?

Re: GPT 5.6 Sol is the best "vision" model OpenAI ever released

#158

Penny sample shown looks like failed EXIF orientation registered by the model/harness. The coins are correctly marked, it's rotated 90 degrees.

Hi! I’m the author of this blog. I had the same intuition, but together with the OpenAI team we figured out that the issue was image resolution. GPT-5.6 doesn’t handle large images well.

OpenAI team sounds like they've misidentified the root cause for this particular case then.

Re: GPT 5.6 Sol is the best "vision" model OpenAI ever released

#159
post #152

Earlier quoted context omitted.

They are smaller models, and you can tell. Small models make dumb common-sense mistakes that big models never do. This is the "smell" many talk about.

Do you have cases where you still see 3.1 pro outperforming 3.7 flash?

Yes, for complex questions of biology, physics, and analysis of anomalies.

3.7 Flash is better at coding, sure, but AI is not just for coding.

Re: GPT 5.6 Sol is the best "vision" model OpenAI ever released

#160

Earlier quoted context omitted.

They are smaller models, and you can tell. Small models make dumb common-sense mistakes that big models never do. This is the "smell" many talk about.

hasn't been an issue since 3.5 for me, what have you seen, say, in the last two months

For complex questions of biology, physics, and analysis of anomalies, 3.1 Pro is still better than 3.7 Flash for me.

3.7 Flash is better at coding, sure, but AI is not just for coding.

Post reply on HN