Wow, there's a lot going on with this pelican riding a bicycle: https://gist.github.com/simonw/c31d7afc95fe6b40506a9562b5e83...
Is that the first SVG pelican with drop shadows?
GPT-5.2
211–220 of 1001 posts
Re: GPT-5.2
#212Is it me, or did it still get at least three placements of components (RAM and PCIe slots, plus it's DisplayPort and not HDMI) in the motherboard image[0] completely wrong? Why would they use that as a promotional image? 0: https://images.ctfassets.net/kftzwdyauwt9/6lyujQxhZDnOMruN3f...
Yep, the point we wanted to make here is that GPT-5.2's vision is better, not perfect. Cherrypicking a perfect output would actually mislead readers, and that wasn't our intent.
Re: GPT-5.2
#213Great! It'll be SOTA for a couple of weeks until the quality degrades due to throttling. I'll stick with plug and play API instead.
Jump in and soak up that extra-discounted compute while the getting is good, kids! Personally, I recently retired so I just occasionally mess around with LLMs for casual hobby projects, so I've only ever used the free tier of all the providers. Having lived through the dot com bubble, I regret not soaking up more of the free and heavily subsidized stuff back then. Trying not to miss out this time. All this compute available for free or below cost won't last too much longer...
Re: GPT-5.2
#214The only table where they showed comparisons against Opus 4.5 and Gemini 3: https://x.com/OpenAI/status/1999182104362668275 https://i.imgur.com/e0iB8KC.png
100% on the AIME (assuming its not in the training data) is pretty impressive. I got like 4/15 when I was in HS...
Re: GPT-5.2
#215From GPT 5.1 Thinking: ARC AGI v2: 17.6% -> 52.9% SWE Verified: 76.3% -> 80% That's pretty good!
Note that GPT 5.2 newly supports a "xhigh" reasoning level, which could explain the better benchmarks. It'll be noteworthy to see the cost-per-task on ARC AGI v2.
Re: GPT-5.2
#216> Unlike the previous GPT-5.1 model, GPT-5.2 has new features for managing what the model "knows" and "remembers to improve accuracy. Dumb nit, but why not put your own press release through your model to prevent basic things like missing quote marks? Reminds me of that time an OAI released wildly inaccurate copy/pasted bar charts.
It does seem to raise fair questions about either the utility of these tools, or adoption inertia. If not even OpenAI feels compelled to integrate this kind of model-check into their pipeline, what's that say about the business world at-large? Is it that it's too onerous to set up, is it that it's too hard to get only true-positive corrections, is it that it's too low value for the effort?
Nothing. OpenAI is a terrible baseline to extrapolate anything from.
Re: GPT-5.2
#217Earlier quoted context omitted.
Yep, the point we wanted to make here is that GPT-5.2's vision is better, not perfect. Cherrypicking a perfect output would actually mislead readers, and that wasn't our intent.
That would be a laudable goal, but I feel like it's contradicted by the text: > Even on a low-quality image, GPT‑5.2 identifies the main regions and places boxes that roughly match the true locations of each component I would not consider it to have "identified the main regions" or to have "roughly matched the true locations" when ~1/3 of the boxes have incorrect labels . The remark "even on a low-quality image" is n…
Re: GPT-5.2
#218Re: GPT-5.2
#219Everything is still based on 4 4o still right? is a new model training just too expensive? They can consult deepseek team maybe for cost constrained new models.
I thought whenever the knowledge cutoff increased that meant they’d trained a new model, I guess that’s completely wrong?
Re: GPT-5.2
#220Earlier quoted context omitted.
Claude’s voice chat isn’t “native” though, is it? It feels like it’s speech-to-text-to-LLM and back.
You can test it by asking it to: change the pitch of its voice, make specific sounds (like laughter), differentiate between words that are spelled the same but pronounced differently (record and record), etc.