Live data from Hacker News

GPT-5.2

openai.com

231–240 of 1001 posts

Re: GPT-5.2

#231

From GPT 5.1 Thinking: ARC AGI v2: 17.6% -> 52.9% SWE Verified: 76.3% -> 80% That's pretty good!

I don't think SWE Verified is an ideal benchmark, as the solutions are in the training dataset.

Re: GPT-5.2

#232

Earlier quoted context omitted.

Yep, the point we wanted to make here is that GPT-5.2's vision is better, not perfect. Cherrypicking a perfect output would actually mislead readers, and that wasn't our intent.

That would be a laudable goal, but I feel like it's contradicted by the text: > Even on a low-quality image, GPT‑5.2 identifies the main regions and places boxes that roughly match the true locations of each component I would not consider it to have "identified the main regions" or to have "roughly matched the true locations" when ~1/3 of the boxes have incorrect labels . The remark "even on a low-quality image" is n…

[deleted]

Re: GPT-5.2

#233

Earlier quoted context omitted.

Arc-AGI is just an iq test. I don’t see the problem with training it to be good at iq tests because that’s a skill that translates well.

Exactly. In principle, at least, the only way to overfit to Arc-AGI is to actually be that smart. Edit: if you disagree, try actually TAKING the Arc-AGI 2 test, then post.

I would not be so sure. You can always prep to the test.

Re: GPT-5.2

#234

Is it me, or did it still get at least three placements of components (RAM and PCIe slots, plus it's DisplayPort and not HDMI) in the motherboard image[0] completely wrong? Why would they use that as a promotional image? 0: https://images.ctfassets.net/kftzwdyauwt9/6lyujQxhZDnOMruN3f...

Yep, the point we wanted to make here is that GPT-5.2's vision is better, not perfect. Cherrypicking a perfect output would actually mislead readers, and that wasn't our intent.

You know what would be great? If it had added some boxes with “might be X or Y, but not sure”.

Re: GPT-5.2

#235

The thing about OpenAI is their models never fit anywhere for me. Yes they maybe smart or even the smartest models but they are alway so fucking slow. The ChatGPT web app is literally usable for me. I ask simple task and it does most extreme shit jsut to get an answer that the same as Claude or Gemini. For example, I asked ChatGPT to take a chart and convert into a table. It went and cut up the image and zoomed in fo…

Are you using 5.1 Thinking? I tended to prefer Claude before this model.

I use models based on the task. They still seem specialized and better at specific tasks. If I have a question I tend to go to it. If I need code, I tend to go to Claude (Code).

I go to ChatGPT for questions I have because I value an accurate answer over a quick answer and, in my experience, it tends to give me more accurate answers because of its (over) willingness to go to the web for search results and question its instincts. Claude is much more likely to make an assumption and its search patterns aren't as thorough. The slow answers don't bother me because it's an expectation I have for how I use it and they've made that use case work really well with background processing and notifications.

Re: GPT-5.2

#236
post #31

For me the last remaining killer feature of ChatGPT is the quality of the voice chat. Do any of the competitors have something like that?

gemini live is a thing - never tried chaptgpt, are they not similar?

Not for my use case. I can open it up, and in restored classical Latin pronunciation say "Hi, my name is X, how are you?" and it will respond (also in Latin) "Hello X, I am well, thanks for asking. I hope you are doing great." Its pronunciation is not great, but intelligible. In the written transcript, it butchers what I say, but its responses look good, although sans macrons indicating phonemic vowel length.

Gemini responds in what I think is Spanish, or perhaps Portuguese.

However I can hand an 8 minute long 48k mono mp3 of a nuanced Latin speaker who nasalizes his vowels, and makes regular use of elision to Gemini-3-pro-preview and it will produce an accurate macronized Latin transcription. It's pretty mind blowing.

Re: GPT-5.2

#237

Earlier quoted context omitted.

Seems pretty false if you look at the model card and web site of Opus 4.5 that is… (check notes) their latest model.

Building a good model generally means it will do well on benchmarks too. The point of the speculation is that Anthropic is not focused on benchmaxxing which is why they have models people like to use for their day-to-day. I use Gemini, Anthropic stole $50 from me (expired and kept my prepaid credits) and I have not forgiven them yet for it, but people rave about claude for coding so I may try the model again through…

You could try Codex cli. I prefer it over Claude code now, but only slightly.

Re: GPT-5.2

#238

> Unlike the previous GPT-5.1 model, GPT-5.2 has new features for managing what the model "knows" and "remembers to improve accuracy. Dumb nit, but why not put your own press release through your model to prevent basic things like missing quote marks? Reminds me of that time an OAI released wildly inaccurate copy/pasted bar charts.

Maybe they did

Re: GPT-5.2

#240
post #204
post #125

Wow, there's a lot going on with this pelican riding a bicycle: https://gist.github.com/simonw/c31d7afc95fe6b40506a9562b5e83...

seems to be eating something

Probably a jellyfish. You're seeing the tentacles
Post reply on HN