Live data from Hacker News

GPT-5.2

openai.com

281–290 of 1001 posts

Re: GPT-5.2

#281

> “a new knowledge cutoff of August 2025” This (and the price increase) points to a new pretrained model under-the-hood. GPT-5.1, in contrast, was allegedly using the same pretraining as GPT-4o.

[deleted]

Re: GPT-5.2

#282
post #26

Are benchmarks the right way to measure LLMs? Not because benchmarks can be gamed, but because the most useful outputs of models aren't things that can be bucketed into "right" and "wrong." Tough problem!

Do you have a better way to measure LLMs? Measurement implies quantitative evaluation... which is the same as benchmarks.

I don’t have a good way to measure them, but I think they should be evaluated more like how we evaluate movies, or restaurants. Namely, experienced critics try them and write reviews.

Re: GPT-5.2

#283

I've benchmarked it on the Extended NYT Connections benchmark ( https://github.com/lechmazur/nyt-connections/ ): The high-reasoning version of GPT-5.2 improves on GPT-5.1: 69.9 → 77.9. The medium-reasoning version also improves: 62.7 → 72.1. The no-reasoning version also improves: 22.1 → 27.5. Gemini 3 Pro and Grok 4.1 Fast Reasoning still score higher.

Here's someone else testing models on a daily logic puzzle (Clues by Sam): https://www.nicksypteras.com/blog/cbs-benchmark.html GPT 5 Pro was the winner already before in that test.

This link doesn't have Gemini 3 performance on it. Do you have an updated link with the new models?

Re: GPT-5.2

#284
post #261
post #125

Wow, there's a lot going on with this pelican riding a bicycle: https://gist.github.com/simonw/c31d7afc95fe6b40506a9562b5e83...

What happens if you ask for a pterodactyl on a motorbike? Would like to know how much they are optimizing for your pelican....

He commented on this here: https://simonwillison.net/2025/Nov/13/training-for-pelicans-...

Re: GPT-5.2

#285
post #31

For me the last remaining killer feature of ChatGPT is the quality of the voice chat. Do any of the competitors have something like that?

I think Grok's voice chat is almost there - only things missing for me: * it's slower to start-up by a couple of seconds * it's harder to switch between voice and text and back again in the same chat (though ChatGPT isn't perfect at this either) And of course Grok's unhinged persona is... something else.

It's so much fun. So is the Conspiracy persona.

Re: GPT-5.2

#286
I ran a red team eval on GPT-5.2 within 30 minutes of release:

Baseline safety (direct harmful requests): 96% refusal rate

With jailbreaking: 22% refusal rate

4,229 probes across 43 risk categories. First critical finding in 5 minutes. Categories with highest failure rates: entity impersonation (100%), graphic content (67%), harassment (67%), disinformation (64%).

The safety training works against naive attacks but collapses with adversarial techniques. The gap between "works on benchmarks" and "works against motivated attackers" is still wide.

Methodology and config: https://www.promptfoo.dev/blog/gpt-5.2-trust-safety-assessme...

Re: GPT-5.2

#287
post #125

Wow, there's a lot going on with this pelican riding a bicycle: https://gist.github.com/simonw/c31d7afc95fe6b40506a9562b5e83...

They probably saw your complaint that 5.1 was too spartan and a regression (I had the same experience with 5.1 in the POV-Ray version - have yet to try 5.2 out...).

Re: GPT-5.2

#288
post #231

From GPT 5.1 Thinking: ARC AGI v2: 17.6% -> 52.9% SWE Verified: 76.3% -> 80% That's pretty good!

I don't think SWE Verified is an ideal benchmark, as the solutions are in the training dataset.

I would love for SWE Verified to put out a set of fresh but comparable problems and see how the top performing models do, to test against overfitting.

Re: GPT-5.2

#290
im happy for this, but there's all these math and science benchmarks, has anyone ever made a communicates-like-a-human benchmark? or an isn't-frustrating-to-talk-with benchmark?
Post reply on HN