Live data from Hacker News

GPT-5.2

openai.com

261–270 of 1001 posts

Re: GPT-5.2

#261
post #125

Wow, there's a lot going on with this pelican riding a bicycle: https://gist.github.com/simonw/c31d7afc95fe6b40506a9562b5e83...

What happens if you ask for a pterodactyl on a motorbike?

Would like to know how much they are optimizing for your pelican....

Re: GPT-5.2

#262

In other news, been using Devstral 2 (Ollama) with OpenCode, and while it's not as good as Claude Code, my initial sense it that it's nonetheless good enough and doesn't require me to send my data off my laptop. I kind of wonder how close we are to alternative (not from a major AI lab) models being good enough for a lot of productive work and data sovereignty being the deciding factor.

Wait, isn't Devstral2 (normal not small) 123b? What type of laptop do you have? MacBooks don't go over 128GiB

Re: GPT-5.2

#263

Earlier quoted context omitted.

We're also in benchmark saturation territory. I heard it speculated that Anthropic emphasizes benchmarks less in their publications because internally they don't care about them nearly as much as making a model that works well on the day-to-day

These models still consistently fail the only benchmark that matters: if I give you a task, can you complete it successfully without making shit up? Thus far they all fail. Code outputs don’t run, or variables aren’t captured correctly, or hallucinations are stated as factual rather than suspect or “I don’t know.” It’s 2000’s PC gaming all over again (“gotta game the benchmark!”).

To say that a model won't solve a problem is unfair. Claude Code, with Opus 4.5, has solved plenty of problems for me.

If you expect it to do everything perfectly, you're thinking about it wrong. If you can't get it to do anything perfectly, you're using it wrong.

Re: GPT-5.2

#264
post #209

Earlier quoted context omitted.

Yep, the point we wanted to make here is that GPT-5.2's vision is better, not perfect. Cherrypicking a perfect output would actually mislead readers, and that wasn't our intent.

When I saw that it labeled DP ports as HDMI I immediately decided that I am not going to touch this until it is at least 5x better with 95% accuracy with basic things. I don't see any advantage in using the tool.

That's a far more dangerous territory. A machine that is obviously broken will not get used. A machine that is subtly broken will propagate errors because it will have achieved a high enough trust level that it will actually get used.

Think 'Therac-25', it worked in 99.5% of the time. In fact it worked so well that reports of malfunctions were routinely discarded.

Re: GPT-5.2

#265

Earlier quoted context omitted.

gemini live is a thing - never tried chaptgpt, are they not similar?

Not for my use case. I can open it up, and in restored classical Latin pronunciation say "Hi, my name is X, how are you?" and it will respond (also in Latin) "Hello X, I am well, thanks for asking. I hope you are doing great." Its pronunciation is not great, but intelligible. In the written transcript, it butchers what I say, but its responses look good, although sans macrons indicating phonemic vowel length. Gemini…

I have to ask: What usecase requires you to speak Latin to the llm?

Re: GPT-5.2

#266

Earlier quoted context omitted.

We're also in benchmark saturation territory. I heard it speculated that Anthropic emphasizes benchmarks less in their publications because internally they don't care about them nearly as much as making a model that works well on the day-to-day

Arc-AGI is just an iq test. I don’t see the problem with training it to be good at iq tests because that’s a skill that translates well.

[deleted]

Re: GPT-5.2

#267
post #263

Earlier quoted context omitted.

These models still consistently fail the only benchmark that matters: if I give you a task, can you complete it successfully without making shit up? Thus far they all fail. Code outputs don’t run, or variables aren’t captured correctly, or hallucinations are stated as factual rather than suspect or “I don’t know.” It’s 2000’s PC gaming all over again (“gotta game the benchmark!”).

To say that a model won't solve a problem is unfair. Claude Code, with Opus 4.5, has solved plenty of problems for me. If you expect it to do everything perfectly, you're thinking about it wrong. If you can't get it to do anything perfectly, you're using it wrong.

That means you're probably asking it to do very simple things.

Re: GPT-5.2

#268

gpt-5.2 and gpt-5.2-chat-latest the same token price? Isn't the latter non-thinking and more akin to -nano or -mini?

No. It is the same model without reasoning.

So is maybe gpt-5.2 with reasoning set to 'none' identical to gpt-5.2-chat-latest in capabilities but perhaps with a different system (system) prompt? I notice chat-latest doesn't accept temperature or reasoning (which makes sense) parameters, so something is certainly different underneath?

Re: GPT-5.2

#269

Trying it now in Vscode Insiders with Github Copilot (codex crashes with HTTP 400 server errors), and it eventually started using sed and grep in shells instead of using the better tools it has access to. I guess this is not an issue to perform well in benchmarks.

to be fair I've seen the other sota models do this as well

Re: GPT-5.2

#270
“…where it outperforms industry professionals at well-specified knowledge work tasks spanning 44 occupations.”

What a sociopathic way to sell

Post reply on HN