Live data from Hacker News

GPT-5.4

openai.com

311–320 of 868 posts

Re: GPT-5.4

#312
post #111

Earlier quoted context omitted.

The model was released less than an hour ago, and somehow you've been able to form such a strong opinion about it. Impressive!

I am actually super impressed with Codex-5.3 extra high reasoning. Its a drop in replacement (infact better than Claude Opus 4.6. lately claude being super verbose going in circles in getting things resolved). I stopped using claude mostly and having a blast with Codex 5.3. looking forward to 5.4 in codex.

I still love Opus but it's just too expensive / eats usage limits.

I've found that 5.3-Codex is mostly Opus quality but cheaper for daily use.

Curious to see if 5.4 will be worth somewhat higher costs, or if I'll stick to 5.3-Codex for the same reasons.

Re: GPT-5.4

#315
post #238
post #222

Earlier quoted context omitted.

That last benchmark seemed like an impressive leg up against Opus until I saw the sneaky footnote that it was actually a Sonnet result. Why even include it then, other than hoping people don't notice?

It's only that one number that is for sonnet.

except for the webarena-verified

Re: GPT-5.4

#317
>Today, we’re releasing GPT‑5.3 Instant

>Today, we’re releasing GPT‑5.4 in ChatGPT (as GPT‑5.4 Thinking),

>Note that there is not a model named GPT‑5.3 Thinking

They held out for eight months without a confusing numbering scheme :)

Re: GPT-5.4

#318

It's interesting that they charge more for the > 200k token window, but the benchmark score seems to go down significantly past that. That's judging from the Long Context benchmark score they posted, but perhaps I'm misunderstanding what that implies.

[flagged]

I guess that you pay more for worse quality to unlock use cases that could maybe be solved by better context management.

Re: GPT-5.4

#319

Earlier quoted context omitted.

Ironically this would actually be a good thing. As we can see from Iran Claude doesn’t quite have these bugs ironed out yet…

This is the exact attitude that lead to a chat bot being used to identify a school for girls as a valid target. The chatbot cannot be held responsible. Whoever is using chatbots for selecting targets is incompetent and should likely face war crime charges.

What attitude exactly are you talking about? The one that says that if you’re going to morally sell out it would be better if you at least tried not to kill children?

Re: GPT-5.4

#320

Earlier quoted context omitted.

The self reported safety score for violence dropped from 91% to 83%.

What the hell is a "safety score for violence"?

It's making sure AI condemns violence perpetuated by people without power and sanctifies violence of those who have it.
Post reply on HN