Live data from Hacker News

GPT-5.4

openai.com

181–190 of 868 posts

Re: GPT-5.4

#181

Anyone else feel that it’s exhausting keeping up with the pace of new model releases. I swear every other week there’s a new release!

If you think about it there shouldn't really be a reason to care as long as things don't get worse.

Presumably this is where it'll evolve to with the product just being the brand with a pricing tier and you always get {latest} within that, whatever that means (you don't have to care). They could even shuffle models around internally using some sort of auto-like mode for simpler questions. Again why should I care as long as average output is not subjectively worse.

Just as I don't want to select resources for my SaaS software to use or have that explictly linked to pricing, I don't want to care what my OpenAI model or Anthropic model is today, I just want to pay and for it to hopefully keep getting better but at a minimum not get worse.

Re: GPT-5.4

#182
I wouldn't trust any of these benchmarks unless they are accompanied by some sort of proof other than "trust me bro". Also not including the parameters the models were run at (especially the other models) makes it hard to form fair comparisons. They need to publish, at minimum, the code and runner used to complete the benchmarks and logs.

Not including the Chinese models is also obviously done to make it appear like they aren't as cooked as they really are.

Re: GPT-5.4

#183
post #50

I use ChatGPT primarily for health related prompts. Looking at bloodwork, playing doctor for diagnosing minor aches/pains from weightlifting, etc. Interesting, the "Health" category seems to report worse performance compared to 5.2.

I've done the same, and I tested the same prompts with Claude and Google, and they both started hallucinating my blood results and supplement stack ingredients. Hopefully this new model doesn't fall on this. Claude and Google are dangerously unusable on the subject of health, from my experience.

Re: GPT-5.4

#184
post #81

5.4 vs 5.3-Codex? Which one is better for coding?

Related question:

- Do they have the same context usage/cost particularly in a plan?

They've kept 5.3-Codex along with 5.4, but is that just for user-preference reasons, or is there a trade-off to using the older one? I'm aware that API cost is better, but that isn't 1:1 with plan usage "cost."

Re: GPT-5.4

#185

Earlier quoted context omitted.

Why do so many people in the comments want 4o so bad?

Someone correct me if I'm wrong, but seemingly a lot of the people who found a "love interest" in LLMs seems to have preferred 4o for some reason. There was a lot of loud voices about that in the subreddit r/MyBoyfriendIsAI when it initially went away.

I think it's time for an https://hotornot.com for AI models.

Re: GPT-5.4

#186
post #128

Earlier quoted context omitted.

Benchmarks don't capture a lot - relative response times, vibes, what unmeasured capabilities are jagged and which are smooth, etc. I find there's a lot of difference between models - there are things which Grok is better than ChatGPT for that the benchmarks get inverted, and vice versa. There's also the UI and tools at hand - ChatGPT image gen is just straight up better, but Grok Imagine does better videos, and is f…

Gemini 3.1 slaps all other models at subtle concurrency bugs, sql and js security hardening when reviewing . (Obviously haven’t tested gpt 5.4 yet.) It’s a required step for me at this point to run any and all backend changes through Gemini 3.1 pro.

I have a few standard problems I throw at AI to see if they can solve them cleanly, like visualizing a neural network, then sorting each neuron in each layer by synaptic weights, largest to smallest, correctly reordering any previous and subsequent connected neurons such that the network function remains exactly the same. You should end up with the last layer ordered largest to smallest, and prior layers shuffled accordingly, and I still haven't had a model one-shot it. I spent an hour poking and prodding codex a few weeks back and got it done, but it conceptually seems like it should be a one-shot problem.

Re: GPT-5.4

#187
post #69

These releases are lacking something. Yes, they optimised for benchmarks, but it’s just not all that impressive anymore. It is time for a product, not for a marginally improved model.

When did they stop putting competitor models on the comparison table btw? And yeh I mean the benchmark improvements are meh. Context Window and lack of real memory is still an issue.

Re: GPT-5.4

#188
post #15

If you don't want to click in, easy comparison with other 2 frontier models - https://x.com/OpenAI/status/2029620619743219811?s=20

how does 5.4-thinking have a lower FrontierMath score than 5.4-pro?

Re: GPT-5.4

#189
post #131

Earlier quoted context omitted.

They have AI psychosis and think it's their boyfriend. The 5.x series have terrible writing styles, which is one way to cut down on sycophancy.

Somebody on Twitter used Claude code to connect… toys… as mcps to Claude chat. We’ve seen nothing yet.

what.. :o

Re: GPT-5.4

#190
post #15

If you don't want to click in, easy comparison with other 2 frontier models - https://x.com/OpenAI/status/2029620619743219811?s=20

how does 5.4-thinking have a lower FrontierMath score than 5.4-pro?

Well 5.4-pro is the more expensive and more advanced version of 5.4-thinking so why wouldn't it?
Post reply on HN