Live data from Hacker News

GPT-5.4

openai.com

231–240 of 868 posts

Re: GPT-5.4

#231
post #222
post #15

If you don't want to click in, easy comparison with other 2 frontier models - https://x.com/OpenAI/status/2029620619743219811?s=20

That last benchmark seemed like an impressive leg up against Opus until I saw the sneaky footnote that it was actually a Sonnet result. Why even include it then, other than hoping people don't notice?

Sonnet was pretty close to (or better than) Opus in a lot of benchmarks, I don't think it's a big deal

Re: GPT-5.4

#232

The marquee feature is obviously the 1M context window, compared to the ~200k other models support with maybe an extra cost for generations beyond >200k tokens. Per the pricing page, there is no additional cost for tokens beyond 200k: https://openai.com/api/pricing/ Also per pricing, GPT-5.4 ($2.50/M input, $15/M output) is much cheaper than Opus 4.6 ($5/M input, $25/M output) and Opus has a penalty for its beta >200…

Why would some one use codex instead?

[deleted]

Re: GPT-5.4

#233

Earlier quoted context omitted.

> If this rate of progress is steady, though, this year is gonna be crazy. Do you want to make any concrete predictions of what we'll see at this pace? It feels like we're reaching the end of the S-curve, at least to me.

If you look at the difference in quality between gpt-2 and 3, it feels like a big step, but the difference between 5.2 and 5.4 is more massive, it's just that they're both similarly capable and competent. I don't think it's an S curve; we're not plateauing. Million token context windows and cached prompts are a huge space for hacking on model behaviors and customization, without finetuning. Research is proceeding at…

For 2026, I am really interested in seeing whether local models can remain where they are: ~1 year behind the state of the art, to the point where a reasonably quantized November 2026 local model running on a consumer GPU actually performs like Opus 4.5.

I am betting that the days of these AI companies losing money on inference are numbered, and we're going to be much more dependent on local capabilities sooner rather than later. I predict that the equivalent of Claude Max 20x will cost $2000/mo in March of 2027.

Re: GPT-5.4

#234
post #131

Earlier quoted context omitted.

Somebody on Twitter used Claude code to connect… toys… as mcps to Claude chat. We’ve seen nothing yet.

My computer ethics teacher was obsessed with 'teledildonics' 30 years ago. There's nothing new under the sun.

Was your teacher Ted Nelson?

Re: GPT-5.4

#236

Earlier quoted context omitted.

Someone correct me if I'm wrong, but seemingly a lot of the people who found a "love interest" in LLMs seems to have preferred 4o for some reason. There was a lot of loud voices about that in the subreddit r/MyBoyfriendIsAI when it initially went away.

I think it's time for an https://hotornot.com for AI models.

botornot?

Re: GPT-5.4

#238
post #222
post #15

If you don't want to click in, easy comparison with other 2 frontier models - https://x.com/OpenAI/status/2029620619743219811?s=20

That last benchmark seemed like an impressive leg up against Opus until I saw the sneaky footnote that it was actually a Sonnet result. Why even include it then, other than hoping people don't notice?

It's only that one number that is for sonnet.

Re: GPT-5.4

#239

Earlier quoted context omitted.

prompt> Hi we want to build a missile, here is the picture of what we have in the yard.

{ tools: [ { name: "nuke", description: "Use when sure.", ... { lat: number, long: number } } ] }

Just remember an ethical programmer would never write a function “bombBagdad”. Rather they would write a function “bombCity(target City)”.

Re: GPT-5.4

#240

The "RPG Game" example on the blogpost is one of the most impressive demo's of autonomous engineering I've seen. It's very similar to "Battle Brothers", and the fact that RPG games require art assets, AI for enemy moves, and a host of other logical systems makes it all the more impressive.

[flagged]
Post reply on HN