Live data from Hacker News

GPT-5.4

openai.com

21–30 of 868 posts

Re: GPT-5.4

#21
post #15

If you don't want to click in, easy comparison with other 2 frontier models - https://x.com/OpenAI/status/2029620619743219811?s=20

Definitely don’t want to click in at x either.

Re: GPT-5.4

#22

Earlier quoted context omitted.

Up now. The OP has frequently gotten the scoop for new LLM releases and I am curious what their pipeline is.

curl the URL https://openai.com/index/introducing-gpt-5- ? until you get 200

Probably refresh the api models list every couple minutes instead. No one could have guessed the name of GPT-Codex-Spark

Re: GPT-5.4

#23

The marquee feature is obviously the 1M context window, compared to the ~200k other models support with maybe an extra cost for generations beyond >200k tokens. Per the pricing page, there is no additional cost for tokens beyond 200k: https://openai.com/api/pricing/ Also per pricing, GPT-5.4 ($2.50/M input, $15/M output) is much cheaper than Opus 4.6 ($5/M input, $25/M output) and Opus has a penalty for its beta >200…

GPT 5.3 codex had 400K context window btw

Re: GPT-5.4

#24
Bit concerning that we see in some cases significantly worse results when enabling thinking. Especially for Math, but also in the browser agent benchmark.

Not sure if this is more concerning for the test time compute paradigm or the underlying model itself.

Maybe I'm misunderstanding something though? I'm assuming 5.4 and 5.4 Thinking are the same underlying model and that's not just marketing.

Re: GPT-5.4

#26
post #21
post #15

If you don't want to click in, easy comparison with other 2 frontier models - https://x.com/OpenAI/status/2029620619743219811?s=20

Definitely don’t want to click in at x either.

Ditto, but I did anyways and enjoyed that OpenAI doesn't include the dogwater that is Grok on their scorecard.

Re: GPT-5.4

#29
can anyone compare the $200/mo codex usage limits with the $200/mo claude usage limits? It’s extremely difficult to get a feel for whether switching between the two is going to result in hitting limits more or less often, and it’s difficult to find discussion online about this.

In practice, if I buy $200/mo codex, can I basically run 3 codex instances simultaneously in tmux, like I can with claude code pro max, all day every day, without hitting limits?

Re: GPT-5.4

#30

Bit concerning that we see in some cases significantly worse results when enabling thinking. Especially for Math, but also in the browser agent benchmark. Not sure if this is more concerning for the test time compute paradigm or the underlying model itself. Maybe I'm misunderstanding something though? I'm assuming 5.4 and 5.4 Thinking are the same underlying model and that's not just marketing.

[flagged]
Post reply on HN