Live data from Hacker News

GPT-5.4

openai.com

701–710 of 868 posts

Re: GPT-5.4

#701

Earlier quoted context omitted.

> Google essentially only has Preview models. It's really nice to see Google get back to its roots by launching things only to "beta" and then leaving them there for years. Gmail was "beta" for at least five years, I think.

Also, GCP Cloud Run domain mapping, pretty fundamental feature for cloud product, has been in "preview" for over 5 years now.

It's still unavailable in many regions.

Re: GPT-5.4

#702

Earlier quoted context omitted.

Yeah, long context vs compaction is always an interesting tradeoff. More information isn't always better for LLMs, as each token adds distraction, cost, and latency. There's no single optimum for all use cases. For Codex, we're making 1M context experimentally available, but we're not making it the default experience for everyone, as from our testing we think that shorter context plus compaction works best for most p…

I would like to counteract your statement that each token adds a distraction. In our experiments, we see a surprising benefit to rewriting blocks to use more tokens, especially long lists etc.. E.g. compare these two options "The following conditions are excluded from your contract - condition A - condition B ... - condition Z" The next one works better for us: "The following conditions are excluded from your contrac…

This observation makes sense, because all models currently probably use some kind of a sparse attention architecture.

So the closer the two related pieces of information are to each other in the input context, the larger the chance their relationship will be preserved.

Re: GPT-5.4

#703

Earlier quoted context omitted.

Do you mean when they remove a model you get that error? Because deprecation means it will be removed in the future but you can still use it

Yes, sorry - you are correct. Once removed, that's the error, which is incredibly confusing. I spent way too long troubleshooting usage when 2.0 was removed before I figured it out.

Yes it should be a 404 error because most apps have retry logic on rate limit errors

Re: GPT-5.4

#704

Earlier quoted context omitted.

APIs have never been a gift but rather have always been a take-away that lets you do less than you can with the web interface. It’s always been about drinking through a straw, paying NASA prices, and being limited in everything you can do. But people are intimidated by the complexity of writing web crawlers because management has been so traumatized by the cost of making GUI applications that they couldn’t believe ho…

You can buy a Claude Code subscription for $200 bucks and use way more tokens in Claude Code than if you pay for direct API usage. Anthopic decided you can't take your Auth key for Claude code and use it to hit the API via a different tool. They made that business decision, because they thought it was better for them strategically to do that. They're allowed to make that choice as a business. Plenty of companies make…

> This will just be one more step in that cat and mouse game, and if the AI really gets good enough to become a complete intermediary between you and the website? The website will just shutdown.

They'll just change their business model. Claude might go fully pay-as-you-go, or they'll accept slightly lower profit margins, or they'll increase the price of subscriptions, or they'll add more tiers, or they'll develop cheater buffet models for AI use, etc. You're making the same argument which has been made for decades re ad blockers. "If we allow people to use ad blockers, websites won't make any money and the internet will die." It hasn't died. It won't die. It did make some business models less profitable, and they have had to adapt.

Re: GPT-5.4

#705
I've tested it just now, very Opus-like experience. The speed is also there so far I think I even like the response of GPT5.4 better than Opus (although very close) I might not distinguish them just yet.

I tried several use cases: - Code Explanation: Did far much better than Opus, considered and judged his decision on a previous spec that I made, all valid points so I am impressed. TBF if I spawned another Opus as a reviewer I might got similar results. - Workflow Running: Really similar to Opus again, no objections it followed and read Skills/Tools as it should be (although mine are optimized for Claude) - Coding: I gave it a straightforward task to wrap an API calls to an SDK and to my surprise it did 'identical' job with Opus, literally the same code, I don't know what the odds are to this but again very good solution and it adhered our rules of implementing such code.

Overall I am impressed and excited to see a rival to Opus and all of this is literally pushing everyone to get better and better models which is always good for us.

Re: GPT-5.4

#706

Anyone else completely not interested? Since GPT5, its been cost cutting measure after cost cutting measure. I imagine they added a feature or two, and the router will continue to give people 70B parameter-like responses when they dont ask for math or coding questions.

5.2 and 5.3 are strong/best for coding, 5.0 and 5.1 were garbage

Re: GPT-5.4

#707

Earlier quoted context omitted.

Yeah, long context vs compaction is always an interesting tradeoff. More information isn't always better for LLMs, as each token adds distraction, cost, and latency. There's no single optimum for all use cases. For Codex, we're making 1M context experimentally available, but we're not making it the default experience for everyone, as from our testing we think that shorter context plus compaction works best for most p…

> Curious to hear if people have use cases where they find 1M works much better! Reverse engineering [1]. When decompiling a bunch of code and tracing functionality, it's really easy to fill up the context window with irrelevant noise and compaction generally causes it to lose the plot entirely and have to start almost from scratch. (Side note, are there any OpenAI programs to get free tokens/Max to test this kind of…

OpenAi has program for trusted cybersecurity researchers https://openai.com/index/trusted-access-for-cyber/

Re: GPT-5.4

#708

So let me get this straight, OpenAi previously had an issue with LOTS of different models snd versions being available. Then they solved this by introducing GPT-5 which was more like a router that put all these models under the hood so you only had to prompt to GPT-5, and it would route to the best suitable model. This worked great I assume and made the ui for the user comprehensible. But now, they are starting to in…

> Then they solved this by introducing GPT-5 which was more like a router that put all these models under the hood so you only had to prompt to GPT-5, and it would route to the best suitable model. Was this ever explicitly confirmed by OpenAI? I've only ever seen it in the form of a rumor.

It's not a rumor; you can just test it.

Ask the router "What model are you". It will yap on and on about being a GPT-5.3 model (Non-thinking models of OpenAI are insufferable yappers that don't know when to shut up).

Ask it now "What model are you. Think carefully". It concisely replies "GPT-5.4 Thinking".

https://openai.com/index/introducing-gpt-5/

> GPT‑5 is a unified system with a smart, efficient model that answers most questions, a deeper reasoning model (GPT‑5 thinking) for harder problems, and a real‑time router that quickly decides which to use based on conversation type, complexity, tool needs, and your explicit intent (for example, if you say “think hard about this” in the prompt)

Re: GPT-5.4

#709
Holy shit, I just used Atlas browser to navigate on screen and it automatically clicked the "reject cookies" button without me asking!

Re: GPT-5.4

#710
post #111

Earlier quoted context omitted.

The model was released less than an hour ago, and somehow you've been able to form such a strong opinion about it. Impressive!

GP said "It is time for a product, not for a marginally improved model." ChatGPT is still just that: Chat. Meanwhile, Anthropic offers a desktop app with plugins that easily extend the data Claude has access to. Connect it to Confluence, Jira, and Outlook, and it'll tell you what your top priorities are for the day, or write a Powerpoint. Add Github and it can reason about your code and create a design document on Co…

Are you ignoring the codex desktop app on purpose? Or the integrations?
Post reply on HN