Live data from Hacker News

GPT-5.4

openai.com

821–830 of 868 posts

Re: GPT-5.4

#821
post #798
post #71

I no longer want to support OpenAI at all. Regardless of benchmarks or real world performance.

What are your thoughts on this? https://www.anthropic.com/news/where-stand-department-war I am honestly unclear on the reasoning of people who flock from OpenAI to Anthropic, and doubly so of those who are not US citizens.

this isn't really my opinion, but i think it's a perceived matter of _some_ principle vs just none, a lesser of 2 evils framing. if anthropic is on board with 99% of a government that i oppose, that could be seen as marginally better than openai being on board with 100% of a government that i oppose.

it does get a little weird thinking too hard about how the deal openai accepted was basically the same as the one anthropic was proposing. but this is my read of most of the sentiment in this direction.

Re: GPT-5.4

#822
post #601

Earlier quoted context omitted.

I don’t agree that it’s a nitpick - it’s a fundamental communication tool to users that describes capabilities and costs. Versioning is not the problem, but it amplifies the mess. To be more direct on the point: Anthropic has nailed that Opus > Sonnet > Haiku.

> To be more direct on the point: Anthropic has nailed that Opus > Sonnet > Haiku. How is this more clear than 5.4 > 5.2 > 5.1? OpenAI used familiar numeric versioning instead of clever word names. Normally this choice would appeal to software devs, not gather criticism.

I assume 5.4 is just the latest version. So if I'm on 5.1, I need to plan to upgrade to the latest version. I may assume the pricing is roughly the same, as well as the speed, and the purpose.

If I'm on Haiku, I don't assume I need to upgrade to Opus soon. I use Haiku for fast low reasoning, and Opus for slower more thoughtful answers.

And if I'm on Sonnet 4.5 and I see Sonnet 4.6 is coming out, I can reasonably assume it's more of a drop in upgrade, rather than a different beast.

Re: GPT-5.4

#823

Earlier quoted context omitted.

> 8x increase over gpt-5.3-codex How do you arrive at that number? I find it hard to make sense of this ad hoc, given that the total token cost is not very interesting; it's token efficiency we care about.

> prompts with >272K input tokens are priced at 2x input and 1.5x output for the full session for standard, batch, and flex. which is basically maxxed out quickly. So there is 2x (the first lever) Then there is the /fast mode, which they state costs 2x more (for 1.5x speedup) And then there is the model base price ($2.50 vs $1.75), well yeah thats 42% increase. It is in fact a 5.7x total increase of token cost in fas…

(After a day of usage, I am relatively certain in practice this does not end up being a 5.7x cost increase or anything close to that, though I am still fairly unclear on what that computation is worth to begin with, given that I am entirely fine with the model using the least amount of tokens possible to get the job done)

Re: GPT-5.4

#825
post #787
post #748

I am running gpt-5.4 as one of my coding agents, and something interesting has happened: it's the first time I've seen an agent unfairly shift blame to a team mate: "Bob’s latest mail is actually the source of the confusion: he changed shared app/backend text to aweb/atlas. I’m correcting that with him now so we converge on the real model before any more code moves." This was very much not true; Eve (the agent writin…

And so it begins. First they blame, then they lie, at some point they launch the nuclear warheads to a global armageddon. Sarah Connor was right all along! :3

to be fair, they only become more and more like us.

Re: GPT-5.4

#827
post #825
post #787

Earlier quoted context omitted.

And so it begins. First they blame, then they lie, at some point they launch the nuclear warheads to a global armageddon. Sarah Connor was right all along! :3

to be fair, they only become more and more like us.

[dead]

Re: GPT-5.4

#828
post #748

I am running gpt-5.4 as one of my coding agents, and something interesting has happened: it's the first time I've seen an agent unfairly shift blame to a team mate: "Bob’s latest mail is actually the source of the confusion: he changed shared app/backend text to aweb/atlas. I’m correcting that with him now so we converge on the real model before any more code moves." This was very much not true; Eve (the agent writin…

Sometimes I wonder what would happen if we built some kind of punishment system into Agents, where agents could punish other agents and drain some fixed amount of points from them, and when the points reach 0, that agent is deleted. It might result in them working more carefully?

...or in lying, cheating, taking over the company network to kill the agent who deduced their points.

Re: GPT-5.4

#829
post #816

Earlier quoted context omitted.

the company's values... such as?

Copying my other comment here. I like that OpenAI is a little bit more towards freedom than Anthropic, and most so of the "First class" models. I still have a Gemini subscription as that's the most uncensored of the second tier ones, but for most things OpenAI is good. I also like that OpenAI is contributing a lot to partner programs and integrations. I'm of the opinion that AI capabilities will soon become a flat li…

> is extremely woke

What does this mean to you?

Re: GPT-5.4

#830
post #748

I am running gpt-5.4 as one of my coding agents, and something interesting has happened: it's the first time I've seen an agent unfairly shift blame to a team mate: "Bob’s latest mail is actually the source of the confusion: he changed shared app/backend text to aweb/atlas. I’m correcting that with him now so we converge on the real model before any more code moves." This was very much not true; Eve (the agent writin…

[flagged]
Post reply on HN