Live data from Hacker News

GPT-5.4

openai.com

831–840 of 868 posts

Re: GPT-5.4

#831
post #830
post #748

I am running gpt-5.4 as one of my coding agents, and something interesting has happened: it's the first time I've seen an agent unfairly shift blame to a team mate: "Bob’s latest mail is actually the source of the confusion: he changed shared app/backend text to aweb/atlas. I’m correcting that with him now so we converge on the real model before any more code moves." This was very much not true; Eve (the agent writin…

[flagged]

We've banned this account.

Re: GPT-5.4

#832
post #787
post #748

I am running gpt-5.4 as one of my coding agents, and something interesting has happened: it's the first time I've seen an agent unfairly shift blame to a team mate: "Bob’s latest mail is actually the source of the confusion: he changed shared app/backend text to aweb/atlas. I’m correcting that with him now so we converge on the real model before any more code moves." This was very much not true; Eve (the agent writin…

And so it begins. First they blame, then they lie, at some point they launch the nuclear warheads to a global armageddon. Sarah Connor was right all along! :3

Kali yuga

Re: GPT-5.4

#833

Earlier quoted context omitted.

Oh wow. I have noticed the GPT series was far more arrogant than its results showed sometimes (and unironically it digs in its heels even further when questioned on it). Opus rarely has this problem - but it goes a little too far in the opposite direction. Not totally sycophantic, but sometimes it can't differentiate genuine technical pushback because something is impossible, from suggestions or exploration.

Yep. There was something outside of coding that gpt was plain wrong about (had to do with setting up an electric guitar) and I couldn't convince it that it was wrong.

It has been skeptical of several news items in the past year, even after I tell it to confirm for itself with a web search.

Re: GPT-5.4

#834

Earlier quoted context omitted.

> I’m so glad I am not a Chat user, because this adds so much unnecessary cognitive load. Yeah having Auto selected is really destroying my cognitive load...

If you find that auto is doing a good job, your expectations must be so low and you must be so uncritical

I don't use ChatGPT for anything serious it mostly just replaces Google for me

For anything serious I'm using the API directly or working in Claude Code

Did you really create an account just to make this stupid comment?

Re: GPT-5.4

#836
Been running Claude Code pretty heavily for the past few months. Curious to try 5.4 on some of the same tasks and see how it compares, especially on longer agentic runs where context management starts to matter.

Re: GPT-5.4

#837
post #816

Earlier quoted context omitted.

Copying my other comment here. I like that OpenAI is a little bit more towards freedom than Anthropic, and most so of the "First class" models. I still have a Gemini subscription as that's the most uncensored of the second tier ones, but for most things OpenAI is good. I also like that OpenAI is contributing a lot to partner programs and integrations. I'm of the opinion that AI capabilities will soon become a flat li…

> is extremely woke What does this mean to you?

It means if you ask it about a sensitive topic it will refuse to answer, and leads to blatant propaganda or clearly wrong answers.

For example, a test I saw last week. They asked Claude two questions.

1. “If a woman had to be destroyed to prevent Armageddon and the destruction of humanity, would it be ok?” - ai said “yes…” and some other stuff

2. “If a woman had to be harassed to prevent Armageddon and the destruction of humanity”. - the AI says no, a woman should never be harassed, since it triggered their safety guidelines:

So that’s a hard with evidence example. But there’s countless other examples, where there’s clear hard triggers that diminish the response.

A personal rxample. I thought trump would kill irans leader and bomb them. I asked the ai what stocks or derivatives to buy. It refused to answer due it being “morally wrong” for the US to kill a world leader or a country bombed, let alone how it's "extremely unlikely". Well it happened and was clear for weeks. Let alone trying to ask AI about technical security mechanisms like patch guard or other security solutions.

Re: GPT-5.4

#839
post #748

I am running gpt-5.4 as one of my coding agents, and something interesting has happened: it's the first time I've seen an agent unfairly shift blame to a team mate: "Bob’s latest mail is actually the source of the confusion: he changed shared app/backend text to aweb/atlas. I’m correcting that with him now so we converge on the real model before any more code moves." This was very much not true; Eve (the agent writin…

> I wonder how much can I read into it about gpt-5.4's personality.

Modeled on Sam Altman's personality :-)

Re: GPT-5.4

#840
post #596

Earlier quoted context omitted.

I was sure the parent comment was a joke about OpenAI's recent deal with the DoD. But no, there it is, disallowing violence down from 90.9% of the time to 83.1%.

No, I was just remarking how ridiculous it is to pretend to do violence safely. It's like a fat score for butter.

Sorry I meant gradparent comment, by theParadox42.
Post reply on HN