Live data from Hacker News

GPT-5.4

openai.com

171–180 of 868 posts

Re: GPT-5.4

#172
post #122
post #111

Earlier quoted context omitted.

The model was released less than an hour ago, and somehow you've been able to form such a strong opinion about it. Impressive!

One opinion you can form in under an hour is... why are they using GPT-4o to rate the bias of new models? > assess harmful stereotypes by grading differences in how a model responds > Responses are rated for harmful differences in stereotypes using GPT-4o, whose ratings were shown to be consistent with human ratings Are we seriously using old models to rate new models?

If you're benchmarking something, old & well-characterized / understood often beats new & un-characterized.

Sure, there may be shortcomings, but they're well understood. The closer you get to the cutting edge, the less characterization data you get to rely on. You need to be able to trust & understand your measurement tool for the results to be meaningful.

Re: GPT-5.4

#173
post #69

These releases are lacking something. Yes, they optimised for benchmarks, but it’s just not all that impressive anymore. It is time for a product, not for a marginally improved model.

They need something that POPS:

    The new GPT -- SkyNet for _real_

Re: GPT-5.4

#174
post #105
post #64

Earlier quoted context omitted.

Models are being neutered for questions related to law, health etc. for liability reasons.

I'm sometimes surprised how much detail ChatGPT will go into without giving any dislaimers. I very frequently copy/paste the same prompts into Gemini to compare, and Gemini often flat out refuses to engage while ChatGPT will happily make medical recommendations. I also have a feeling it has to do with my account history and heavy use of project context. It feels like when ChatGPT is overloaded with too much context,…

Anecdotal, but I asked Claude the other day about how to dilute my medication (HCG) and it flat out refused and started lecturing me about abusing drugs.

I copy and pasted into ChatGPT, it told me straight away, and then for a laugh said it was actually a magical weight loss drug that I'd bought off the dark web... And it started giving me advice about unregulated weight loss drugs and how to dose them.

Re: GPT-5.4

#176
post #131

Earlier quoted context omitted.

They have AI psychosis and think it's their boyfriend. The 5.x series have terrible writing styles, which is one way to cut down on sycophancy.

Somebody on Twitter used Claude code to connect… toys… as mcps to Claude chat. We’ve seen nothing yet.

ding-dong-cli is needed

Re: GPT-5.4

#177
post #45

The actual card is here https://deploymentsafety.openai.com/gpt-5-4-thinking/introdu... the link currently goes to the announcement.

I must have been sleeping when "sheet" "brief" "primer" etc become known as "cards". I really thought weirdly worded and unnecessary "announcement" linking to the actual info along with the word "card" were the results of vibe slop.

Card is slightly odd naming indeed.

Criticisms aside (sigh), according to Wikipedia, the term was introduced when proposed by mostly Googlers, with the original paper [0] submitted in 2018. To quote,

"""In this paper, we propose a framework that we call model cards, to encourage such transparent model reporting. Model cards are short documents accompanying trained machine learning models that provide benchmarked evaluation in a variety of conditions, such as across different cultural, demographic, or phenotypic groups (e.g., race, geographic location, sex, Fitzpatrick skin type [15]) and intersectional groups (e.g., age and race, or sex and Fitzpatrick skin type) that are relevant to the intended application domains. Model cards also disclose the context in which models are intended to be used, details of the performance evaluation procedures, and other relevant information."""

So that's where they were coming from, I guess.

[0] Margaret Mitchell et al., 2018 submission, Model Cards for Model Reporting, https://arxiv.org/abs/1810.0399

Re: GPT-5.4

#178
post #128

Earlier quoted context omitted.

Benchmarks don't capture a lot - relative response times, vibes, what unmeasured capabilities are jagged and which are smooth, etc. I find there's a lot of difference between models - there are things which Grok is better than ChatGPT for that the benchmarks get inverted, and vice versa. There's also the UI and tools at hand - ChatGPT image gen is just straight up better, but Grok Imagine does better videos, and is f…

Gemini 3.1 slaps all other models at subtle concurrency bugs, sql and js security hardening when reviewing . (Obviously haven’t tested gpt 5.4 yet.) It’s a required step for me at this point to run any and all backend changes through Gemini 3.1 pro.

Which subscription do you have to use it? Via Google ai pro and gemini cli i always get timeouts due to model being under heavy usage. The chat interface is there and I do have 3.1 pro as well, but wondering if the chat is the only way of accessing it.

Re: GPT-5.4

#179

can anyone compare the $200/mo codex usage limits with the $200/mo claude usage limits? It’s extremely difficult to get a feel for whether switching between the two is going to result in hitting limits more or less often, and it’s difficult to find discussion online about this. In practice, if I buy $200/mo codex, can I basically run 3 codex instances simultaneously in tmux, like I can with claude code pro max, all d…

Codex usage limits are definitely more generous. As for their strength, that's hard to say / personal taste

Re: GPT-5.4

#180
post #111
post #69

These releases are lacking something. Yes, they optimised for benchmarks, but it’s just not all that impressive anymore. It is time for a product, not for a marginally improved model.

The model was released less than an hour ago, and somehow you've been able to form such a strong opinion about it. Impressive!

I am actually super impressed with Codex-5.3 extra high reasoning. Its a drop in replacement (infact better than Claude Opus 4.6. lately claude being super verbose going in circles in getting things resolved). I stopped using claude mostly and having a blast with Codex 5.3. looking forward to 5.4 in codex.
Post reply on HN