Live data from Hacker News

GPT-5.4

openai.com

151–160 of 868 posts

Re: GPT-5.4

#151
83% win rate over industry professionals across 44 occupations.

I'd believe it on those specific tasks. Near-universal adoption in software still hasn't moved DORA metrics. The model gets better every release. The output doesn't follow. Just had a closer look on those productivity metrics this week: https://philippdubach.com/posts/93-of-developers-use-ai-codi...

Re: GPT-5.4

#154
post #5

"GPT‑5.4 interprets screenshots of a browser interface and interacts with UI elements through coordinate-based clicking to send emails and schedule a calendar event." They show an example of 5.4 clicking around in Gmail to send an email. I still think this is the wrong interface to be interacting with the internet. Why not use Gmail APIs? No need to do any screenshot interpretation or coordinate-based clicking.

A model that gets good at computer use can be plugged in anywhere you have a human. A model that gets good at API use cannot. From the standpoint of diffusion into the economy/labor market, computer use is much higher value.

Re: GPT-5.4

#155
post #69

These releases are lacking something. Yes, they optimised for benchmarks, but it’s just not all that impressive anymore. It is time for a product, not for a marginally improved model.

The product is putting the skills / harness behind the api instead of the agent locally on your computer and iterating on that between model updates. Close off the garden.

Not that I want it, just where I imagine it going.

Re: GPT-5.4

#156
post #96

Earlier quoted context omitted.

[flagged]

Can't you continue to use to older model, if you prefer the pricing? But they also claim this new model uses fewer tokens, so it still might ultimately be cheaper even if per token cost is higher.

I'm not against the pricing, just seems uncommon to frame it in the way they did, as opposed to the usual 'assume the customer expects more performance will cost more'

I guess they have to sell to investors that the price to operate is going down, while still needing more from the user to be sustainable

Re: GPT-5.4

#157

Earlier quoted context omitted.

Benchmarks don't capture a lot - relative response times, vibes, what unmeasured capabilities are jagged and which are smooth, etc. I find there's a lot of difference between models - there are things which Grok is better than ChatGPT for that the benchmarks get inverted, and vice versa. There's also the UI and tools at hand - ChatGPT image gen is just straight up better, but Grok Imagine does better videos, and is f…

> If this rate of progress is steady, though, this year is gonna be crazy. Do you want to make any concrete predictions of what we'll see at this pace? It feels like we're reaching the end of the S-curve, at least to me.

If you look at the difference in quality between gpt-2 and 3, it feels like a big step, but the difference between 5.2 and 5.4 is more massive, it's just that they're both similarly capable and competent. I don't think it's an S curve; we're not plateauing. Million token context windows and cached prompts are a huge space for hacking on model behaviors and customization, without finetuning. Research is proceeding at light speed, and we might see the first continual/online learning models in the near future. That could definitively push models past the point of human level generality, but at the very least will help us discover what the next missing piece is for AGI.

Re: GPT-5.4

#158

Earlier quoted context omitted.

[flagged]

Maybe it's finally a bigger pretrain?

I feel like that would have been highlighted then. "As this is a bigger pretrain, we have to raise prices".

They're framing it pretty directly "We want you to think bigger cost means better model"

Re: GPT-5.4

#160

What is the point of gpt codex?

-codex variant models in earlier version were just fine tuned for coding work, and had a little better performance for related tool calling and maybe instruction calling. in 5.4 it looks like the just collapsed that capability into the single frontier family model

They’ll likely come out with a 5.4-Codex at some point, that’s what they did with 5 and 5.2
Post reply on HN