Live data from Hacker News

GPT-5.4

openai.com

711–720 of 868 posts

Re: GPT-5.4

#711

Earlier quoted context omitted.

True. Everytime when i ask something gpt, it use to spit out long stories. Claude ans gemini are always straight to point.

I bullied it into giving me concise answers, now it starts every answer with "just quickly" or something similar but it gets straight to the point

I always add no nonsense no bullshit at the end of my prompt. Its annoying how itries to please the user.

Re: GPT-5.4

#713
post #581

Earlier quoted context omitted.

yeah claude is great... but only if you pay $100-$200 a month

Many people buy two separate Claude pro subscriptions and that makes the limit become a non-issue. It works surprisingly well when you tend to hit the 5 hourly limit after a few hours, and hit the weekly limit after 4-5 days. $40 vs $100 is significant for a lot of people.

I hit limit of Pro in about 30 minutes, 1 hour max. And only when I use a single session, and when I don't use it extensively, ie waits for my responses, and I read and really understand what it wants, what it does. That's still just 1-2 hours/5 hours.

What do you do to avoid that?

Re: GPT-5.4

#714

Earlier quoted context omitted.

Preview Road (only choice, and last preview was deprecated without warning)

If the last preview was 'deprecated', it's still usable. So you have two choices. Peeve of mine when people say 'deprecated' but really they mean 'discontinued' or 'deleted'. Things don't instantly disappear when they're deprecated.

Take it up with the organizations that use deprecated and break things immediately

Re: GPT-5.4

#715
post #679

Earlier quoted context omitted.

This is just the stochastic nature of LLM's at play. I think all of the SOTA models are roughly equivalent, but without enough samples people end up reading into it too much.

There's a certain amount of variance in the way that people utilize these agents. Put five people in a room and ask them to compose the same prompt and you have five distinct prompts. Couple this with the fact that models respond better/worse to certain prompts depending on the stylistic composition of the prompt itself. And since people tend to write in the same style, you'd get people who have more luck with one mo…

> Couple this with the fact that models respond better/worse to certain prompts depending on the stylistic composition of the prompt itself.

Do we really know this, or is it just gut feeling? Did somebody really proved this statistically with a great certainty?

Re: GPT-5.4

#716

So let me get this straight, OpenAi previously had an issue with LOTS of different models snd versions being available. Then they solved this by introducing GPT-5 which was more like a router that put all these models under the hood so you only had to prompt to GPT-5, and it would route to the best suitable model. This worked great I assume and made the ui for the user comprehensible. But now, they are starting to in…

You can't keep asking for 100b every 6 months if you don't give the impression of progress

Re: GPT-5.4

#717

Earlier quoted context omitted.

Oh, come on, if it can't run local models that compete with proprietary ones it's not good enough yet!

Qwen 3.5 small models are actually very impressive and do beat out larger proprietary models.

Qwen version 3.5 might be the last serious version (for some time at least), see Something is afoot in the land of Qwen (2 days ago) https://news.ycombinator.com/item?id=47249343

Also interesting experiences shared in that thread, even someone using it on a rented H200.

Re: GPT-5.4

#718

Earlier quoted context omitted.

The real problem that OpenAI had was that their model naming was completely incomprehensible. 4.5, o3, 4o, 4.1 which is newer than 4.5. It was a complete clusterfuck. The blowback on that issue seems to have led them to misidentify the issue, but nobody was really asking for a single router model. Having a number of sequentially numbered and clearly labelled models is not actually a problem.

Having both o4 and 4o. Really. What the fuck?

There was no o4.

Re: GPT-5.4

#719
The style of the output is a marked qualitative improvement. More concise, less dot points, less bolding/italics, less cringe. Well done on that front.

Re: GPT-5.4

#720
post #713

Earlier quoted context omitted.

Many people buy two separate Claude pro subscriptions and that makes the limit become a non-issue. It works surprisingly well when you tend to hit the 5 hourly limit after a few hours, and hit the weekly limit after 4-5 days. $40 vs $100 is significant for a lot of people.

I hit limit of Pro in about 30 minutes, 1 hour max. And only when I use a single session, and when I don't use it extensively, ie waits for my responses, and I read and really understand what it wants, what it does. That's still just 1-2 hours/5 hours. What do you do to avoid that?

You're probably having long sessions, i.e. repeated back-and-forth in one conversation. Also check if you pollute context with unneeded info. It can be a problem with large and/or not well structured codebases.
Post reply on HN