Live data from Hacker News

GPT-4.5

openai.com

721–730 of 1001 posts

Re: GPT-4.5

#722

Earlier quoted context omitted.

Now the real question about AI automation starts. Is it cheaper to pay a human to do the task or a AI company?

Humans have all sorts of issues you have to deal with. Being hungover, not sleeping well, having a personality, being late to work, not being able to work 24/7, very limited ability to copy them. If there's a soulless generic office-droidGPT that companies could hire that would never talk back and would do all sorts of menial work without needing breaks or to use the bathroom, I don't know that we humans stand a chan…

Once we get to that stage, unless you're a capitalist, remember that your job is next in line to be replaced.

Re: GPT-4.5

#723

Earlier quoted context omitted.

what crazy progress? how much do you spend on tokens every month to witness the crazy progress that I'm not seeing? I feel like I'm taking crazy pills. The progress is linear at best

Large parts of my coding are now done by Claude/Cursor. I give it high level tasks and it just does it. It is honestly incredible, and if I would have see this 2 years ago I wouldn't have believed it.

What kind of coding do you do? How much of it is formulaic?

Re: GPT-4.5

#724

Earlier quoted context omitted.

> "Early testing shows that interacting with GPT‑4.5 feels more natural. Its broader knowledge base, improved ability to follow user intent, and greater “EQ” make it useful for tasks like improving writing, programming, and solving practical problems. We also expect it to hallucinate less." "Early testing doesn't show that it hallucinates less, but we expect that putting that sentence nearby will lead you to draw a c…

The link has data. The link shows a significant reduction. grep hallucination, or, https://imgur.com/a/mkDxe78 .

I really doubt LLM benchmarks are reflective of real world user experience ever since they claimed GPT-4o hallucinated less than the original GPT-4.

Re: GPT-4.5

#725

My 2 cents (disclaimer: I am talking out of my ass) here is why GPTs actually suck at fluid knowledge retrievel (which is kinda their main usecase, with them being used as knowledge engines) - they've mentioned that if you train it on 'Tom Cruise was born July 3, 1962', it won't be able to answer the question "Who was born on July 3, 1962", if you don't feed it this piece of information. It can't really internally co…

An LLM on its own isn't necessarily great for fluid knowledge retrieval, as in directly from its training data. But they're pretty good when you add RAG to it. For instance, asking Copilot "Who was born on July 3, 1962" gave the response: > One notable person born on July 3, 1962, is Tom Cruise, the famous American actor known for his roles in movies like Risky Business, Jerry Maguire, and Rain Man. > Are you a fan o…

Wow it googled the date!

Re: GPT-4.5

#726

If this cannot eliminate hallucinations or at least reduce them to be statistically unlikely to be happen, and I assume it has more params than GPT4's trillion parameters, that means the scaling law is dead isn't it?

I mean the scaling laws were always logarithms, and logarithms become arbitrarily close to flat if you can't drive them with exponential growth, and even if you do it's barely linear. The scaling laws always predicted that model scaling would stop/slow being practical at some point.

Right but the quantum leap in capabilities that came from GPT2->GPT3->GPT3.5Turbo (which I personally felt didn't fare as well at coding as the former)->GPT4 won't be replicated anytime soon with the pure text/chat generation models.

Re: GPT-4.5

#727
post #24

Considering both this blog post and the livestream demos, I am underwhelmed. Having just finished the stream, I had a real "was that all" moment, which on one hand shows how spoiled I've gotten by new models impressing me, but on another feels like OpenAI really struggles to stay ahead of their competitors. What has been shown feels like it could be achieved using a custom system prompt on older versions of OpenAIs m…

I suspect they may launch a GPT4.5Turbo with a price cut... GPT4/GPT432k etc were all pricier than the GPT4Turbo models which also came with the added context length.. but with this huge jump in price, even 4.5Turbo if it does come out would be pricier

Re: GPT-4.5

#728

Earlier quoted context omitted.

I think it's fairer to compare it to the original GPT-4 which might the equivalent in term of "size" (though we don't have actual numbers for either). GPT-4: Input $30.00 / 1M tokens ; Output $60.00 / 1M tokens So 4.5 is 2.5x more expensive. I think they announced this as their last non-reasoning model, so it was maybe with the goal of stretching pre-training as far as they could, just to see what new capabilities wo…

2x that price for the 32k context via API at launch. So nearly the same price, but you get 4x the context

Honestly if long context (that doesn't start to degrade quickly) is what you're after, I would use Grok 3 (not sure when the api version releases though). Over the last week or so I've had a massive thread of conversation with it that started with plenty of my project's relevant code (as in couple hundred lines), and several days later, after like 20 question-aswer blocks, you ask it something and it aswers "since you're doing that this way, and you said you want x, y and z, here are your options blabla"... It's like thinking Gemini but better. Also, unlike Gemini (and others) it seems to have a much more recent data cutoff. Try asking about some language feature / library / framework that has been released recently (say 3 months ago) and most of the models shit the bed, use older versions of the thing or just start to imitate what the code might look like. For example try asking Gemini if it can generate Tailwind 4 code, it will tell you that it's training cutoff is like October or something and Tailwind 4 "isn't released yet" and that it can try to imitate what the code might look like. Uhhhhhh, thanks I guess??

Re: GPT-4.5

#729
post #167

First impression of GPT-4.5: 1. It is very very slow, for some applications where you want real time interactions is just not viable, the text attached below took 7s to generate with 4o, but 46s with GPT4.5 2. The style it writes is way better: it keeps the tone you ask and makes better improvements on the flow. One of my biggest complaints with 4o is that you want for your content to be more casual and accessible bu…

In my experience, Gemini Flash has been the best at writing, and GPT 3.5 onwards has been terrible.

GPT-3 and GPT-2 were actually remarkably good at it, arguably better than a skilled human. I had a bit of fun ghostwriting with these and got a little fan base for a while.

It seems that GPT-4.5 is better than 4 but it's nowhere near the quality of GPT-3 davinci. Davinci-002 has been nerfed quite a bit, but in the end it's $2/MTok for higher quality output.

It's clear this is something users want, but OpenAI and Anthropic seem to be going in the opposite direction.

Re: GPT-4.5

#730

Earlier quoted context omitted.

> A gift to science This is hardly recognizable as science. edit: Sorry, didn't feel this was a controversial opinion. What I meant to say was that for so-called science, this is not reproducible in any way whatsoever. Further, this page in particular has all the hallmarks of _marketing_ copy, not science. Sometimes a failure is just a failure, not necessarily a gift. People could tell scaling wasn't working well bef…

People could tell scaling wasn't working well before the release of GPT 4.5 Who could tell? Who has tried scaling up to this level?

OpenAI took a bullet for the team, by perhaps scaling the model to something bigger than the 1.6T params GPT4 possibly had and basically telling its competitors its not gonna be worth scaling much beyond those number of params in GPT4, without a change in the model architecture
Post reply on HN