Live data from Hacker News

GPT-4.5

openai.com

661–670 of 1001 posts

Re: GPT-4.5

#661
post #9

One comparison I found interesting... I think GPT-4o has a more balanced answer! > What are your thoughts on space exploration? GPT-4.5: Space exploration isn't just valuable—it's essential. People often frame it as a luxury we pursue after solving Earth-bound problems. But space exploration actually helps us address those very challenges: climate change (via satellite monitoring), resource scarcity (through asteroid…

"X isn't just Y - it's Z. [Waffle]. By doing X, you can YY. Remember, ZZ. [Final superfluous sentence]"

God I hate reading what crapgpt writes.

Re: GPT-4.5

#662

Earlier quoted context omitted.

Your original comment opened with: You are lying. This is an ad hominem which assumes intent unknown to anyone other than the person to whom you replied. Subsequently railing against comment rankings and enumerating curt summaries of other comments does not help either.

Lying is defined as "used with reference to a situation involving deception or founded on a mistaken impression." What am I missing here? Those weren't curt summaries, they were quotes! And not pull quotes, they were the unedited beginning of each claim!

>> This is an ad hominem which assumes intent unknown to anyone other than the person to whom you replied.

> What am I missing here?

Intent. Neither you nor I know what the person to whom you replied had.

> Those weren't curt summaries, they were quotes! And not pull quotes, they were the unedited beginning of each claim!

Maybe the more important part of that sentence was:

  Subsequently railing against comment rankings ...
But you do you.

I commented as I did in hope it helped address what I interpreted as confusion regarding how the posts were being received. If it did not help, I apologize.

Re: GPT-4.5

#663

Earlier quoted context omitted.

People being wrong (especially on the internet) doesn't mean they are lying. Lying is being wrong intentionally. Also, the person you replied to comments on the wording tricks they use. After suddenly bringing new data and direction in the discussion, even calling them "wrong" would have been a stretch. I kindly suggest that you (and we all!) to keep discussing with an assumption of good faith.

"Early testing doesn't show that it hallucinates less, but we expect that putting ["we expect it will hallucinate less"] nearby will lead you to draw a connection there yourself"." The link, the link we are discussing shows testing, with numbers. They say "early testing doesn't show that it hallucinates less", to provide a basis for a claim of bad faith. You are claiming that mentioning this is out of bounds if it co…

Oh boy. Do I need to tell you how to communicate?

That comment is making fun of their wording. Maybe extracting too much meaning from their wordplay? Maybe.

Afterwards, evidence is presented that they did not have to do this, which makes that point not so important, and even wrong.

The commenter was not lying, and they were correct about how masterfully deceiving that sequence of sentences are. They arrived at a wrong conclusion though.

Kindly point that out. Say, "hey, the numbers tell a different story, perhaps they didn't mean/need to make a wordplay there".

Re: GPT-4.5

#665
post #24

Considering both this blog post and the livestream demos, I am underwhelmed. Having just finished the stream, I had a real "was that all" moment, which on one hand shows how spoiled I've gotten by new models impressing me, but on another feels like OpenAI really struggles to stay ahead of their competitors. What has been shown feels like it could be achieved using a custom system prompt on older versions of OpenAIs m…

The niche of GPT-4.5 is lower hallucations than any existing model. Whether that niche justifies the price tag for a subset of usecases remains to be seen.

Actually, this comment of mine was incorrect, or at least we don't have enough information to conclude this. The metric OpenAI are reporting is the total number of incorrect responses on SimpleQA (and they're being beaten by Claude Haiku on this metric...), which is a deceptive metric because it doesn't account for non-responses. A better metric would be the ratio of Incorrects to the total number of attempts.

Re: GPT-4.5

#667

Earlier quoted context omitted.

I've worked for very large software companies, some of the biggest products ever made, and never in 25 years can I recall us shipping an update we didn't know was an improvement. The idea that you'd ship something to hundreds of millions of users and say "maybe better, we're not sure, let us know" is outrageous.

Maybe accidental, but I feel you’ve presented a straw man. We’re not discussing something that _may be_ better. It _is_ better. It’s not as big an improvement as previous iterations have been, but it’s still improvement. My claim is that reasonable people might still ship it.

It _is_ better in the general case on most benchmarks. There are also very likely specific use cases for which it is worse and very likely that OpenAI doesn't know what all of those are yet.

Re: GPT-4.5

#668
post #339

Earlier quoted context omitted.

GPT-4.5 may be an awesome model, some say!

Claude just got a version bump from 3.5 to 3.7. Quite a few people have been asking when OpenAI will get a version bump as well, as GPT 4 has been out "what feels like forever" in the words of a specialist I speak with. Releasing GPT 4.5 might simply be a reaction to Claude 3.7.

Feels like when Slackware bumped their Linux version from 4 to 7 just to show they were not falling behind the rest.

Wow, I'm old.

Re: GPT-4.5

#669

In many ways I'm not an OpenAI fan (but I need to recognize their many merits). At the same time, I believe people are missing what they tried to do with GPT 4.5: it was needed and important to explore the pre-training scaling law in that direction. A gift to science, however selfist it could be.

> A gift to science This is hardly recognizable as science. edit: Sorry, didn't feel this was a controversial opinion. What I meant to say was that for so-called science, this is not reproducible in any way whatsoever. Further, this page in particular has all the hallmarks of _marketing_ copy, not science. Sometimes a failure is just a failure, not necessarily a gift. People could tell scaling wasn't working well bef…

People could tell scaling wasn't working well before the release of GPT 4.5

Who could tell? Who has tried scaling up to this level?

Re: GPT-4.5

#670

If this cannot eliminate hallucinations or at least reduce them to be statistically unlikely to be happen, and I assume it has more params than GPT4's trillion parameters, that means the scaling law is dead isn't it?

I interpret this to mean we're in the ugly part of the old scaling law, where `ln(x)` for `x > $BIGNUMBER` starts to becoming punishing, not that the scaling law is in any way empirically refuted. Maybe someone can crunch the numbers and figure out if the benchmarks empirically validate the scaling law or not, relative to GPT-4o (assuming e.g. 200 million params vs 5T params).
Post reply on HN