Live data from Hacker News

GPT-4.5

openai.com

851–860 of 1001 posts

Re: GPT-4.5

#851

Earlier quoted context omitted.

I really doubt LLM benchmarks are reflective of real world user experience ever since they claimed GPT-4o hallucinated less than the original GPT-4.

I begin to believe LLM benchmarks are like european car mileage specs. They say its 4 Liter / 100km but everyone knows it's at least 30% off (same with WLTP for EVs).

Those numbers are not off. They are tested on tracks.

You need to remove your shoe and drive with like two toes to get the speed just right, though.

Test drivers I have done this with takes off their shoes or use ballerina shoes.

Re: GPT-4.5

#852
post #846

I've been using 4.5 for the better part of the day. I also have access to o3-mini-high and o1-pro. I don't get it. For general purposes and for writing, 4.5 is no better than o3-mini. It may even be worse. I'd go so far as to say that Deepseek is actually better than 4.5 for most general purpose use cases. I seriously don't understand what they're trying to achieve with this release.

this model does have a niche use-case: since its so large it does have a lot more knowledge and hallucinates much less. for example as a test question I asked it to list the best restaurants in my small town. and all of them existed. none of the other llms get this right.

That's also a use case where the consensus among those in the know is that you shouldn't be relying on the model's size in the first place.

You know what gets the list of restaurants in my home town right? Llama 3.2 1b q4 running on my desktop with web search enabled.

Re: GPT-4.5

#853
post #401
post #144

Earlier quoted context omitted.

The quotation marks in the grandparent comment are scare (sneer) quotes and not actual quotation. https://en.m.wikipedia.org/wiki/Scare_quotes > Whether quotation marks are considered scare quotes depends on context because scare quotes are not visually different from actual quotations.

That's not a scare quote. It's just a proposed subtext of the quote. Sarcastic, sure, but no a scare quote, which is a specific kind of thing. (from your linked wikipedia: "... around a word or phrase to signal that they are using it in an ironic, referential, or otherwise non-standard sense.")

Right. I don't agree with the quote, but it's more like a subtext thing and it seemed to me to be pretty clear from context.

Though, as someone who had a flagged comment a couple years ago for a supposed "misquote" I did in a similar form in style, I think hn's comprehension of this form of communication is not super strong. Also the style more often than not tends towards low quality smarm and probably should be resorted to sparingly.

Re: GPT-4.5

#854
Interesting times that are changing quickly. It looks like the high end pay model that OpenAI is implementing may not be sustainable. Too many new players are making LLM breakthroughs and OpenAI's lead is shrinking and it may be overvalued.

Re: GPT-4.5

#855
A 30x price increase with zero named benefits?

This sure looks like the runway is about to come far short of takeoff. I’m reminded of Ed Zitron’s recent predictions…

Re: GPT-4.5

#856
post #817

Earlier quoted context omitted.

That's some top-tier sales work right there. I suck at and hate writing the mildly deceptive corporate puffery that seems to be in vogue. I wonder if GPT-4.5 can write that for me or if it's still not as good at it as the expert they paid to put that little gem together.

Good sales lines are like prompt injection for the human mind.

Gold

Re: GPT-4.5

#857

Is it official then? Most of us have been waiting for this moment for a while. The transformer architecture as it is currently understood can't be milked any further. Many of us knew this since last year. GPT-5 delays eventually led to non-tech voices to suggest likewise. But we all held our final decision until the next big release from OpenAI as Sam Altman has been making claims about AGI entering the workforce thi…

As someone who is terrified of agentic ASI, I desperately hope this is true. We need more time to figure out alignment.

I'm not sure this will ever be solved. It requires both a technical solution and social consensus. I don't see consensus on "alignment" happening any time soon. I think it'll boil down to "aligned with the goals of the nation-state", and lots of nation states have incompatible goals.

Re: GPT-4.5

#858

Is it official then? Most of us have been waiting for this moment for a while. The transformer architecture as it is currently understood can't be milked any further. Many of us knew this since last year. GPT-5 delays eventually led to non-tech voices to suggest likewise. But we all held our final decision until the next big release from OpenAI as Sam Altman has been making claims about AGI entering the workforce thi…

Honestly, I'm not sure how you can make all those claims when:

1. OpenAI still has the most capable model in o3

2. We've seen some huge increases in capability in 2024, some shocking

3. We're only 3 months into 2025

4. Blackwell hasn't been used to train a model yet

Re: GPT-4.5

#859
post #177

Earlier quoted context omitted.

You're the one buying him the underwear. Don't index funds outperform managed investing? I think especially after accounting for fees, but possibly even after accounting that 50% of money managers are below average.

Depends who's pitch deck you're reading. Warren Buffett didn't get rich waiting on index funds.

I think Warren Buffet doesn't just buy stocks. He also influences the direction of the companies he buys.

Re: GPT-4.5

#860

Is it official then? Most of us have been waiting for this moment for a while. The transformer architecture as it is currently understood can't be milked any further. Many of us knew this since last year. GPT-5 delays eventually led to non-tech voices to suggest likewise. But we all held our final decision until the next big release from OpenAI as Sam Altman has been making claims about AGI entering the workforce thi…

It's worth pointing out that GPT-4.5 seems focused on better pre-training and doesn't include reasoning.

I think GPT-5 - if/when it happens - will be 4.5 with reasoning, and as such it will feel very different.

The barrier, is the computational cost of it. Once 4.5 gets down to similar costs to 4.0 - which could be achieved through various optimization steps (what happened to the ternary stuff that was published last year that meant you could go many times faster without expensive GPUs?), and better/cheaper/more efficient hardware, you can throw reasoning into the mix and suddenly have a major step up in capability.

I am a user, not a researcher of builder. I do think we're in a hype bubble, I do think that LLMs are not The Answer, but I also think there is more mileage left in this path than you seem to. I think automated RL (not HF), reasoning, and better/optimal architectures and hardware mean there is a lot more we can get out of the stochastic parrots, yet.

Post reply on HN