Live data from Hacker News

GPT-4.5

openai.com

81–90 of 1001 posts

Re: GPT-4.5

#81
post #24

Considering both this blog post and the livestream demos, I am underwhelmed. Having just finished the stream, I had a real "was that all" moment, which on one hand shows how spoiled I've gotten by new models impressing me, but on another feels like OpenAI really struggles to stay ahead of their competitors. What has been shown feels like it could be achieved using a custom system prompt on older versions of OpenAIs m…

> How could they justify that asking price?

They're still selling $1 for <$1. Like personal food delivery before it, consumers will eventually need to wake up to this fact - these things will get expensive, fast.

Re: GPT-4.5

#82
I'm one week in on heavy grok usage. I didn't think I'd say this, but for personal use, I'm considering cancelling my OpenAI plan.

The one thing I wish grok had was more separation of the UI from X itself. The interface being so coupled to X puts me off and makes it feel like a second-hand citizen. I like ChatGPTs minimalist UI.

Re: GPT-4.5

#84
post #26

GPT 4.5 pricing is insane: Price Input: $75.00 / 1M tokens Cached input: $37.50 / 1M tokens Output: $150.00 / 1M tokens GPT 4o pricing for comparison: Price Input: $2.50 / 1M tokens Cached input: $1.25 / 1M tokens Output: $10.00 / 1M tokens It sounds like it's so expensive and the difference in usefulness is so lacking(?) they're not even gonna keep serving it in the API for long: > GPT‑4.5 is a very large and comput…

If it really costs them 30x more surely they must plan on putting pretty significant usage limits on any rollout to the Plus tier and if that is the case i'm not sure what the point is considering it seems primarily a replacement/upgrade for 4o.

The cognitive overhead of choosing between what will be 6 different models now on chatGPT and trying to map whether a query is "worth" using a certain model and worrying about hitting usage limits is getting kind of out of control.

Re: GPT-4.5

#85
post #7

A bit better at coding than ChatGPT 4o but not better than o3-mini - there is a chart near the bottom of the page that is easy to overlook: - ChatGPT 4.5 on AWS Bench verified: 38.0% - ChatGPT 4o on AWS Bench verified: 30.7% - OpenAI o3-mini on AWS Bench verified: 61.0% BTW Anthropic Claude 3.7 is better than o3-mini at coding at around 62-70% [1]. This means that I'll stick with Claude 3.7 for the time being for my…

>BTW Anthropic Claude 3.7 is better than o3-mini at coding at around 62-70% [1]. This means that I'll stick with Claude 3.7 for the time being for my open source alternative to Claude-code That's not a fair comparison as o3-mini is significantly cheaper. It's fine if your employer is paying, but on a personal project the cost of using Claude through the API is really noticeable.

> That's not a fair comparison as o3-mini is significantly cheaper. It's fine if your employer is paying...

I use it via Cursor editor's built-in support for Claude 3.7. That caps the monthly expense to $20. There probably is a limit in Claude for these queries. But I haven't run into it yet. And I am a heavy user.

Re: GPT-4.5

#86

Earlier quoted context omitted.

I think it's fairer to compare it to the original GPT-4 which might the equivalent in term of "size" (though we don't have actual numbers for either). GPT-4: Input $30.00 / 1M tokens ; Output $60.00 / 1M tokens So 4.5 is 2.5x more expensive. I think they announced this as their last non-reasoning model, so it was maybe with the goal of stretching pre-training as far as they could, just to see what new capabilities wo…

The marginal costs for running a GPT-4-class LLM are much lower nowadays due to significant software and hardware innovations since then, so costs/pricing are harder to compare.

Agreed, however it might make sense that a much-larger-than-GPT-4 LLM would also, at launch, be more expensive to run than the OG GPT-4 was at launch.

(And I think this is probably also scarecrow pricing to discourage casual users from clogging the API since they seem to be too compute-constrained to deliver this at scale)

Re: GPT-4.5

#87
post #16

Sounds like it's a distill of O1? After R1, I don't care that much about non-reasoning models anymore. They don't even seem excited about it on the livestream. I want tiny, fast and cheap non-reasoning models I can use in APIs and I want ultra smart reasoning models that I can query a few times a day as an end user (I don't mind if it takes a few minutes while I refill a coffee). Oh, and I want that advanced voice mo…

My impression is that it’s a massive increase in the parameter count. This is likely the spiritual successor to GPT4 and would have been called GPT5 if not for the lackluster performance. The speculation is that there simply isn’t enough data on the internet to support yet another 10x jump in parameters.

O1-mini is a distill of O1. This definitely isn’t the same thing.

Re: GPT-4.5

#88
post #26

GPT 4.5 pricing is insane: Price Input: $75.00 / 1M tokens Cached input: $37.50 / 1M tokens Output: $150.00 / 1M tokens GPT 4o pricing for comparison: Price Input: $2.50 / 1M tokens Cached input: $1.25 / 1M tokens Output: $10.00 / 1M tokens It sounds like it's so expensive and the difference in usefulness is so lacking(?) they're not even gonna keep serving it in the API for long: > GPT‑4.5 is a very large and comput…

I suppose this was their final hurrah after two failed attempts at training GPT-5 with the traditional pre-training paradigm. Just confirms reasoning models are the only way forward.

> Compared to OpenAI o1 and OpenAI o3‑mini, GPT‑4.5 is a more general-purpose, innately smarter model. We believe reasoning will be a core capability of future models, and that the two approaches to scaling—pre-training and reasoning—will complement each other. As models like GPT‑4.5 become smarter and more knowledgeable through pre-training, they will serve as an even stronger foundation for reasoning and tool-using agents.

Re: GPT-4.5

#89
post #52
post #7

A bit better at coding than ChatGPT 4o but not better than o3-mini - there is a chart near the bottom of the page that is easy to overlook: - ChatGPT 4.5 on AWS Bench verified: 38.0% - ChatGPT 4o on AWS Bench verified: 30.7% - OpenAI o3-mini on AWS Bench verified: 61.0% BTW Anthropic Claude 3.7 is better than o3-mini at coding at around 62-70% [1]. This means that I'll stick with Claude 3.7 for the time being for my…

It's the other way around on their new SWE-Lancer benchmark, which is pretty interesting: GPT-4.5 scores 32.6%, while o3-mini scores 10.8%.

To put that in context, Claude 3.5 Sonnet (new), a model we have had for months now and which from all accounts seems to have been cheaper to train and is cheaper to use, is still ahead of GPT-4.5 at 36.1% vs 32.6% in SWE-Lancer Diamond [0]. The more I look into this release, the more confused I get.

[0] https://arxiv.org/pdf/2502.12115

Re: GPT-4.5

#90
post #16

Sounds like it's a distill of O1? After R1, I don't care that much about non-reasoning models anymore. They don't even seem excited about it on the livestream. I want tiny, fast and cheap non-reasoning models I can use in APIs and I want ultra smart reasoning models that I can query a few times a day as an end user (I don't mind if it takes a few minutes while I refill a coffee). Oh, and I want that advanced voice mo…

It isn't even vaguely a distill of o1. The reasoning models are, from what we can tell, relatively small. This model is massive and they probably scaled the parameter count to improve factual knowledge retention. They also mentioned developing some new techniques for training small models and then incorporating those into the larger model (probably to help scale across datacenters), so I wonder if they are doing a bi…

You can 'distill' with data from a smaller, better model into a larger, shittier one. It doesn't matter. This is what they said they did on the livestream.
Post reply on HN