Considering both this blog post and the livestream demos, I am underwhelmed. Having just finished the stream, I had a real "was that all" moment, which on one hand shows how spoiled I've gotten by new models impressing me, but on another feels like OpenAI really struggles to stay ahead of their competitors. What has been shown feels like it could be achieved using a custom system prompt on older versions of OpenAIs m…
My first thought seeing this and looking at benchmarks was that if it wasn’t for reasoning, then either pundits would be saying we’ve hit a plateau, or at the very least OpenAI is clearly in 2nd place to Anthropic in model performance. Of course we don’t live in such a world, but I thought of this nonetheless because for all the connotations that come with a 4.5 moniker this is kind of underwhelming.
GPT-4.5
161–170 of 1001 posts
Re: GPT-4.5
#162A bit better at coding than ChatGPT 4o but not better than o3-mini - there is a chart near the bottom of the page that is easy to overlook: - ChatGPT 4.5 on AWS Bench verified: 38.0% - ChatGPT 4o on AWS Bench verified: 30.7% - OpenAI o3-mini on AWS Bench verified: 61.0% BTW Anthropic Claude 3.7 is better than o3-mini at coding at around 62-70% [1]. This means that I'll stick with Claude 3.7 for the time being for my…
Re: GPT-4.5
#163Earlier quoted context omitted.
From Sam's twitter: > After that, a top goal for us is to unify o-series models and GPT-series models by creating systems that can use all our tools, know when to think for a long time or not, and generally be useful for a very wide range of tasks. > In both ChatGPT and our API, we will release GPT-5 as a system that integrates a lot of our technology, including o3. We will no longer ship o3 as a standalone model. Yo…
Ah, great point. Yes, the wording here would imply that they're basically planning on building scaffolding around multiple models instead of having one more capable Swiss Army Knife model. I would feel a bit bummed if GPT-5 turned out not to be a model, but rather a "product".
Somehow working together in the same latent space? That could be neat.
Re: GPT-4.5
#164I can't wait for fireship.io and the comment section here to tell me what to think about this
You appear to have the direction of causation reversed. (In that fireship does the same)
Re: GPT-4.5
#165Earlier quoted context omitted.
I suppose this was their final hurrah after two failed attempts at training GPT-5 with the traditional pre-training paradigm. Just confirms reasoning models are the only way forward.
What it confirms, I think, is, that we are going to need a lot more chips.
Re: GPT-4.5
#166I feel like OpenAI is pursuing AGI when Anthropic/Claude is pursuing making AI awesome for practical things like coding. I only ever using OpenAI's coding now as a double check against Claude. Does OpenAI have their eyes on the ball?
>I feel like OpenAI is pursuing AGI I don't think so, the "AGI guy" was Ilya Sutskever, he is gone, he wanted to make OpenAI "less comercial", AGI is just a buzzword for Altmann.
Re: GPT-4.5
#1671. It is very very slow, for some applications where you want real time interactions is just not viable, the text attached below took 7s to generate with 4o, but 46s with GPT4.5
2. The style it writes is way better: it keeps the tone you ask and makes better improvements on the flow. One of my biggest complaints with 4o is that you want for your content to be more casual and accessible but GPT / DeepSeek wants to write like Shakespeare did.
Some comparisons on a book draft: GPT4o (left) and GPT4.5 (green). I also adjusted the spacing around the paragraphs, to better diff match. I still am wary of using ChatGPT to help me write, even with GPT 4.5, but the improvement is very noticeable.
Re: GPT-4.5
#168GPT 4.5 pricing is insane: Price Input: $75.00 / 1M tokens Cached input: $37.50 / 1M tokens Output: $150.00 / 1M tokens GPT 4o pricing for comparison: Price Input: $2.50 / 1M tokens Cached input: $1.25 / 1M tokens Output: $10.00 / 1M tokens It sounds like it's so expensive and the difference in usefulness is so lacking(?) they're not even gonna keep serving it in the API for long: > GPT‑4.5 is a very large and comput…
> We look forward to learning more about its strengths, capabilities, and potential applications in real-world settings. If GPT‑4.5 delivers unique value for your use case, your feedback (opens in a new window) will play an important role in guiding our decision. "We don't really know what this is good for, but spent a lot of money and time making it and are under intense pressure to announce new things right now. If…
Re: GPT-4.5
#169In other words, these performance stats with Gemini 2.0 Flash pricing looks reasonable. At these prices, zero usecases for anyone I think. This is a dead on arrival model.
Re: GPT-4.5
#170Claude 3.6 (new 3.5) and 3.7 non-reasoning are much better at pretty much everything, and much cheaper. What's Anthropic's secret sauce?