GPT-4.5
21–30 of 1001 posts
Re: GPT-4.5
#22Re: GPT-4.5
#23I am beginning to think these human eval tests are a waste of time at best, and negative value at worst. Maybe I am being snobby, but I don't think the average human is able to properly evaluate usefulness, truthfulness, or other metrics that I actually care about. I am sure this is good for openAI since if more people like what the hear, they are more likely come back. I don't want my AI more obsequious, I want it m…
How is it supposed to be more correct and capable if these human eval tests are a waste of time?
Once you ask it to do more than add two numbers together, it gets a lot more difficult and subjective to determine whether it's correct and how correct.
Re: GPT-4.5
#24What has been shown feels like it could be achieved using a custom system prompt on older versions of OpenAIs models, and I struggle to see anything here that truly required ground-up training on such a massive scale. Hearing that they were forced to spread their training across multiple data centers simultaneously, coupled with their recent release of SWE-Lancer [0] which showed Anthropic (Claude 3.5 Sonnet (new) to be exact) handily beating them, I was really expecting something more than "slightly more casual/shorter output", which again, I fail to see how that wasn't possible by prompting GPT-4o.
Looking at pricing [1], I am frankly astonished.
> Input: $75.00 / 1M tokens > Cached input: $37.50 / 1M tokens > Output: $150.00 / 1M tokens
How could they justify that asking price? And, if they have some amazing capabilities that make a 30-fold pricing increase justifiable, why not show it? Like, OpenAI are many things, but I always felt they understood price vs performance incredibly well, from the start with gpt-3.5-turbo up to now with o3-mini, so this really baffles me. If GPT-4.5 can justify such immense cost in certain tasks, why hide that and if not, why release this at all?
Re: GPT-4.5
#25This seems very rushed because of DeepSeek's R1 and Anthropic's Claude 3.7 Sonnet. Pretty underwhelming, they didn't even show programming? In the livestream, they struggled to come up with reasons why I should prefer GPT-4.5 over GPT-4o or o1.
Re: GPT-4.5
#26GPT 4o pricing for comparison: Price Input: $2.50 / 1M tokens Cached input: $1.25 / 1M tokens Output: $10.00 / 1M tokens
It sounds like it's so expensive and the difference in usefulness is so lacking(?) they're not even gonna keep serving it in the API for long:
> GPT‑4.5 is a very large and compute-intensive model, making it more expensive than and not a replacement for GPT‑4o. Because of this, we’re evaluating whether to continue serving it in the API long-term as we balance supporting current capabilities with building future models. We look forward to learning more about its strengths, capabilities, and potential applications in real-world settings. If GPT‑4.5 delivers unique value for your use case, your feedback (opens in a new window) will play an important role in guiding our decision.
I'm still gonna give it a go, though.
Re: GPT-4.5
#27Not seeing it available in the app or on ChatGPT.com with a pro subscription.
Re: GPT-4.5
#28Per Altman on X: "we will add tens of thousands of GPUs next week and roll it out to the plus tier then". Meanwhile a month after launch rtx 5000 series is completely unavailable and hardly any restocks and the "launch" consisted of microcenters getting literally tens of cards. Nvidia really has basically abandoned consumers.