One comparison I found interesting... I think GPT-4o has a more balanced answer! > What are your thoughts on space exploration? GPT-4.5: Space exploration isn't just valuable—it's essential. People often frame it as a luxury we pursue after solving Earth-bound problems. But space exploration actually helps us address those very challenges: climate change (via satellite monitoring), resource scarcity (through asteroid…
GPT-4.5
31–40 of 1001 posts
Re: GPT-4.5
#32This seems very rushed because of DeepSeek's R1 and Anthropic's Claude 3.7 Sonnet. Pretty underwhelming, they didn't even show programming? In the livestream, they struggled to come up with reasons why I should prefer GPT-4.5 over GPT-4o or o1.
Re: GPT-4.5
#33No wonder Sam wasn’t part of the presentation.
Re: GPT-4.5
#34I am beginning to think these human eval tests are a waste of time at best, and negative value at worst. Maybe I am being snobby, but I don't think the average human is able to properly evaluate usefulness, truthfulness, or other metrics that I actually care about. I am sure this is good for openAI since if more people like what the hear, they are more likely come back. I don't want my AI more obsequious, I want it m…
> I want it more correct and capable. How is it supposed to be more correct and capable if these human eval tests are a waste of time? Once you ask it to do more than add two numbers together, it gets a lot more difficult and subjective to determine whether it's correct and how correct.
I've read reports that some of the changes that are preferred by human evaluators actually hurt the performance on the more objective tests.
Re: GPT-4.5
#35Re: GPT-4.5
#36Re: GPT-4.5
#37We're far from the days of "this is not a person, we do not want to make it addictive" and getting a firm foot on the territory of "here's your new AI friend".
This is very visible in the example comparing 4o with 4.5 when the user is complaining about failing a test, where 4o's response is what one would expect from a "typical AI response" with problem-solving bullets, and 4.5 is sending what you'd expect from a pal over instant messaging.
It seems Anthropic and Grok have both been moving in this direction as well. Are we going to see an escalation of foundation models impersonating "a friendly person" rather than "a helpful assistant"?
Personally I find this worrying and (as someone who builds upon SOTA model APIs) I really hope this behavior is not going to seep into API responses, or will at least be steerable through the system/developer prompt.
Re: GPT-4.5
#38Per Altman on X: "we will add tens of thousands of GPUs next week and roll it out to the plus tier then". Meanwhile a month after launch rtx 5000 series is completely unavailable and hardly any restocks and the "launch" consisted of microcenters getting literally tens of cards. Nvidia really has basically abandoned consumers.
Re: GPT-4.5
#39Re: GPT-4.5
#40At this point I think the ultimate benchmark for any new LLM is whether or not it can come up with a coherent naming scheme for itself. Call it “self awareness.”
The people naming them really took the "just give the variable any old name, it doesn't matter" advice from Programming 101 to heart.