A bit better at coding than ChatGPT 4o but not better than o3-mini - there is a chart near the bottom of the page that is easy to overlook: - ChatGPT 4.5 on AWS Bench verified: 38.0% - ChatGPT 4o on AWS Bench verified: 30.7% - OpenAI o3-mini on AWS Bench verified: 61.0% BTW Anthropic Claude 3.7 is better than o3-mini at coding at around 62-70% [1]. This means that I'll stick with Claude 3.7 for the time being for my…
Does the benchmark reflect your opinion on 3.7? I've been using 3.7 via Cursor and it's noticeably worse than 3.5. I've heard using the standalone model works fine, didn't get a chance to try it yet though.
GPT-4.5
181–190 of 1001 posts
Re: GPT-4.5
#182> I’m going through a tough time after failing a test > > Aw, I'm really sorry to hear that. Failing a test can feel pretty tough and discouraging, but remember, it's just one moment—not a reflection of who you are or your capabilities. Want to talk about what happened, or do you just need a distraction? I'm here either way.
vs 4o:
>I’m going through a tough time after failing a test > >I'm really sorry to hear that you're going through a tough time. Here are a few >things you might consider: > >1. Reflect and Understand: Try to analyze what might have gone wrong. Was it a >lack of preparation, time management during the test, or perhaps >misunderstanding the material? Understanding the root cause can help you address >it in the future. > >2. Seek Support: Talk to someone you trust—friends, family, or a school >counselor. They can offer a different perspective and sometimes just talking >about your feelings can be alleviating. > >3. Speak with Your Teacher or Professor: They can provide insights into where >you might have gone wrong and offer advice on how to improve. They might also >offer opportunities for make-up work or extra credit. > >4. Develop a Study Plan: Based on your reflections, create a plan to improve >your study habits. This might include setting regular study times, breaking down >material into manageable portions, or finding new study methods that work better >for you. > >5. Stay Positive and Forgive Yourself: Mistakes and setbacks are a part of life. >It’s important to be kind to yourself and recognize that failure is a stepping >stone to success. > >6. Focus on the Bigger Picture: Remember that one test is just one part of your >educational journey. There will be many more opportunities to do well. > >If you need further support or resources, consider reaching out to educational >support services at your institution, or mental health resources if you’re >feeling particularly overwhelmed. You’re not alone in this, and things can get >better with time and effort.
Is it just me or is the 4o response insanely better? I'm not the type of person to reach for a LLM for help about this kind of thing, but if I were, the 4o respond seems vastly better to the point I'm surprised they used that as their main "EQ" example.
Re: GPT-4.5
#183Probably what happened, is that in doing so, they had to scale either the model size or the training cost to untenable levels.
In my experience, LLMs really suck at fluid knowledge retrieval tasks, like book recommendation - I asked GPT4 to recommend me some SF novels with certain characteristics, and what it spat out was a mix of stuff that didn't really match, and stuff that was really reaching - when I asked the same question on Reddit, all the answers were relevant and on point - so I guess there's still something humans are good for.
Which is a shame, because I'm pretty sure relevant product recommendation is a many billion dollar business - after all that's what Google has built it's empire on.
Re: GPT-4.5
#184GPT 4.5 pricing is insane: Price Input: $75.00 / 1M tokens Cached input: $37.50 / 1M tokens Output: $150.00 / 1M tokens GPT 4o pricing for comparison: Price Input: $2.50 / 1M tokens Cached input: $1.25 / 1M tokens Output: $10.00 / 1M tokens It sounds like it's so expensive and the difference in usefulness is so lacking(?) they're not even gonna keep serving it in the API for long: > GPT‑4.5 is a very large and comput…
> We look forward to learning more about its strengths, capabilities, and potential applications in real-world settings. If GPT‑4.5 delivers unique value for your use case, your feedback (opens in a new window) will play an important role in guiding our decision. "We don't really know what this is good for, but spent a lot of money and time making it and are under intense pressure to announce new things right now. If…
And maybe Tesla is going to deliver truly full self driving tech any day now.
And Star Citizen will prove to have been worth it along along, and Bitcoin will rain from the heavens.
It's very difficult to remain charitable when people seem to always be chasing the new iteration of the same old thing, and we're expected to come along for the ride.
Re: GPT-4.5
#185Claude 3.6 (new 3.5) and 3.7 non-reasoning are much better at pretty much everything, and much cheaper. What's Anthropic's secret sauce?
Post benchmark links.
Re: GPT-4.5
#186This is probably a dumb question, but are we just gonna be stuck on always having X.5 versions of GPT forever? If there's never an X.0, it feels like it's basically meaningless.
Re: GPT-4.5
#187I’m wondering whether this seemingly underwhelming bump on 4o magnifies when/if reasoning is added.
Re: GPT-4.5
#188Per Altman on X: "we will add tens of thousands of GPUs next week and roll it out to the plus tier then". Meanwhile a month after launch rtx 5000 series is completely unavailable and hardly any restocks and the "launch" consisted of microcenters getting literally tens of cards. Nvidia really has basically abandoned consumers.
AI GPUs are bottlenecked mostly by high-bandwidth memory (HBM) chips and CoWoS (packaging tech used to integrate HBM with the GPU die), which are in short supply and aren't found in consumer cards at all
Re: GPT-4.5
#189Claude 3.6 (new 3.5) and 3.7 non-reasoning are much better at pretty much everything, and much cheaper. What's Anthropic's secret sauce?
Re: GPT-4.5
#190Earlier quoted context omitted.
> We look forward to learning more about its strengths, capabilities, and potential applications in real-world settings. If GPT‑4.5 delivers unique value for your use case, your feedback (opens in a new window) will play an important role in guiding our decision. "We don't really know what this is good for, but spent a lot of money and time making it and are under intense pressure to announce new things right now. If…
Said the quiet part out loud! Or as we say these days, “transparently exposed the chain of thought tokens”.