For day to day coding, I've found Anthropic to be killing it with Sonnet 3.7 and now Sonnet 4, and Claude Code feeling like it has even bigger advantages over when it's used in Cursor (And I can't explain why). I don't even try to use the OpenAI models because it's felt like night and day. Hopefully GPT-5 helps them catch up. Although I'm sure there are 100 people that have their own personal "hopefully GPT-5 fixes m…
GPT-5
81–90 of 1001 posts
Re: GPT-5
#82OpenAI taking a page out of Apple's book and only comparing against themselves
Re: GPT-5
#83It's very interesting how memetic the language around different models is. Elon seems to have coined "PhD level intelligence in all topics" and now Sam repeated it in his presentation. Despite it not having an actual meaning. I think OpenAI will coin they've achieved AGI first (as they have incentives to based on the rumored contract with MSFT), and then everyone will claim we've achieved it.
Re: GPT-5
#84Re: GPT-5
#85> comments turned off yikes - the poor executive leadership’s fragile egos cannot take the criticism.
Re: GPT-5
#86It seems like it's actually an ideal "trick" question for an LLM actually, since so much content has been written about it incorrectly. I thought at first they were going to demo this to show that it knew better, but it seems like it's just regurgitating the same misleading stuff. So, not a good look.
Re: GPT-5
#87Is it bad that I hope it's not a significant improvement in coding?
Re: GPT-5
#88Re: GPT-5
#89Re: GPT-5
#90What's going on with their SWE bench graph?[0] GPT-5 non-thinking is labeled 52.8% accuracy, but o3 is shown as a much shorter bar, yet it's labeled 69.1%. And 4o is an identical bar to o3, but it's labeled 30.8%... [0] https://i.postimg.cc/DzkZZLry/y-axis.png