What's going on with this plot's y-axis? https://bsky.app/profile/tylermw.com/post/3lvtac5hues2n
GPT-5
161–170 of 1001 posts
Re: GPT-5
#162Watching the livestream now, the improvement over their current models on the benchmarks is very small. I know they seemed to be trying to temper our expectations leading up to this, but this is much less improvement than I was expecting
im sure i am repeating someone else but sounds like we're coming over the s-curve
Diminished returns.-
... here's hoping it leads to progress.-
Re: GPT-5
#163Sam Altman, in the summer update video: > "[GPT-5] can write an entire computer program from scratch, to help you with whatever you'd like. And we think this idea of software on demand is going to be one of the defining characteristics of the GPT-5 era."
GPT-5 doesn't seem to get you there tho ...
(Disclaimer: But I am 100% sure it will happen eventually)
Re: GPT-5
#164These presenters all give off such a “sterile” vibe
Re: GPT-5
#165Not that this proves GPT-5 sucks, but it made me laugh that I could cheese the rolling ball minigame by holding spacebar.
Re: GPT-5
#166Is it bad that I hope it's not a significant improvement in coding?
Re: GPT-5
#167Input: $1.25 / 1M tokens Cached: $0.125 / 1M tokens Output: $10 / 1M tokens
With 74.9% on SWE-bench, this inches out Claude Opus 4.1 at 74.5%, but at a much cheaper cost.
For context, Claude Opus 4.1 is $15 / 1M input tokens and $75 / 1M output tokens.
> "GPT-5 will scaffold the app, write files, install dependencies as needed, and show a live preview. This is the go-to solution for developers who want to bootstrap apps or add features quickly." [0]
Since Claude Code launched, OpenAI has been behind. Maybe the RL on tool calling is good enough to be competitive now?
Re: GPT-5
#168Seems LLMs really hit the wall.