Pricing seems good, but the open question is still on tool calling reliability. Input: $1.25 / 1M tokens Cached: $0.125 / 1M tokens Output: $10 / 1M tokens With 74.9% on SWE-bench, this inches out Claude Opus 4.1 at 74.5%, but at a much cheaper cost. For context, Claude Opus 4.1 is $15 / 1M input tokens and $75 / 1M output tokens. > "GPT-5 will scaffold the app, write files, install dependencies as needed, and show a…
GPT-5
211–220 of 1001 posts
Re: GPT-5
#212# GPT5 all official links Livestream link: https://www.youtube.com/live/0Uu_VJeVVfo Research blog post: https://openai.com/index/introducing-gpt-5/ Developer blog post: https://openai.com/index/introducing-gpt-5-for-developers API Docs: https://platform.openai.com/docs/guides/latest-model Note the free form function calling documentation: https://platform.openai.com/docs/guides/function-calling#con... GPT5 prompting…
Re: GPT-5
#213What's going on with their SWE bench graph?[0] GPT-5 non-thinking is labeled 52.8% accuracy, but o3 is shown as a much shorter bar, yet it's labeled 69.1%. And 4o is an identical bar to o3, but it's labeled 30.8%... [0] https://i.postimg.cc/DzkZZLry/y-axis.png
Re: GPT-5
#214Re: GPT-5
#215> comments turned off yikes - the poor executive leadership’s fragile egos cannot take the criticism.
Re: GPT-5
#216ChatGPT5 in this demo: > For an airplane wing (airfoil), the top surface is curved and the bottom is flatter. When the wing moves forward: > * Air over the top has to travel farther in the same amount of time -> it moves faster -> pressure on the top decreases. > * Air underneath moves slower -> pressure underneath is higher > * The presure difference creates an upward force - lift Isn't that explanation of why wings…
Re: GPT-5
#217Watching the livestream now, the improvement over their current models on the benchmarks is very small. I know they seemed to be trying to temper our expectations leading up to this, but this is much less improvement than I was expecting
It is at least much cheaper and seems faster. They also announced gpt-5-pro but I haven't seen benchmarks on that yet.
Re: GPT-5
#218What's going on with their SWE bench graph?[0] GPT-5 non-thinking is labeled 52.8% accuracy, but o3 is shown as a much shorter bar, yet it's labeled 69.1%. And 4o is an identical bar to o3, but it's labeled 30.8%... [0] https://i.postimg.cc/DzkZZLry/y-axis.png
Re: GPT-5
#219The marketing copy and the current livestream appear tautological: "it's better because it's better." Not much explanation yet why GPT-5 warrants a major version bump. As usual, the model (and potentially OpenAI as a whole) will depend on output vibe checks.
How is this sustainable.
Re: GPT-5
#220The introduction said to try the following prompt Describe me based on all our chats — make it catchy! It was flattering as all get out, but fairly accurate (IMHO) Mike Warot: The Tinkerer of Tomorrow A hardware hacker with a poet’s soul, Mike blends old-school radio wisdom with cutting-edge curiosity. Whether he's decoding atomic clocks, reinventing FPGA logic with BitGrid, or pondering the electromagnetic vector po…