GPT4 gave her better response than doctors she said.
GPT-5
251–260 of 1001 posts
Re: GPT-5
#252[flagged]
If you email us at hn@ycombinator.com and tell us who you want to contact, we might be able to email them and ask if they would be willing to have you contact them. No guarantees though!
Re: GPT-5
#253What's going on with their SWE bench graph?[0] GPT-5 non-thinking is labeled 52.8% accuracy, but o3 is shown as a much shorter bar, yet it's labeled 69.1%. And 4o is an identical bar to o3, but it's labeled 30.8%... [0] https://i.postimg.cc/DzkZZLry/y-axis.png
Re: GPT-5
#254ChatGPT5 in this demo: > For an airplane wing (airfoil), the top surface is curved and the bottom is flatter. When the wing moves forward: > * Air over the top has to travel farther in the same amount of time -> it moves faster -> pressure on the top decreases. > * Air underneath moves slower -> pressure underneath is higher > * The presure difference creates an upward force - lift Isn't that explanation of why wings…
Source: PhD on aircraft design
Re: GPT-5
#255Pricing seems good, but the open question is still on tool calling reliability. Input: $1.25 / 1M tokens Cached: $0.125 / 1M tokens Output: $10 / 1M tokens With 74.9% on SWE-bench, this inches out Claude Opus 4.1 at 74.5%, but at a much cheaper cost. For context, Claude Opus 4.1 is $15 / 1M input tokens and $75 / 1M output tokens. > "GPT-5 will scaffold the app, write files, install dependencies as needed, and show a…
Re: GPT-5
#256Watching the livestream now, the improvement over their current models on the benchmarks is very small. I know they seemed to be trying to temper our expectations leading up to this, but this is much less improvement than I was expecting
Sam said maybe two years ago that they want to avoid "mic drop" releases, and instead want to stick to incremental steps. This is day one, so there is probably another 10-20% in optimizations that can be squeezed out of it in the coming months.
Re: GPT-5
#257Re: GPT-5
#258What's going on with their SWE bench graph?[0] GPT-5 non-thinking is labeled 52.8% accuracy, but o3 is shown as a much shorter bar, yet it's labeled 69.1%. And 4o is an identical bar to o3, but it's labeled 30.8%... [0] https://i.postimg.cc/DzkZZLry/y-axis.png
Re: GPT-5
#259What's going on with this plot's y-axis? https://bsky.app/profile/tylermw.com/post/3lvtac5hues2n
Re: GPT-5
#260Damn, you guys are toxic. So -- they did not invent AGI yet. Yet, I like what I'm seeing. Major progress on multiple fronts. Hallucination fix is exciting on its own. The React demos were mindblowing.