Live data from Hacker News

GPT-5

openai.com

781–790 of 1001 posts

Re: GPT-5

#781
post #771

It is frequently suggested that once one of the AI companies reaches an AGI threshold, they will take off ahead of the rest. It's interesting to note that at least so far, the trend has been the opposite: as time goes on and the models get better, the performance of the different company's gets clustered closer together. Right now GPT-5, Claude Opus, Grok 4, Gemini 2.5 Pro all seem quite good across the board (ie the…

For those who happen to have a subscription to The Economist, there is a very interesting Money Talks podcast where they interview Anthropic's boss Dario Amodei[1]. There were two interesting takeaways about AGI: 1. Dario makes the remark that the term AGI/ASI is very misleading and dangerous. These terms are ill defined and it's more useful to understand that the capabilities are simply growing exponentially at the…

1. FWIW, I watched clips from several of Dario’s interviews. His expressions and body language convey sincere concerns.

2. Commoditization can be averted with access to proprietary data. This is why all of ChatGPT, Claude, and Gemini push for agents and permissions to access your private data sources now. They will not need to train on your data directly. Just adapting the models to work better with real-world, proprietary data will yield a powerful advantage over time.

Also, the current training paradigm utilizes RL much more extensively than in previous years and can help models to specialize in chosen domains.

Re: GPT-5

#782
Something that's really hitting me is something brought up in this piece:

https://www.interconnects.ai/p/gpt-5-and-bending-the-arc-of-...

When a model comes out, I usually think about it in terms of my own use. This is largely agentic tooling, and I mostly us Claude Code. All the hallucination and eval talk doesn't really catch me because I feel like I'm getting value of these tools today.

However, this model is not _for_ me in the same way models normally are. This is for the 800m or whatever people that open up chatgpt every day and type stuff in. All of them have been stuck on GPT-4o unbeknwst to them. They had no idea SOTA was far beyond that. They probably dont even know that there is a "model" at all. But for all these people, they just got a MAJOR upgrade. It will probably feel like turning the lights on for these people, who have been using a subpar model for the past year.

That said I'm also giving GPT-5 a run in Codex and it's doing a pretty good job!

Re: GPT-5

#783
post #470

I did a little test that I like to do with new models: "I have rectangular space of dimensions 30x30x90mm. Would 36x14x60mm battery fit in it, show in drawing proof". GPT5 failed spectacularly.

This was a fun prompt. I learned things from the models. Gemini 2.5 was wayy better than gpt5 here even though quite incomplete in the first response

Re: GPT-5

#784
I have a canonical test for chatbots -- I ask them who I am. I'm sufficiently unknown in modern times that it's a fair test. Just ask, "Who is Paul Lutus?"

ChatGPT 5's reply is mostly made up -- about 80% is pure invention. I'm described as having written books and articles whose titles I don't even recognize, or having accomplished things at odds with what was once called reality.

But things are slowly improving. In past ChatGPT versions I was described as having been dead for a decade.

I'm waiting for the day when, instead of hallucinating, a chatbot will reply, "I have no idea."

I propose a new technical Litmus test -- chatbots should be judged based on what they won't say.

Re: GPT-5

#785

Something that's really hitting me is something brought up in this piece: https://www.interconnects.ai/p/gpt-5-and-bending-the-arc-of-... When a model comes out, I usually think about it in terms of my own use. This is largely agentic tooling, and I mostly us Claude Code. All the hallucination and eval talk doesn't really catch me because I feel like I'm getting value of these tools today. However, this model is not…

I’m curious what this means. Maybe I’m stupid but I read through the sample gpt-4 vs got-5 and I largely couldn’t tell the difference and sometimes preferred the gpt-4 answer. But like what are the average 800 million people using this for that the average 800 million user will be able to see a difference?

Maybe I’m a far below average user? But I can’t tell the difference between models in causal use.

Unless you’re talking performance, apparently gpt-5 is much faster.

Re: GPT-5

#786

It is frequently suggested that once one of the AI companies reaches an AGI threshold, they will take off ahead of the rest. It's interesting to note that at least so far, the trend has been the opposite: as time goes on and the models get better, the performance of the different company's gets clustered closer together. Right now GPT-5, Claude Opus, Grok 4, Gemini 2.5 Pro all seem quite good across the board (ie the…

> It is frequently suggested that once one of the AI companies reaches an AGI threshold, they will take off ahead of the rest.

This seems to be a result of using overly simplistic models of progress. A company makes a breakthrough, the next breakthrough requires exploring many more paths. It is much easier to catch up than find a breakthrough. Even if you get lucky and find the next breakthrough before everyone catches up, they will probably catch up before you find the breakthrough after that. You only have someone run away if each time you make a breakthrough, it is easier to make the next breakthrough than to catch up.

Consider the following game:

1. N parties take turns rolling a D20. If anyone rolls 20, they get 1 point.

2. If any party is 1 or more points behind, they get only need to roll a 19 or higher to get one point. That is being behind gives you a slight advantage in catching up.

While points accumulate, most of the players end up with the same score.

I ran a simulation of this game for 10,000 turns with 5 players:

Game 1: [852, 851, 851, 851, 851]

Game 2: [827, 825, 827, 826, 826]

Game 3: [827, 822, 827, 827, 826]

Game 4: [864, 863, 860, 863, 863]

Game 5: [831, 828, 836, 833, 834]

Re: GPT-5

#788
post #771

It is frequently suggested that once one of the AI companies reaches an AGI threshold, they will take off ahead of the rest. It's interesting to note that at least so far, the trend has been the opposite: as time goes on and the models get better, the performance of the different company's gets clustered closer together. Right now GPT-5, Claude Opus, Grok 4, Gemini 2.5 Pro all seem quite good across the board (ie the…

For those who happen to have a subscription to The Economist, there is a very interesting Money Talks podcast where they interview Anthropic's boss Dario Amodei[1]. There were two interesting takeaways about AGI: 1. Dario makes the remark that the term AGI/ASI is very misleading and dangerous. These terms are ill defined and it's more useful to understand that the capabilities are simply growing exponentially at the…

There's already so many comparable models, and even local models are starting to approach the performance of the bigger server models.

I also feel like, it's stopped being exponential already. I mean last few releases we've only seen marginal improvements. Even this release feels marginal, I'd say it feels more like a linear improvement.

That said, we could see a winner take all due to the high cost of copying. I do think we're already approaching something where it's mostly price and who released their models last. But the cost to train is huge, and at some point it won't make sense and maybe we'll be left with 2 big players.

Re: GPT-5

#789

In terms of raw prose quality, I'm not convinced GPT-5 sounds "less like AI" or "more like a friend". Just count the number of em-dashes. It's become something of a LLM shibboleth.

Sorry, as someone who uses a lot of em-dashes (and semicolons, and other slightly less common punctuation) I find the whole em-dash thing to be completely unserious.

Re: GPT-5

#790
All people are talking about GPT-5 all over the world, the competition is so intense that every major tech company is racing to develop their own advanced AI models.
Post reply on HN