Live data from Hacker News

GPT-5

openai.com

711–720 of 1001 posts

Re: GPT-5

#711

Watching the livestream now, the improvement over their current models on the benchmarks is very small. I know they seemed to be trying to temper our expectations leading up to this, but this is much less improvement than I was expecting

I mean that's just the consequence of releasing a new model every couple months. If Open AI stayed mostly silent since the GPT-4 release (like they did for most iterations) and only now released 5 then nobody would be complaining about weak gains in benchmarks.

If they had stayed silent since GPT-4, nobody would care what OpenAI was releasing as they would have become completely irrelevant compared to Gemini/Claude.

Re: GPT-5

#712
post #92

It says out now in chatgpt. Did anyone yet hit the usage limits to report back how many messages are possible?

> Did anyone yet hit the usage limits to report back how many messages are possible?

10 messages every 5 hours on GPT-5 for free users, then it uses GPT-5-mini.

80 messages every 3 hours on GPT-5 for Plus users, then it uses GPT-5-mini (In fact, I tested this and was not allowed to use the mini model until I’ve exhausted my GPT-5-Thinking quota. That seems to be a bug.)

200 messages per week on GPT-5-Thinking on Plus and Team.

Unlimited GPT-5 on Team and Pro, subject to abuse guardrails.

Re: GPT-5

#713

It is frequently suggested that once one of the AI companies reaches an AGI threshold, they will take off ahead of the rest. It's interesting to note that at least so far, the trend has been the opposite: as time goes on and the models get better, the performance of the different company's gets clustered closer together. Right now GPT-5, Claude Opus, Grok 4, Gemini 2.5 Pro all seem quite good across the board (ie the…

The reason AGI would create a singularity is because of its ability to self learn. Presently we are still a long way from that. In my opinion we at least are as far away from AGI as 1970s mainframes were from LLMs. I really don’t expect to see AGI in my lifetime.

There are areas where we seem to be much closer to AGI than most people realize. AGI for software development, in particular, seems incredibly close. For example, Claude Code has bewildering capabilities that feel like magic. Mix it with a team of other capable development-oriented AIs and you might be able to build AI software that builds better AI software, all by itself.

Re: GPT-5

#714

It is frequently suggested that once one of the AI companies reaches an AGI threshold, they will take off ahead of the rest. It's interesting to note that at least so far, the trend has been the opposite: as time goes on and the models get better, the performance of the different company's gets clustered closer together. Right now GPT-5, Claude Opus, Grok 4, Gemini 2.5 Pro all seem quite good across the board (ie the…

Perhaps it is not possible to simulate higher-level intelligence using a stochastic model for predicting text. I am not an AI researcher, but I have friends who do work in the field, and they are not worried about LLM-based AGI because of the diminishing returns on results vs amount of training data required. Maybe this is the bottleneck. Human intelligence is markedly different from LLMs: it requires far fewer examp…

To be smarter than human intelligence you need smarter than human training data. Humans already innately know right and wrong a lot of the time so that doesn't leave much room.

Re: GPT-5

#715

GPT-5 knowledge cutoff: Sep 30, 2024 (10 months before release). Compare that to Gemini 2.5 Pro knowledge cutoff: Jan 2025 (3 months before release) Claude Opus 4.1: knowledge cutoff: Mar 2025 (4 months before release) https://platform.openai.com/docs/models/compare https://deepmind.google/models/gemini/pro/ https://docs.anthropic.com/en/docs/about-claude/models/overv...

with web search, is knowledge cutoff really relevant anymore? Or is this more of a comment on how long it took them to do post-training?

I've been having a lot of issues with chatgpt's knowledge of DuckDb being out of date. It doesn't think DuckDb enforces foreign keys, for instance.

Re: GPT-5

#717
> a real-time router that quickly decides which model to use based on conversation type, complexity, tool needs, and explicit intent

I'd love to see factors considered in the algorithm for system-1 vs system 2 thinking.

Is "complexity" the factor that says "hard problem"? Because it's often not the complexity that makes it hard.

Re: GPT-5

#718

It is frequently suggested that once one of the AI companies reaches an AGI threshold, they will take off ahead of the rest. It's interesting to note that at least so far, the trend has been the opposite: as time goes on and the models get better, the performance of the different company's gets clustered closer together. Right now GPT-5, Claude Opus, Grok 4, Gemini 2.5 Pro all seem quite good across the board (ie the…

We don’t seem to be closer to AGI however.

Re: GPT-5

#719

Tech aside (covered well by other commenters), the presentation itself was incredibly dry. Such a stark difference in presenting style here compared to, for example, Apple's or Google's keynotes. They should really put more effort into it.

I thought I was in the wrong live thread.

This seemed like a presentation you'd give to a small org, not a presentation a $500B company would give to release it's newest, greatest thing.

Re: GPT-5

#720
post #600

It is frequently suggested that once one of the AI companies reaches an AGI threshold, they will take off ahead of the rest. It's interesting to note that at least so far, the trend has been the opposite: as time goes on and the models get better, the performance of the different company's gets clustered closer together. Right now GPT-5, Claude Opus, Grok 4, Gemini 2.5 Pro all seem quite good across the board (ie the…

I'm still stuck at the bit where just throwing more and more data to make a very complex encyclopedia with an interesting search interface that tricks us into believing it's human-like gets us to AGI when we have no examples and thus no evidence or understanding of where the GI part comes from. It's all just hyperbole to attract investment and shareholder value and the people peddling the idea of AGI as a tangible po…

Me too. Some of them are frauds, but most of the weird AI-as-messiah people really believe it as far as I can tell.

The tech is neat and it can do some neat things but...it's a bullshit machine fueled by a bullshit machine hype bubble. I do not get it.

Post reply on HN