Live data from Hacker News

GPT-5

openai.com

961–970 of 1001 posts

Re: GPT-5

#961

It is frequently suggested that once one of the AI companies reaches an AGI threshold, they will take off ahead of the rest. It's interesting to note that at least so far, the trend has been the opposite: as time goes on and the models get better, the performance of the different company's gets clustered closer together. Right now GPT-5, Claude Opus, Grok 4, Gemini 2.5 Pro all seem quite good across the board (ie the…

If AGI is ever achieved, it would open the door to recursive self improvement that would presumably rapidly exceed human capability across any and all fields, including AI development. So the AI would be improving itself while simultaneously also making revolutionary breakthroughs in essentially all fields. And, for at least a while, it would also presumably be doing so at an exponentially increasing rate. But I thin…

AI can be trained on some special knowledge of person A and another special knowledge of person B. These two persons may never met before and therefore they can not combine their knowledge to get some new knowledge or insight.

AI can do it fine as it knows A and B. And that is knowledge creation.

Re: GPT-5

#962

I am thoroughly unimpressed by GPT-5. It still can't compose iambic trimeters in ancient Greek with a proper penthemimeral cæsura, and it insists on providing totally incorrect scansion of the flawed lines it does compose. I corrected its metrical sins twice, which sent it into "thinking" mode until it finally returned a "Reasoning failed" error. There is no intelligence here: it's still just giving plausible output.…

It's well-known at this point that LLMs don't handle spelling, syllables, rhythm, meter, or other word-form-based questions well due to tokenization -- sometimes sheer scale (or leaning on code) can get the right answer if they're lucky, but they're literally blind to the individual letters.

(Incidentally, go back in time even five years and this specific expectation of AI capability sounds comically overblown. "Everything's amazing and nobody's happy.")

Re: GPT-5

#964

Ok this[0] sounds very, uh bold to me? Surely this is going to break a ton of workflows etc seemingly with nearly no notice? I'm assuming 'launches' equates with 'fully rolls out' or something but it's not that clear to me. When GPT-5 launches, several older models will be retired, including: - GPT-4o - GPT-4.1 - GPT-4.5 - GPT-4.1-mini - o4-mini - o4-mini-high - o3 - o3-pro If you open a conversation that used one of…

This behavior should be an early warning sign of future potential enshitification and a reason to consider open weight models you can host elsewhere.

If you are building on models that could disappear tomorrow when a company needs to juice the launch of a new model (or increase prices), you are introducing avoidable risk.

Re: GPT-5

#965

It is frequently suggested that once one of the AI companies reaches an AGI threshold, they will take off ahead of the rest. It's interesting to note that at least so far, the trend has been the opposite: as time goes on and the models get better, the performance of the different company's gets clustered closer together. Right now GPT-5, Claude Opus, Grok 4, Gemini 2.5 Pro all seem quite good across the board (ie the…

If AGI is ever achieved, it would open the door to recursive self improvement that would presumably rapidly exceed human capability across any and all fields, including AI development. So the AI would be improving itself while simultaneously also making revolutionary breakthroughs in essentially all fields. And, for at least a while, it would also presumably be doing so at an exponentially increasing rate. But I thin…

That's only assuming there are no fundamental limits or major barriers to computation. Back a hundred years ago at the dawn of flight, one could have said a very similar thing about aircraft performance. And for a time in the 1950s, it looked like aircraft speed was growing exponentially over time. But there haven't been any new airspeed records (at least, officially recorded) since 1986, because it turns out going Mach 3+ is fairly dangerous and approaching some rather severe materials and propulsion limitations, making it not at all economical.

I would also not be surprised if the process of developing something comparable to human intelligence, assuming the extreme computation, energy, and materials issues of packing that much computation and energy into a single system could be overcome, the AI also develops something comparable to human desire and/or mental health issues. There is a not-zero chance we could end up with AI that doesn't want to do what we ask it to do or doesn't work all the time because it wants to do other things.

You can't just assume exponential growth is a forgone conclusion.

Re: GPT-5

#966
First impressions: the emoji trigger happiness of 4o is totally gone. Bolding still happens.

There appear to be 4 ways to run a query now: a) GPT5, b) GPT5 and toggle "extra thinking" on, c) "GPT5 with thinking", and d) "GPT5 with thinking" then click "quick answer" which aborts thinking (this mode is possibly identical with GPT5)

I don't find this much simpler than 4o, o3, etc. It's just reordering the hierarchies. Now the model name is no longer descriptive at all and one has to add which mode one ran it in.

Re: GPT-5

#967

They really nerfed Plus[0]. 80 messages every 3 hours for normal GPT-5. And only 200 messages per week for GPT-5 Thinking. It seems like terrible value. Before it was: 100 o3 per week 100 o4-mini-high per day 300 o4-mini per day 50 4.5 per week [0] https://help.openai.com/en/articles/11909943-gpt-5-in-chatgp...

Oh nice so if you just used o3 (like me) there's an increase If it's not good I'll just unsubscribe. 100 o3 was enough for me and I used o3 almost exclusively. So I'm not super worried. If it's no longer SOTA and Gemini or Claude are better for everything I'll just cancel.

Re: GPT-5

#968

I'm not really convinced, the benchmark blunder was really strange but the demos were quite underwhelming, and it appears this was reflected by a huge market correction in the betting markets as to who will have the best AI by end of the year. What excites me now is that Gemini 3.0 or some answer from Google is coming soon and that will be the one I will actually end up using. It seems like the last mover in the LLM…

Polymarket betters are not impressed. Based upon the market odds, OpenAI had a 35% chance to have the best model (at year end), but those odds have dropped to 18% today. (I'm mostly making this comment to document what happened for the history books.) https://polymarket.com/event/which-company-has-best-ai-model...

Looking at LMarena which polymarket uses, I'm not surprised. Based on the little data there is (3k duels, it's possibly worse than Gemini, it lost more to Gemini 2.5 Pro than it won in direct duels). Not sure why the ELO is still higher, possibly GPT5 did more clearly better against bad models, which I don't care about.

Re: GPT-5

#969

It is frequently suggested that once one of the AI companies reaches an AGI threshold, they will take off ahead of the rest. It's interesting to note that at least so far, the trend has been the opposite: as time goes on and the models get better, the performance of the different company's gets clustered closer together. Right now GPT-5, Claude Opus, Grok 4, Gemini 2.5 Pro all seem quite good across the board (ie the…

It's also worth considering that past some threshold, it may be very difficult for us as users to discern which model is better. I don't think thats what's going on here, but we should be ready for it. For example, if you are an ELO 1000 chess player would you yourself be able to tell if Magnus Carlson or another grandmaster were better by playing them individually? To the extent that our AGI/SI metrics are based on…

It's even more difficult because, while all the benchmarks provide some kind of 'averaged' performance metric for comparison, in my experience most users have pretty specific regular use cases, and pretty specific personal background knowledge. For instance I have a background in ML, 15 years experience in full stack programming, and primarily use LLMs for generating interface prototypes for new product concepts. We use a lot of react and chakraui for that, and I consistently get the best results out of Gemini pro for that. I tried all the available options and settled on that as the best for me and my use case. It's not the best for marketing boilerplate, or probably a million other use cases, but for me, in this particular niche it's clearly the best. Beyond that the benchmarks are irrelevant.

Re: GPT-5

#970
post #664

What's going on with this plot's y-axis? https://bsky.app/profile/tylermw.com/post/3lvtac5hues2n

Also this coding deception rate bar tries to decieve us. https://imgur.com/a/QkriFco

It’s beyond parody that they did something like this on a slide about deception. You couldn’t make this stuff up.
Post reply on HN