Live data from Hacker News

GPT-5

openai.com

761–770 of 1001 posts

Re: GPT-5

#761

It is frequently suggested that once one of the AI companies reaches an AGI threshold, they will take off ahead of the rest. It's interesting to note that at least so far, the trend has been the opposite: as time goes on and the models get better, the performance of the different company's gets clustered closer together. Right now GPT-5, Claude Opus, Grok 4, Gemini 2.5 Pro all seem quite good across the board (ie the…

The reason AGI would create a singularity is because of its ability to self learn. Presently we are still a long way from that. In my opinion we at least are as far away from AGI as 1970s mainframes were from LLMs. I really don’t expect to see AGI in my lifetime.

should be next year in math domain tbh

Re: GPT-5

#762
My favorite thing to ask is ascii art: _ _ _ __ ___ _ __ ___ __ _ __| (_) ___ | '_ \ / _ \| '_ _ \ / _ |/ _ | |/ __| | | | | (_) | | | | | | (_| | (_| | | (__ |_| |_|\___/|_| |_| |_|\__,_|\__,_|_|\___|

What does this say?

GPT 5:

When read normally without the ASCII art spacing, it’s the stylized text for:

markdown Copy Edit _ _ _ __ ___ _ __ ___ __ _ __| (_) ___ | '_ \ / _ \| '_ ` _ \ / _` |/ _` | |/ __| | | | | (_) | | | | | | (_| | (_| | | (__ |_| |_|\___/|_| |_| |_|\__,_|\__,_|_|\___| Which is the ASCII art for:

rust — the default “Rust” welcome banner in ASCII style.

Re: GPT-5

#763
I asked it how to run the image and expose a port. it was just terrible in cursor. thought a Dockerfile wasn't in the repo, called no tools, then hallucinated a novel on dockefile best practices.

Re: GPT-5

#764

It is frequently suggested that once one of the AI companies reaches an AGI threshold, they will take off ahead of the rest. It's interesting to note that at least so far, the trend has been the opposite: as time goes on and the models get better, the performance of the different company's gets clustered closer together. Right now GPT-5, Claude Opus, Grok 4, Gemini 2.5 Pro all seem quite good across the board (ie the…

It's also worth considering that past some threshold, it may be very difficult for us as users to discern which model is better. I don't think thats what's going on here, but we should be ready for it. For example, if you are an ELO 1000 chess player would you yourself be able to tell if Magnus Carlson or another grandmaster were better by playing them individually? To the extent that our AGI/SI metrics are based on…

> it may be very difficult for us as users to discern which model is better

But one thing will stay consistent with LLMs for some time to come: they are programmed to produce output that looks acceptable, but they all unintentionally tend toward deception. You can iterate on that over and over, but there will always be some point where it will fail, and the weight of that failure will only increase as it deceives better.

Some things that seemed safe enough: Hindenburg, Titanic, Deepwater Horizon, Chernobyl, Challenger, Fukushima, Boeing 737 MAX.

Re: GPT-5

#765

It is frequently suggested that once one of the AI companies reaches an AGI threshold, they will take off ahead of the rest. It's interesting to note that at least so far, the trend has been the opposite: as time goes on and the models get better, the performance of the different company's gets clustered closer together. Right now GPT-5, Claude Opus, Grok 4, Gemini 2.5 Pro all seem quite good across the board (ie the…

The race has always been very close IMO. What Google had internally before ChatGPT first came out was mind blowing. ChatGPT was a let down comparatively (to me personally anyway).

Since then they've been about neck and neck with some models making different tradeoffs.

Nobody needs to reach AGI to take off. They just need to bankrupt their competitors since they're all spending so much money.

Re: GPT-5

#766
post #281

> 400,000 context window > 128,000 max output tokens > Input $1.25 > Output $10.00 Source: https://platform.openai.com/docs/models/gpt-5 If this performs well in independent needle-in-haystack and adherence evaluations, this pricing with this context window alone would make GPT-5 extremely competitive with Gemini 2.5 Pro and Claude Opus 4.1, even if the output isn't a significant improvement over o3. If the output qu…

You also have to count the cost of having to verify your identity to use the API

To clarify, you need to verify identity to use the GPT-5 API?

I understand for image generation, but why for text generation?

Re: GPT-5

#767

It is frequently suggested that once one of the AI companies reaches an AGI threshold, they will take off ahead of the rest. It's interesting to note that at least so far, the trend has been the opposite: as time goes on and the models get better, the performance of the different company's gets clustered closer together. Right now GPT-5, Claude Opus, Grok 4, Gemini 2.5 Pro all seem quite good across the board (ie the…

It's also worth considering that past some threshold, it may be very difficult for us as users to discern which model is better. I don't think thats what's going on here, but we should be ready for it. For example, if you are an ELO 1000 chess player would you yourself be able to tell if Magnus Carlson or another grandmaster were better by playing them individually? To the extent that our AGI/SI metrics are based on…

My guess is that more than the raw capabilities of a model, users would be drawn more to the model's personality. A "better" model would then be one that can closely adopt the nuances that a user likes. This is a largely uninformed guess, let's see if it holds up well with time.

Re: GPT-5

#768

I'm not really convinced, the benchmark blunder was really strange but the demos were quite underwhelming, and it appears this was reflected by a huge market correction in the betting markets as to who will have the best AI by end of the year. What excites me now is that Gemini 3.0 or some answer from Google is coming soon and that will be the one I will actually end up using. It seems like the last mover in the LLM…

Polymarket betters are not impressed. Based upon the market odds, OpenAI had a 35% chance to have the best model (at year end), but those odds have dropped to 18% today. (I'm mostly making this comment to document what happened for the history books.) https://polymarket.com/event/which-company-has-best-ai-model...

How on Earth does that market have Anthropic at 2%, in a dead heat with the likes of Meta? If the market was about yesterday rather than 5 months from now I think Claude would be pretty clearly the front runner. Why does the market so confidently think they’ll drop to dead last in the next little while?

Re: GPT-5

#769

What's going on with this plot's y-axis? https://bsky.app/profile/tylermw.com/post/3lvtac5hues2n

It makes it look like the presentation is rushed or made last minute. Really bad to see this as the first plot in the whole presentation. Also, I would have loved to see comparisons with Opus 4.1. Edit: Opus 4.1 scores 74.5% ( https://www.anthropic.com/news/claude-opus-4-1 ). This makes it sound like Anthropic released the upgrade to still be the leader on this important benchmark.

They never compare with other vendors

Re: GPT-5

#770
So far GPT-5 has not been able to pass my personal "Turing test" which has been unsuccessful for the past several years starting through various versions of Dall-e up to the latest model. I want it to create an image of Santa Claus pulling the sleigh with a reindeer in the sleigh holding the reins, driving the sleigh. No matter how I modify the prompt it is still unable to create this image that my daughter requested a few years ago. This is an image that is easily imagined and drawn by a small child yet the most advanced AI models still can't produce it. I think this is a good example that these models are unable to "imagine" something that falls outside of the realm of it's training data.
Post reply on HN