It is frequently suggested that once one of the AI companies reaches an AGI threshold, they will take off ahead of the rest. It's interesting to note that at least so far, the trend has been the opposite: as time goes on and the models get better, the performance of the different company's gets clustered closer together. Right now GPT-5, Claude Opus, Grok 4, Gemini 2.5 Pro all seem quite good across the board (ie the…
The reason AGI would create a singularity is because of its ability to self learn. Presently we are still a long way from that. In my opinion we at least are as far away from AGI as 1970s mainframes were from LLMs. I really don’t expect to see AGI in my lifetime.
GPT-5
761–770 of 1001 posts
Re: GPT-5
#762What does this say?
GPT 5:
When read normally without the ASCII art spacing, it’s the stylized text for:
markdown Copy Edit _ _ _ __ ___ _ __ ___ __ _ __| (_) ___ | '_ \ / _ \| '_ ` _ \ / _` |/ _` | |/ __| | | | | (_) | | | | | | (_| | (_| | | (__ |_| |_|\___/|_| |_| |_|\__,_|\__,_|_|\___| Which is the ASCII art for:
rust — the default “Rust” welcome banner in ASCII style.
Re: GPT-5
#763Re: GPT-5
#764It is frequently suggested that once one of the AI companies reaches an AGI threshold, they will take off ahead of the rest. It's interesting to note that at least so far, the trend has been the opposite: as time goes on and the models get better, the performance of the different company's gets clustered closer together. Right now GPT-5, Claude Opus, Grok 4, Gemini 2.5 Pro all seem quite good across the board (ie the…
It's also worth considering that past some threshold, it may be very difficult for us as users to discern which model is better. I don't think thats what's going on here, but we should be ready for it. For example, if you are an ELO 1000 chess player would you yourself be able to tell if Magnus Carlson or another grandmaster were better by playing them individually? To the extent that our AGI/SI metrics are based on…
But one thing will stay consistent with LLMs for some time to come: they are programmed to produce output that looks acceptable, but they all unintentionally tend toward deception. You can iterate on that over and over, but there will always be some point where it will fail, and the weight of that failure will only increase as it deceives better.
Some things that seemed safe enough: Hindenburg, Titanic, Deepwater Horizon, Chernobyl, Challenger, Fukushima, Boeing 737 MAX.
Re: GPT-5
#765It is frequently suggested that once one of the AI companies reaches an AGI threshold, they will take off ahead of the rest. It's interesting to note that at least so far, the trend has been the opposite: as time goes on and the models get better, the performance of the different company's gets clustered closer together. Right now GPT-5, Claude Opus, Grok 4, Gemini 2.5 Pro all seem quite good across the board (ie the…
Since then they've been about neck and neck with some models making different tradeoffs.
Nobody needs to reach AGI to take off. They just need to bankrupt their competitors since they're all spending so much money.
Re: GPT-5
#766> 400,000 context window > 128,000 max output tokens > Input $1.25 > Output $10.00 Source: https://platform.openai.com/docs/models/gpt-5 If this performs well in independent needle-in-haystack and adherence evaluations, this pricing with this context window alone would make GPT-5 extremely competitive with Gemini 2.5 Pro and Claude Opus 4.1, even if the output isn't a significant improvement over o3. If the output qu…
You also have to count the cost of having to verify your identity to use the API
I understand for image generation, but why for text generation?
Re: GPT-5
#767It is frequently suggested that once one of the AI companies reaches an AGI threshold, they will take off ahead of the rest. It's interesting to note that at least so far, the trend has been the opposite: as time goes on and the models get better, the performance of the different company's gets clustered closer together. Right now GPT-5, Claude Opus, Grok 4, Gemini 2.5 Pro all seem quite good across the board (ie the…
It's also worth considering that past some threshold, it may be very difficult for us as users to discern which model is better. I don't think thats what's going on here, but we should be ready for it. For example, if you are an ELO 1000 chess player would you yourself be able to tell if Magnus Carlson or another grandmaster were better by playing them individually? To the extent that our AGI/SI metrics are based on…
Re: GPT-5
#768I'm not really convinced, the benchmark blunder was really strange but the demos were quite underwhelming, and it appears this was reflected by a huge market correction in the betting markets as to who will have the best AI by end of the year. What excites me now is that Gemini 3.0 or some answer from Google is coming soon and that will be the one I will actually end up using. It seems like the last mover in the LLM…
Polymarket betters are not impressed. Based upon the market odds, OpenAI had a 35% chance to have the best model (at year end), but those odds have dropped to 18% today. (I'm mostly making this comment to document what happened for the history books.) https://polymarket.com/event/which-company-has-best-ai-model...
Re: GPT-5
#769What's going on with this plot's y-axis? https://bsky.app/profile/tylermw.com/post/3lvtac5hues2n
It makes it look like the presentation is rushed or made last minute. Really bad to see this as the first plot in the whole presentation. Also, I would have loved to see comparisons with Opus 4.1. Edit: Opus 4.1 scores 74.5% ( https://www.anthropic.com/news/claude-opus-4-1 ). This makes it sound like Anthropic released the upgrade to still be the leader on this important benchmark.