GPT-5.4
331–340 of 868 posts
Re: GPT-5.4
#332Earlier quoted context omitted.
Sonnet was pretty close to (or better than) Opus in a lot of benchmarks, I don't think it's a big deal
wat
https://artificialanalysis.ai indicates that sonnect 4.6 beats opus 4.6 on GDPval-AA, Terminal-Bench Hard, AA Long context Reasoning, IFBench.
see: https://artificialanalysis.ai/?models=claude-sonnet-4-6%2Ccl...
Re: GPT-5.4
#333Re: GPT-5.4
#334Earlier quoted context omitted.
If you look at the difference in quality between gpt-2 and 3, it feels like a big step, but the difference between 5.2 and 5.4 is more massive, it's just that they're both similarly capable and competent. I don't think it's an S curve; we're not plateauing. Million token context windows and cached prompts are a huge space for hacking on model behaviors and customization, without finetuning. Research is proceeding at…
For 2026, I am really interested in seeing whether local models can remain where they are: ~1 year behind the state of the art, to the point where a reasonably quantized November 2026 local model running on a consumer GPU actually performs like Opus 4.5. I am betting that the days of these AI companies losing money on inference are numbered, and we're going to be much more dependent on local capabilities sooner rathe…
My biggest worry is that the private jet class of people end up with absurdly powerful AI at their fingertips, while the rest of us are left with our BigMac McAIs.
Re: GPT-5.4
#335OpenAI now has three price points: GPT 5.1, GPT 5.2 and now GPT 5.4. There version numbers jump across different model lines with codex at 5.3, what they now call instant also at 5.3.
Anthropic are really the only ones who managed to get this under control: Three models, priced at three different levels. New models are immediately available everywhere.
Google essentially only has Preview models! The last GA is 2.5. As a developer, I can either use an outdated model or have zero insurances that the model doesn't get discontinued within weeks.
Re: GPT-5.4
#336Earlier quoted context omitted.
Ironically this would actually be a good thing. As we can see from Iran Claude doesn’t quite have these bugs ironed out yet…
This is the exact attitude that lead to a chat bot being used to identify a school for girls as a valid target. The chatbot cannot be held responsible. Whoever is using chatbots for selecting targets is incompetent and should likely face war crime charges.
Has it been stated authoritatively somewhere that this was an AI-driven mistake?
There are myrid ways that mistake could have been made that don't require AI. These kinds of mistakes were certainly made by all kinds of combatants in the pre-AI era.
Re: GPT-5.4
#337What a model mess! OpenAI now has three price points: GPT 5.1, GPT 5.2 and now GPT 5.4. There version numbers jump across different model lines with codex at 5.3, what they now call instant also at 5.3. Anthropic are really the only ones who managed to get this under control: Three models, priced at three different levels. New models are immediately available everywhere. Google essentially only has Preview models! Th…
Re: GPT-5.4
#338Re: GPT-5.4
#339Earlier quoted context omitted.
This is the exact attitude that lead to a chat bot being used to identify a school for girls as a valid target. The chatbot cannot be held responsible. Whoever is using chatbots for selecting targets is incompetent and should likely face war crime charges.
"that lead to a chat bot being used to identify a school for girls as a valid target" Has it been stated authoritatively somewhere that this was an AI-driven mistake? There are myrid ways that mistake could have been made that don't require AI. These kinds of mistakes were certainly made by all kinds of combatants in the pre-AI era.
Yeah yeah, they probably had a human in the loop, that’s not really the point though.