Earlier quoted context omitted.
With Anthropic you always have 3 models to choose from: Opus-latest, Sonnet-latest, and Haiku-latest, from the best/slowest to the worst/fastest. The version numbers are mostly irrelevant as afaik price per token doesn't change between versions.
Three random names isn't ideal. I'm often need to double check which is which. This is why we use numbers
GPT-5.4
561–570 of 868 posts
Re: GPT-5.4
#562Re: GPT-5.4
#563Earlier quoted context omitted.
So what is your motivation for doing this, incidentally? Can you be explicit about it? I am genuinely curious. Especially when it’s to the point of, you know, nagging/policing people to do it the way you’d prefer, when you could just redirect your router requests from x.com to xcancel.com
It's not particularly about x.com, hundreds of site like x, youtube, facebook, linkedin, tiktok etc surreptitious add tracking parameters to their links. The iOS Messages app even hides these tracking parameters. I don't like being surreptitiously tracked online and judging by the success of my free app, there are millions of people like me.
i’m not being facetious, honest question, especially considering ads are the only thing paying these people these days
Re: GPT-5.4
#564I find it quite funny how this blog post has a big "Ask ChatGPT" box at the bottom. So you might think you could ask a question about the contents of the blog post, so you type the text "summarise this blog post". And it opens a new chat window with the link to the blog post followed by "summarise this blog post". Only to be told "I can't access external URLs directly, but if you can paste the relevant text or descri…
Re: GPT-5.4
#565Results from my Extended NYT Connections benchmark: GPT-5.4 extra high scores 94.0 (GPT-5.2 extra high scored 88.6). GPT-5.4 medium scores 92.0 (GPT-5.2 medium scored 71.4). GPT-5.4 no reasoning scores 32.8 (GPT-5.2 no reasoning scored 28.1).
Re: GPT-5.4
#566The "RPG Game" example on the blogpost is one of the most impressive demo's of autonomous engineering I've seen. It's very similar to "Battle Brothers", and the fact that RPG games require art assets, AI for enemy moves, and a host of other logical systems makes it all the more impressive.
However, I think what actually happened is that a skilled engineer made that game using codex. They could have made 100s of prompts after carefully reviewing all source code over hours or days.
The tycoon game is impressive for being made in a single prompt. They include the prompt for this one. They call it "lightly specified", but it's a pretty dense todo list for how to create assets, add many features from RollerCoaster Tycoon, and verify it works. I think it can probably pull a lot of inspiration from pretraining since RCT is an incredibly storied game.
The bridge flyover is hilariously bad. The bridge model ... has so many things wrong with it, the camera path clips into the ground and bridge, and the water and ground are z fighting. It's basically a C homework assignment that a student made in blender. It's impressive that it was able to achieve anything on such a visual task, but the bar is still on the floor. A game designer etc. looking for a prototype might actually prefer to greybox rather than have AI spend an hour making the worst bridge model ever.
Re: GPT-5.4
#567I've only used 5.4 for 1 prompt (edit: 3@high now) so far (reasoning: extra high, took really long), and it was to analyse my codebase and write an evaluation on a topic. But I found its writing and analysis thoughtful, precise, and surprisingly clearly written, unlike 5.3-Codex. It feels very lucid and uses human phrasing. It might be my AGENTS.md requiring clearer, simpler language, but at least 5.4's doing a good…
The latest research these days is that including an AGENTS.md file only makes outcomes worse with frontier models.
> somehow is the optimal strategy
My strategy of not spending an ounce of effort learning how to use AI beyond installing the Codex desktop app and telling it what to do keeps paying off lol.
Re: GPT-5.4
#568Earlier quoted context omitted.
If you're trying to use LLMs in an enterprise context, you would understand. Switching models sometimes requires tweaking prompts. That can be a complete mess, when there are dozens or hundreds of prompts you have to test.
This sounds made up. Much like “prompt engineering” Let’s hear an actual example
Sample size was 1000 jobs per prompt/model. We run them once per month to detect regression as well.
Re: GPT-5.4
#569Earlier quoted context omitted.
This is such a stale take. In the past 3 years I’ve worked on multiple products with AI at their core, not as some add-on. Just because the corpo-land dullards[0] can’t execute on anything more complex than shoehorning a chatbot into their offerings doesn’t mean there aren’t plenty of people and companies doing far more interesting things. [0] In this case, and with heavy irony, including OpenAI, although it sounds l…
Kinda reminds me of crypto. There are certainly very interesting things happening in the crypto space. But the most visible parts of the crypto universe are the stupid parts (buying PNGs for millions, for example)
Re: GPT-5.4
#570Earlier quoted context omitted.
I picked up Claude today after being away and using only ChatGPT and Gemini for a while. I was pretty impressed with how they’ve improved user experience. If I had to guess, I’d say Anthropic has better product people who put more attention to detail in these areas.
I agree! I recently migrated from ChatGPT to Claude and it is just superior in every way. It doesn't blather on the at the end ask me for clarification. It's succinct and clarifies vital information before providing a solution.