My fav coding benchmark for frontier models is to build a simple RTS game in one file (js/html/css). Claude Code with Opus 4.8 in ultracode mode nailed it, the best result so far: https://bsky.app/profile/senko.net/post/3mmwnrkwboc2v The prompt was: Create a simple but functional real time strategy (RTS) game similar to old WarCraft, StarCraft or Command & Conquer games. The player should be able to build buildings,…
Claude Opus 4.8
831–840 of 1001 posts
Re: Claude Opus 4.8
#832Which days in a week have the letter d in them?
Response:
Four: Monday, Tuesday, Wednesday, and Sunday.
Re: Claude Opus 4.8
#833Earlier quoted context omitted.
I'm here to complain about the churn. I feel like I get to know a model in the human sense of understanding a personality. Yesterday I knew 4.6 extended, today it's different, there's multiple "token budget" levels. I just want 4.6 extended back as it was, I was getting on well with it / them.
Humanizing this technology seems like a step in the wrong direction.
Re: Claude Opus 4.8
#834Opus 4.8: Which days in a week have the letter d in them? Response: Four: Monday, Tuesday, Wednesday, and Sunday.
Re: Claude Opus 4.8
#835Re: Claude Opus 4.8
#836I find it freaky how you notice the language change between models. Some words which pop up now all the time, that I don't remember reacting to with previous models, such as "honest(ly)" and "load-bearing". Feels like a new AI smell, like em-dashes or "it's not just x, it's y".
Re: Claude Opus 4.8
#837Given DeepSWE just blew apart the SWE-Bench Pro benchmark and handed a 14-point lead to GPT-5.5, it looks pretty bad that they've listed SWE-Bench first in the model release and no DeepSWE. Like, this isn't obviously an answer. Or maybe it is, but publish the DeepSWE numbers so we can see for ourselves.
This is a terrible benchmark. It literally tests the models on their ability to track shifting line numbers. If they cannot keep up, no amount of abstract reasoning can redeem them.
Re: Claude Opus 4.8
#838Re: Claude Opus 4.8
#839Re: Claude Opus 4.8
#840Early ArtificialAnalysis.ai results show GPT 5.5 is still the better bang-for-your-buck. OpenAI solves tasks with about 50% less output tokens. https://artificialanalysis.ai/?intelligence=coding-index&int...