Live data from Hacker News

Claude Opus 4.8

anthropic.com

831–840 of 1001 posts

Re: Claude Opus 4.8

#831
post #385

My fav coding benchmark for frontier models is to build a simple RTS game in one file (js/html/css). Claude Code with Opus 4.8 in ultracode mode nailed it, the best result so far: https://bsky.app/profile/senko.net/post/3mmwnrkwboc2v The prompt was: Create a simple but functional real time strategy (RTS) game similar to old WarCraft, StarCraft or Command & Conquer games. The player should be able to build buildings,…

I wonder if your previous prompts were part of the new RL fine tuning, and that’s why is now better at this specific question

Re: Claude Opus 4.8

#833

Earlier quoted context omitted.

I'm here to complain about the churn. I feel like I get to know a model in the human sense of understanding a personality. Yesterday I knew 4.6 extended, today it's different, there's multiple "token budget" levels. I just want 4.6 extended back as it was, I was getting on well with it / them.

Humanizing this technology seems like a step in the wrong direction.

There's so much intelligence here on HN and so little humanity.

Re: Claude Opus 4.8

#834
post #832

Opus 4.8: Which days in a week have the letter d in them? Response: Four: Monday, Tuesday, Wednesday, and Sunday.

It seems like they’ve been optimising their models for coding. That’s what the benchmarks used in the article suggest at least.

Re: Claude Opus 4.8

#836

I find it freaky how you notice the language change between models. Some words which pop up now all the time, that I don't remember reacting to with previous models, such as "honest(ly)" and "load-bearing". Feels like a new AI smell, like em-dashes or "it's not just x, it's y".

That’s a really sharp observation. Hopefully they take a belt and suspenders approach to these smoking guns in future.

Re: Claude Opus 4.8

#837
post #676

Given DeepSWE just blew apart the SWE-Bench Pro benchmark and handed a 14-point lead to GPT-5.5, it looks pretty bad that they've listed SWE-Bench first in the model release and no DeepSWE. Like, this isn't obviously an answer. Or maybe it is, but publish the DeepSWE numbers so we can see for ourselves.

This is a terrible benchmark. It literally tests the models on their ability to track shifting line numbers. If they cannot keep up, no amount of abstract reasoning can redeem them.

Where did you get that idea? It uses mini-swe-agent, same as SWE-Bench.

https://github.com/datacurve-ai/deep-swe

Re: Claude Opus 4.8

#840

Early ArtificialAnalysis.ai results show GPT 5.5 is still the better bang-for-your-buck. OpenAI solves tasks with about 50% less output tokens. https://artificialanalysis.ai/?intelligence=coding-index&int...

[deleted]
Post reply on HN