Live data from Hacker News

Claude Opus 4.8

anthropic.com

511–520 of 1001 posts

Re: Claude Opus 4.8

#511
post #485
post #385

My fav coding benchmark for frontier models is to build a simple RTS game in one file (js/html/css). Claude Code with Opus 4.8 in ultracode mode nailed it, the best result so far: https://bsky.app/profile/senko.net/post/3mmwnrkwboc2v The prompt was: Create a simple but functional real time strategy (RTS) game similar to old WarCraft, StarCraft or Command & Conquer games. The player should be able to build buildings,…

What is ultracode mode?

it's a brand new mode

Re: Claude Opus 4.8

#513
post #77

A rambling comment: I think this is the first time we've had a third minor version bump on a frontier Anthropic model. (I count the 0.5s as major here, because they've been issued non-sequentially and also corresponded to massive capability leaps, eg, Sonnet 3.5, Opus 4.5). So now the Opus 4.5 family has successors 4.6, 4.7, and 4.8, each posting fairly modest claimed gains. My own experience w/ 4.6 and 4.7 are that…

Well, it seems like collectively we are all struggling to perceive model progress, given that it seems like every reply to you is reporting different experiences with which of the models has subjectively performed best for them.

Re: Claude Opus 4.8

#514
post #385

My fav coding benchmark for frontier models is to build a simple RTS game in one file (js/html/css). Claude Code with Opus 4.8 in ultracode mode nailed it, the best result so far: https://bsky.app/profile/senko.net/post/3mmwnrkwboc2v The prompt was: Create a simple but functional real time strategy (RTS) game similar to old WarCraft, StarCraft or Command & Conquer games. The player should be able to build buildings,…

Kinda buggy, but impressively nonetheless. How long did it take?

Re: Claude Opus 4.8

#515

"Users will find Opus 4.8 to be a modest but tangible improvement on its predecessor." This is a refreshing attitude! I've also verified that you can now turn off adaptive thinking in the web UI, which is great. I've had a lot of problems with thinking not triggering and the model producing sub-par output. Glad we can finally turn it off. (I hope being able to turn off adaptive thinking is new, if I could have turned…

"We've cut our costs A LOT"

Re: Claude Opus 4.8

#516

Earlier quoted context omitted.

why are the models the same price? https://platform.claude.com/docs/en/about-claude/pricing ``` Model Base Input Tokens 5m Cache Writes 1h Cache Writes Cache Hits & Refreshes Output Tokens Claude Opus 4.8 $5 / MTok $6.25 / MTok $10 / MTok $0.50 / MTok $25 / MTok Claude Opus 4.7 $5 / MTok $6.25 / MTok $10 / MTok $0.50 / MTok $25 / MTok Claude Opus 4.6 $5 / MTok $6.25 / MTok $10 / MTok $0.50 / MTok $25 / MTok Claude Op…

Why shouldn’t they be? They are probably the same size and cost the same to run. They are not doing full training runs (eg Mythos) so don’t need to recover insane training costs.

I'd be kind of shocked if a model that came out six months ago is the same size and cost to run as one that just came out today.

Re: Claude Opus 4.8

#517

Earlier quoted context omitted.

I won't be surprised if the next gen frontier models are the last. There's orders of magnitude of low hanging juice to squeeze out of smaller models. It is almost guaranteed that a 60-90B model can outperform current SOTA in coding tasks within 2-3 years (design not certain, probably unlikely). It is far less clear that a 1.2T model will be meaningfully better enough to justify training it. As far as reasoning is con…

And anyway, with quantum, there will be no need for frontier companies as you might be able to even run a 1T param model on a consumer quantum computer.

I'm assuming this is a joke, but:

- why'd a quantum computer help running an LLM?

- of course there'd be need for frontier companies - nobody else has the resources to train frontier models.

Re: Claude Opus 4.8

#518

Earlier quoted context omitted.

> My own experience w/ 4.6 and 4.7 are that I don't firmly grasp any capabilities improvements over my memory of 4.5, but it's all so fuzzy that it's truly difficult to tell. I've actually intentionally switched back to 4.5. I hated 4.7 so much that I decided to jump back all the way to 4.5. Now that I've been using 4.5 for a few weeks, I find it significantly more reliable but a bit more forgetful than 4.6/4.7. I'm…

If you are using Claude code, just set effort to xhigh. This one change will probably solve 80% of the problems you have noticed.

Isn't xhigh on opus 4.7 very expensive on tokens?

Re: Claude Opus 4.8

#519

"Users will find Opus 4.8 to be a modest but tangible improvement on its predecessor." This is a refreshing attitude! I've also verified that you can now turn off adaptive thinking in the web UI, which is great. I've had a lot of problems with thinking not triggering and the model producing sub-par output. Glad we can finally turn it off. (I hope being able to turn off adaptive thinking is new, if I could have turned…

What's refreshing about it given the context that 4.7 was a regression in many ways (including as measured by benchmarks)? 4.8 is also 2x more expensive for a "modest" performance bump. How refreshing. This is just cope.

> 4.8 is also 2x more expensive for a "modest" performance bump. How refreshing.

Where are you seeing it's 2x more expensive? https://platform.claude.com/docs/en/about-claude/pricing

Re: Claude Opus 4.8

#520

"Users will find Opus 4.8 to be a modest but tangible improvement on its predecessor." This is a refreshing attitude! I've also verified that you can now turn off adaptive thinking in the web UI, which is great. I've had a lot of problems with thinking not triggering and the model producing sub-par output. Glad we can finally turn it off. (I hope being able to turn off adaptive thinking is new, if I could have turned…

What's refreshing about it given the context that 4.7 was a regression in many ways (including as measured by benchmarks)? 4.8 is also 2x more expensive for a "modest" performance bump. How refreshing. This is just cope.

Price hasn’t changes at all, though.
Post reply on HN