Live data from Hacker News

Claude Opus 4.8

anthropic.com

301–310 of 1001 posts

Re: Claude Opus 4.8

#301

Earlier quoted context omitted.

> My own experience w/ 4.6 and 4.7 are that I don't firmly grasp any capabilities improvements over my memory of 4.5, but it's all so fuzzy that it's truly difficult to tell. I've actually intentionally switched back to 4.5. I hated 4.7 so much that I decided to jump back all the way to 4.5. Now that I've been using 4.5 for a few weeks, I find it significantly more reliable but a bit more forgetful than 4.6/4.7. I'm…

If you are using Claude code, just set effort to xhigh. This one change will probably solve 80% of the problems you have noticed.

This. XHigh and the 'plan' mode for complex tasks is absolutely a must have.

Still, the context window is sometimes too small for my usage.

Re: Claude Opus 4.8

#302
post #73

I generated pelicans riding bicycles on both thinking level low and thinking level high: https://gist.github.com/simonw/68560eddb0b268a8417f80ceb7304... The high one is notably better - the bicycle frame is the correct shape, unlike thinking level low. For comparison, here's Opus 4.7: https://gist.github.com/simonw/afcb19addf3f38eb1996e1ebe749c...

You've peed in the pool Simon, this has to be a part of the internal evals by now! You got to try something new - maybe a panda in a canoe?

Click the link

Re: Claude Opus 4.8

#303

Give us Mythos! This piecemealing doesn't help Anthropic at all, especially psychologically! They are playing a dangerous game, and I see many people leaving Claude Code for good - both due to the subsidy games, and for Anthropic not dogfooding and using unreleased models internally and giving us subpar ones. Benchmarks are nice, but the real-world experience is quite different - neither can you notice these slight i…

I am also pushing my office to use chatgpt. Misanthropic thinks they are some kind of novel org doing whole humanity a favor...

Re: Claude Opus 4.8

#304
post #262

On my tests[0] it does a bit worse, and it's almost 2x expensive than Opus 4.7... I was surprised to see that it failed a Data extraction test (it gets it right 2/3 times, but one time it randomly returns null for a value instead). It makes sense a bit that it fails more Trivia/Domain-specific knowledge tasks (I think models are more and more trained towards agentic use-case than general intelligence). [0]: https://a…

Wait, doesn’t the blog post say the price is the same as 4.7?

> Claude Opus 4.8 is available everywhere today. Pricing for regular usage is unchanged from Opus 4.7: $5 per million input tokens and $25 per million output tokens. Pricing for fast mode is $10 per million input tokens and $50 per million output tokens.

Where do you see the 2x cost?

Re: Claude Opus 4.8

#305

The rapid release cadence and rate of innovation of Anthropic (and OpenAI) is impressive. And obviously it's because these are startups solely dedicated to AI so they can move quickly. Big Tech (like Google) won't be able to keep up with the pace of them (too much bureaucracy and red tape at Google). Classic Innovator's Dilemma. The longer a company exists, the more people, processes, and rules are added, which inevi…

I think big tech can catch up. Both Google and Meta have carved out startup like environments internally that move extremely fast. Neither OAI nor Anthropic can afford to rest on their laurels.

Re: Claude Opus 4.8

#307

"Users will find Opus 4.8 to be a modest but tangible improvement on its predecessor." This is a refreshing attitude! I've also verified that you can now turn off adaptive thinking in the web UI, which is great. I've had a lot of problems with thinking not triggering and the model producing sub-par output. Glad we can finally turn it off. (I hope being able to turn off adaptive thinking is new, if I could have turned…

I liked the "modest but tangible improvement" too! There is a cynical take here but I think I'm gonna hold it in...

Re: Claude Opus 4.8

#308

Probably explains why Opus was trash for the last week - https://marginlab.ai/trackers/claude-code/ . Curious if the new baseline will rise now in-line with the new benchmarks.

Nice. Can you release that for older models too? I've been using a mixture of releases recently, and cannot tell the difference between any of them.

Re: Claude Opus 4.8

#309
Ugh...

Invalid request The request couldn't be completed. View details API Error: 400 messages.1.content.7: `thinking` or `redacted_thinking` blocks in the latest assistant message cannot be modified. These blocks must remain as they were in the original response.

I would rather not. 4.6 was fine. 4.7 got to be fine 1 week after the release. Now 4.8. No difference, same thing.

But the app is broken and nothing works. So now I have to regress to different clients and wait it out while it becomes workable again.

Re: Claude Opus 4.8

#310
post #234

> Not only that, but we plan to release a new class of model with even higher intelligence than Opus. As part of Project Glasswing, a small number of organizations are currently using Claude Mythos Preview for cybersecurity work. Models of this capability level require stronger cyber safeguards before they can be generally released. We’re making swift progress on developing these safeguards and expect to be able to b…

Seems like they might be hinting that if you are not a billionaire or multi-billion dollar company you will just get a limited and nerfed Claude Code slash command /mythos-security-audit or something. Hope this isn’t the case and that normal average Joe’s of the world don’t get policed out of access.

It does sound like an even higher API price tier for sure.
Post reply on HN