Live data from Hacker News

Claude Opus 4.8

anthropic.com

961–970 of 1001 posts

Re: Claude Opus 4.8

#961
post #909

Earlier quoted context omitted.

pretty spot on. In my experience, Opus 4.0 was fantastic, major jump from 3.7. it was creative, super slow and expensive, and would sometime forget what it was doing, but it was getting the job done. 4.1 they made it much faster, so a lot of infra improvements. 4.5 was the time it could work on longer task, didn't make a lot of obvious mistakes of 4.0, and i think this was about the time the opus went mainstream, and…

> "4.6 was such a bad model," It's just amusing reading all these posts with different viewpoints, just in this thread there are multiple people saying 4.6 was so much better than 4.7 and that they switched back to 4.6.

that is a fair point, everything i said above was in my experience.

* in our experience, in our evals and codebase, 4.6 was a bad model. This is over 60k developers, so statistically significant.

Re: Claude Opus 4.8

#963

Earlier quoted context omitted.

I like that benchmark. You should throw the results up on GitHub pages so people can try out the games.

Yeah! Host on GitHub pages, so it's easy to click a link and play!

I put a version on Hallway: https://hallway.com/workspaces/4ddaa042-13b1-4fa5-bcf4-3d646...

Easy to edit and share.

Re: Claude Opus 4.8

#965
post #135
post #77

A rambling comment: I think this is the first time we've had a third minor version bump on a frontier Anthropic model. (I count the 0.5s as major here, because they've been issued non-sequentially and also corresponded to massive capability leaps, eg, Sonnet 3.5, Opus 4.5). So now the Opus 4.5 family has successors 4.6, 4.7, and 4.8, each posting fairly modest claimed gains. My own experience w/ 4.6 and 4.7 are that…

I'm curious to poll HN on this issue. Do you feel like we've had meaningful/noticeable gains in terms of your programming workflows between 4.5 and 4.7? My 2¢, I personally feel like all of the productivity gains since 4.5's release (in November 2025!!) have come from improvements to the harnesses (cc, cursor cli, codex, opencode, whatever) AND from the context window expansion from 200k to 1M. But the actual "raw" i…

For my day to day tasks 4.6 feels sufficient.

I have limited enterprise budget and Claude 4.7 costs 7x more. So unless there's close to 7x improvement, it doesn't make sense to switch to 4.7.

I actually gave both 4.6 a really complex task. It kept on thinking for several minutes before I hit the brakes. I then gave 4.7 the same task, and didn't notice any difference in behavior. Clearly not worth the 7x premium.

I hope 4.6 becomes cheaper/free at some point because I'm starting to see a push towards optimizing token expenditures across the board. While frontier models are still the default for developing new workflows, everybody is starting to ask how to automate repetitive tasks without using tokens.

Re: Claude Opus 4.8

#966
post #749

I use 4.6, because 4.7 is super lazy, deflects responsibility, and assumes it is good and I am bad, and avoids checking reality. It looks like it's trained on lazy humans instead of good engineers. Should I try 4.8? I am happy with 4.6. I am not happy with 4.7.

I still use 4.7. I don’t know what I’m doing wrong but 4.7 frequently tells me to it’s time to sleep at all hours of the day while working. I’ve tried clearing all my memory/agents files. I’m hoping the “go to sleep” behavior has been rlhf’d away in 4.8.

Oh cool, I thought I had wound up leaving something weird in my CLAUDE.md file that it was always saying Good Night.

Re: Claude Opus 4.8

#967

Give us Mythos! This piecemealing doesn't help Anthropic at all, especially psychologically! They are playing a dangerous game, and I see many people leaving Claude Code for good - both due to the subsidy games, and for Anthropic not dogfooding and using unreleased models internally and giving us subpar ones. Benchmarks are nice, but the real-world experience is quite different - neither can you notice these slight i…

Either Mythos is garbage or it's legitimately better than Opus but requires a subsidy Anthropic can't possibly afford.

They've been going through a funding run, if they had a better product they would've released it to show investors the awesome numbers.

Re: Claude Opus 4.8

#968

Early ArtificialAnalysis.ai results show GPT 5.5 is still the better bang-for-your-buck. OpenAI solves tasks with about 50% less output tokens. https://artificialanalysis.ai/?intelligence=coding-index&int...

This is true but OpenAI has been slowly boiling the frog here, too. $20 and $100 on their plan doesn't get you nearly as far as it did two, three months ago. I use their (newish) 5x $100 plan and I routinely run out of weekly limits about a two days before the end of the week. This has also goaded me into upgrading to $200 once before... and then had them hand out limits resets to everyone. Argh.

same with $200 plan

they silently raised the costs

also feels degraded

Re: Claude Opus 4.8

#969

Early ArtificialAnalysis.ai results show GPT 5.5 is still the better bang-for-your-buck. OpenAI solves tasks with about 50% less output tokens. https://artificialanalysis.ai/?intelligence=coding-index&int...

This is true but OpenAI has been slowly boiling the frog here, too. $20 and $100 on their plan doesn't get you nearly as far as it did two, three months ago. I use their (newish) 5x $100 plan and I routinely run out of weekly limits about a two days before the end of the week. This has also goaded me into upgrading to $200 once before... and then had them hand out limits resets to everyone. Argh.

I feel you. I also switched to $200 plan because I have agents running half of most days.

I was previously coasting on that 2x both Claude and OpenAI were offering.

Re: Claude Opus 4.8

#970

Does anyone troll these releases and cherry pick random metrics other companies would cherry pick to show how amazing their models are? There's like 8 million benchmarks. Every release, every model randomly picks 5-10 where they win in everything except 1, to make it look like they aren't randomly cherry picking benchmarks they probably benchmaxxed for.

Ultimately I think the only way you can trust benchmarks is if you build them yourself and keep them secret from the AI labs. There are different levels of "cheating" on benchmarks. The worst would be just literally putting them in the loss function during RL, I assume the major labs are not cheating at that level. And I am sure they are making a genuine effort to keep the benchmark content out of the training data.…

> Ultimately I think the only way you can trust benchmarks is if you build them yourself and keep them secret from the AI labs.

I agree.

At the same time, one of the first things we see in the HN comments when a new model is released are pelicans on a bike. Makes you wonder where the priorities of the AI "community" lie when karma farming is the main motivation for model "evaluation".

Post reply on HN