Live data from Hacker News

Claude Opus 4.8

anthropic.com

721–730 of 1001 posts

Re: Claude Opus 4.8

#721

This is the first time I saw a model pop-up on HN and didn't really care. Model exhaustion? It looks interesting but not exciting. While I'd normally _love_ incremental improvements --- I think the recent ones are far too minor to get excited about or change up a workflow. Besides, benchmarks tend to exaggerate the gap between versions. At this point I'd almost rather Anthropic wait and really wow us with a 5.0 relea…

I have model fatigue

I have… non-deterministic black box that seemingly requires me to re-work myself to get decent results every 4 weeks fatigue

Re: Claude Opus 4.8

#722
LGTM. With "ultra" effort Opus 4.8 was able to reproduce and fix a rare bug in our reactive dataflow that has been haunting me for 4 months. I've had >10 attempts to reproduce and fix with Opus 4.7. What made it hard was that it randomly occurred in only a subset of CI runners and never occurred with local testing across multiple machines. It was a real concurrency bug in the core dataflow.

Re: Claude Opus 4.8

#723

The table comparing eval scores shows the following: Agentic Terminal Coding (Terminal-Bench 2.1) Opus 4.8 74.6% GPT 5.5 78.2% Then, when you scroll all the way down to the bottom Footnotes section it says "Terminal-Bench 2.1: We reported scores for all models using the Terminus-2 public harness. GPT-5.5’s reported score with the Codex CLI harness is 83.4%."

Seems reasonable? Presumably Claude also performs better under the Claude Code harness.

Why not state that?

Re: Claude Opus 4.8

#724

Earlier quoted context omitted.

> My own experience w/ 4.6 and 4.7 are that I don't firmly grasp any capabilities improvements over my memory of 4.5, but it's all so fuzzy that it's truly difficult to tell. I've actually intentionally switched back to 4.5. I hated 4.7 so much that I decided to jump back all the way to 4.5. Now that I've been using 4.5 for a few weeks, I find it significantly more reliable but a bit more forgetful than 4.6/4.7. I'm…

Same here. Went back to 4.5 and was happy I did it. The only frustration was that I can tell the model has declined compared to the first few weeks it was released. I also recently moved to 4.6 since I started hitting the context limit too often with my current project.

/model claude-opus-4-6[1m]

allows you to specify you want the 1 million context 4.6

Re: Claude Opus 4.8

#725
post #489

I believe analogy with smartphone will be best for this case. In 2010s iphone was the king, all those Chinese devices ware cheaper but not even close to smoothnest and usability of US tech, now after 15 years later everything is changed, now iphone feels like old grandpa to Chinese tech. Same will happend to LLM's just much faster.

EVs too

Re: Claude Opus 4.8

#726

Opus 4.7 was acting extremely stupid today. Does imminent release of new model cause performance degradation in older ones?

Technology is amazing! We’ve managed to make software that has brain fart days and morale problems!

Re: Claude Opus 4.8

#727
I was happily plodding away with it earlier when it threw this out in the middle of a response in Claude code:

--- So — what did you actually see before you hit Ctrl-C? That's the信号 I'm most curious about, and it tells us what to ---

That's the sort of behavior I'd expect from a one or two year old model quantized down to about 1 bit - right word, wrong language in a response. Google translate tells me that's Chinese for signal. I wonder what caused that to happen.

Re: Claude Opus 4.8

#728
post #643

Earlier quoted context omitted.

I won't be surprised if the next gen frontier models are the last. There's orders of magnitude of low hanging juice to squeeze out of smaller models. It is almost guaranteed that a 60-90B model can outperform current SOTA in coding tasks within 2-3 years (design not certain, probably unlikely). It is far less clear that a 1.2T model will be meaningfully better enough to justify training it. As far as reasoning is con…

There is endless returns to frontier intelligence, just because most people can't make use of it doesn't mean someone can't make a ton of money off of it. Most software engineers will just need cheap tokens. But things like physics and drug discovery have no foreseeable upper bound.

Or governance of large organizations... There are a huge number of factors to consider, counterfactuals, studies, lots of non-obvious second and third order effects, etc. We're barely able to get basic governance without creating huge problems (low density zoning rubber stamped across the nation creating a housing crisis, for example), so the bar isn't high.

We pay CEOs an enormous amount because a small improvement in performance of an org because of them can make a massive difference in organizational value.

Re: Claude Opus 4.8

#729
post #698
post #73

I generated pelicans riding bicycles on both thinking level low and thinking level high: https://gist.github.com/simonw/68560eddb0b268a8417f80ceb7304... The high one is notably better - the bicycle frame is the correct shape, unlike thinking level low. For comparison, here's Opus 4.7: https://gist.github.com/simonw/afcb19addf3f38eb1996e1ebe749c...

It's pretty safe to say that AI will be used on the battlefield making real life and death decisions before it will be able to render a decent pelican on a bike in SVG.

the battlefield sounds much easier. worst case scenario you kill somebody, but that's what you're trying to do anyways.

if you kill somebody while trying to render a pelican on a bicycle it's a real problem.

Re: Claude Opus 4.8

#730
post #14

I can't help but think of Iphone updates since about 2018. The thinnest, fastest, longest battery life Iphone ever. It seems mostly the same and I probably won't be able to tell other than the name, but everyone buys it anyway. This is good psychology for the labs. When Buffett invested in Apple he loved citing how most people would rather give up their second car than their Iphone.

If you upgrade your 8 year old phone the many incremental upgrades will be very noticeable. From my personal experience the LLM space is also moving at a faster pace than the phone industry at the moment, but at least from a financial perspective I would expect it to slow down sooner rather than later.
Post reply on HN