Live data from Hacker News

Claude Opus 4.5

anthropic.com

471–480 of 525 posts

Re: Claude Opus 4.5

#471
Opus 4.5's scaling is impressive on benchmarks, but the usual caveats apply: benchmark saturation is real, and we're seeing diminishing returns on evals that test pattern-matching vs. genuine reasoning. The more relevant question: has anyone stress-tested this on novel problems or complex multi-step reasoning outside training data distributions? Marketing often showcases 'advanced math' and 'code generation' where the solutions exist in training data. The claim of 'reasoning improvement' needs validation on genuinely unfamiliar problem classes.

Re: Claude Opus 4.5

#472
post #469

Interesting that the number of hn comments on big model announcements seems to be dropping. I recall previous ones easily surpassing 1k Maybe models are starting to get good enough/ levelling off?

It's fatigue. This is the third major model announcement in the last week.

On the other hand, this is the one I'm most excited by. I wouldn't have commented at all if it wasn't for your comment. But I'm excited to start using this.

Re: Claude Opus 4.5

#473
After experimenting with Gemini 3, I still felt like Sonnet 4.5 had the edge. So I'm very excited to start playing with this in the wild.

Re: Claude Opus 4.5

#474
post #394
post #227

Earlier quoted context omitted.

> Is this empirical evidence? Look, I'm not defending the big labs, I think they're terrible in a lot of ways. And I'm actually suspending judgement on whether there is ~some kind of nerf happening. But the anecdote you're describing is the definition of non-empirical. It is entirely subjective, based entirely on your experience and personal assessment.

It's not non-empirical. He was careful to give it the same experiment twice. The dependent variable is his judgment, sure, but why shouldn't we trust that if he's an experienced SWE?

Sample size is way too small.

Unless he was able to sample with temperature 0 (and get fully deterministic results both times), this can just be random chance. And experience as SWE doesn't imply experience with statistics and experiment design.

Re: Claude Opus 4.5

#475

Earlier quoted context omitted.

I added Opus 4.5 to my benchmark of 30 alternatives to your now-classic pelican-bicycle prompt (e.g., “Generate an SVG of a dragonfly balancing a chandelier”). Nine models are now represented: https://gally.net/temp/20251107pelican-alternatives/index.ht...

Gemini 3.0 Pro Preview is incredible compared to the others, at least for SVGs.

I was about to say the same; suspiciously good, even. Feels like it's either memorised a bunch of SVG files, or has a search tool and is finding complete items off the web to include either in whole or in part.

Given that it also sometimes goes weird, I suspect it's more likely to be the former.

While the latter would be technically impressive, it's also the whole "this is just collage!" criticism that diffusion image generators faced from people that didn't understand diffusion image generators.

Re: Claude Opus 4.5

#476

This is gonna be game-changing for the next 2-4 weeks before they nerf the model. Then for the next 2-3 months people complaining about the degradation will be labeled “skill issue”. Then a sacrificial Anthropic engineer will “discover” a couple obscure bugs that “in some cases” might have lead to less than optimal performance. Still largely a user skill issue though. Then a couple months later they’ll release Opus 4…

I’m disappointed that this type of discourse has now entered HN. I expected a more evidence-based less “nerf cycle” discussion over here.

Re: Claude Opus 4.5

#478

Earlier quoted context omitted.

> There are well documented cases of performance degradation: https://www.anthropic.com/engineering/a-postmortem-of-three-... There was one well-documented case of performance degradation which arose from a stupid bug, not some secret cost cutting measure.

I never claimed that it was being done in secrecy. Here is another example: https://groq.com/blog/inside-the-lpu-deconstructing-groq-spe... . I have seen multiple people mention openrouter multiple times here on HN: https://hn.algolia.com/?dateRange=all&page=0&prefix=true&que... Again, I'm not claiming malicious intent. But model performance depends on a number of factors and the end-user just sees benchmarks for a s…

All those are completely irrelevant. Quantization is just a cost optimization.

People are claiming that Anthropic et all changes the quality of the model after the initial release, which is entirely different and the industry as a whole has denied. When a model is released under a certain version, the model doesn’t change.

The only people who believe this are in the vibe coding community, believing that there’s some kind of big conspiracy, but any time you mention “but benchmarks show the performance stays consistent” you’re told you’re licking corporate ass.

Re: Claude Opus 4.5

#479
post #469

Interesting that the number of hn comments on big model announcements seems to be dropping. I recall previous ones easily surpassing 1k Maybe models are starting to get good enough/ levelling off?

"we gained 2.7% in these artificial benchmarks and here is a picture of a pelican on a bicycle, get excited and give us $7 trillion please"

Re: Claude Opus 4.5

#480

Earlier quoted context omitted.

Thanks. Still looking for some kind of total code by phone thing.

take a look at https://apps.apple.com/us/app/bitrig/id6747835910

Can you run the apps without going through Apple? Do you need a developer account?
Post reply on HN