Live data from Hacker News

Claude Code daily benchmarks for degradation tracking

marginlab.ai

91–100 of 372 posts

Re: Claude Code daily benchmarks for degradation tracking

#91
post #89

Earlier quoted context omitted.

Or maybe, juste maybe, that's when they started testing…

Wayback machine has nothing for this site before today, and article is "last updated Jan 29". A benchmark like this ought to start fresh from when it is published. I don't entirely doubt the degradation, but the choice of where they went back to feels a bit cherry-picked to demonstrate the value of the benchmark.

Which makes sense, you gotta wait until you get enough data before you can communicate on the said data…

If anything it's coherent with the fact that they very likely didn't have data earlier than January the 8th.

Re: Claude Code daily benchmarks for degradation tracking

#92
post #60

Earlier quoted context omitted.

noob question: why would increased demand result in decreased intelligence?

It would happen if they quietly decide to serve up more aggressively distilled / quantised / smaller models when under load.

They advertise the Opus 4.5 model. Secretly substituting a cheaper one to save costs would be fraud.

Re: Claude Code daily benchmarks for degradation tracking

#94
post #79
post #66

Earlier quoted context omitted.

An operator at load capacity can either refuse requests, or move the knobs (quantization, thinking time) so requests process faster. Both of those things make customers unhappy, but only one is obvious.

This is intentional? I think delivering lower quality than what was advertised and benchmarked is borderline fraud, but YMMV.

[flagged]

Re: Claude Code daily benchmarks for degradation tracking

#95
This strategy seems inspired by TikTok's approach for retaining new uploaders.

TikTok used to give new uploaders a visibility boost (i.e., an inflated number of likes and comments) on their first couple of uploads, to get them hooked on the the service.

In Anthropic/Claude's case, the strategy is (allegedly) to give new users access to the premium models on sign-up, and then increasingly cut the product with output from cheaper models.

Re: Claude Code daily benchmarks for degradation tracking

#96

My personal conspiracy theory is that they choose who to serve a degraded model to based on social graph analysis and sentiment analysis, maximizing for persuasion while minimizing compute.

Sounds more like a sound business plan than a conspiracy theory.

It sounds like fraud to me

Re: Claude Code daily benchmarks for degradation tracking

#97

My personal conspiracy theory is that they choose who to serve a degraded model to based on social graph analysis and sentiment analysis, maximizing for persuasion while minimizing compute.

IMO this strategy seems inspired by TikTok's approach for retaining new uploaders.

TikTok used to give new uploaders a visibility boost (i.e., an inflated number of likes and comments) on their first couple of uploads, to get them hooked on the the service.

In Anthropic/Claude's case, the strategy is (allegedly) to give new users access to the premium models on sign-up, and then increasingly cut the product with output from cheaper models.

Of course, your suggestion (better service for users who know how to speak Proper English) would be the cherry on top of this strategy.

From what I've seen on HackerNews, Anthropic is all-in on social media manipulation and social engineering, so I suspect that your assumption holds water.

Re: Claude Code daily benchmarks for degradation tracking

#98
post #88
post #79

Earlier quoted context omitted.

This is intentional? I think delivering lower quality than what was advertised and benchmarked is borderline fraud, but YMMV.

Personally, I'd rather get queued up on a long wait time I mean not ridiculously long but I am ok waiting five minutes to get correct it at least more correct responses. Sure, I'll take a cup of coffee while I wait (:

i’d wait any amount of time lol.

at least i would KNOW it’s overloaded and i should use a different model, try again later, or just skip AI assistance for the task altogether.

Re: Claude Code daily benchmarks for degradation tracking

#99
post #79
post #66

Earlier quoted context omitted.

An operator at load capacity can either refuse requests, or move the knobs (quantization, thinking time) so requests process faster. Both of those things make customers unhappy, but only one is obvious.

This is intentional? I think delivering lower quality than what was advertised and benchmarked is borderline fraud, but YMMV.

Per Anthropic’s RCA linked in Ops post for September 2025 issues:

“… To state it plainly: We never reduce model quality due to demand, time of day, or server load. …”

So according to Anthropic they are not tweaking quality setting due to demand.

Re: Claude Code daily benchmarks for degradation tracking

#100
post #29

Why I do not believe this shows Anthropic serves folks a worse model: 1. The percentage drop is too low and oscillating, it goes up and down. 2. The baseline of Sonnet 4.5 (the obvious choice for when they have GPU busy for the next training) should be established to see Opus at some point goes Sonnet level. This was not done but likely we would see a much sharp decline in certain days / periods. The graph would look…

I believe the science, but I've been using it daily and it's been getting worse, noticeably.
Post reply on HN