Live data from Hacker News

Claude Code daily benchmarks for degradation tracking

marginlab.ai

131–140 of 372 posts

Re: Claude Code daily benchmarks for degradation tracking

#132

There was a moment about a week ago where Claude went down for about an hour. And right after it came back up it was clear a lot of people had given up and were not using it. It was probably 3x faster than usual. I got more done in the next hour with it than I do in half a day usually. It was definitely a bit of a glimpse into a potential future of “what if these things weren’t resource constrained and could just fly…

Noticed the exact same thing a few days ago. So much so that I went on twitter and HN to search for “claude speed boost” to see if there was a known new release. Felt like the time I upgraded from a 2400 baud modem to a 14.4 as a kid - everything was just lightning fast (for a brief shining moment).

Re: Claude Code daily benchmarks for degradation tracking

#133

Earlier quoted context omitted.

Per Anthropic’s RCA linked in Ops post for September 2025 issues: “… To state it plainly: We never reduce model quality due to demand, time of day, or server load. …” So according to Anthropic they are not tweaking quality setting due to demand.

And according to Google, they always delete data if requested. And according to Meta, they always give you ALL the data they have on you when requested.

What would you like?

Re: Claude Code daily benchmarks for degradation tracking

#134
post #70

I have yet to experience any degradation in coding tasks I use to evaluate Opus 4.5, but I did see a rather strange and reproducible worsening in prompt adherence as part of none coding tasks since the third week of January. Very simple queries, even those easily answered via regular web searching, have begun to consistently not result accurate results with Opus 4.5, despite the same prompts previously yielding accur…

I've noticed a degradation in Opus 4.5, also with Gemini-3-Pro. For me, it was a sudden rapid decline in adherence to specs in Claude Code. On an internal benchmark we developed, Gemini-3-Pro also dramatically declined. Going from being clearly beyond every other model (as benchmarks would lead you to believe) to being quite mediocre. Delivering mediocre results in chat queries and coding also missing the mark.

I didn't "try 100 times" so it's unclear if this is an unfortunate series of bad runs on Claude Code and Gemini CLI or actual regression.

I shouldn't have to benchmark this sort of thing but here we are.

Re: Claude Code daily benchmarks for degradation tracking

#135
post #100

Earlier quoted context omitted.

I believe the science, but I've been using it daily and it's been getting worse, noticeably.

Any chance you’re just learning more about what the model is and is not useful for?

I dunno about everyone else but when I learn more about what a model is and is not useful for, my subjective experience improves, not degrades.

Re: Claude Code daily benchmarks for degradation tracking

#136

This strategy seems inspired by TikTok's approach for retaining new uploaders. TikTok used to give new uploaders a visibility boost (i.e., an inflated number of likes and comments) on their first couple of uploads, to get them hooked on the the service. In Anthropic/Claude's case, the strategy is (allegedly) to give new users access to the premium models on sign-up, and then increasingly cut the product with output f…

Yes, but the difference is TikTok didn't sell a particular service version.

Anthropic did sell a particular model version.

Re: Claude Code daily benchmarks for degradation tracking

#137

Very interesting. I would be curious to understand how granular these updates are being applied to CC + what might be causing things like this. I feel like I can notice a very small degradation but have compensated with more detailed prompts (which I think, perhaps naively, is offsetting this issue).

> more detailed prompts (which I think, perhaps naively, is offsetting this issue).

Is exacerbating this issue ... if the load theory is correct.

Re: Claude Code daily benchmarks for degradation tracking

#140
Pretty sure someone at Google, OpenAI, and Anthropic met up at a park, leaving their phones in their car, and had a conversation that January 2026, they were all going to silently degrade their models.

They were fighting an arms race that was getting incredibly expensive and realized they could get away with spending less electricity and there was nothing the general population could do about it.

Grok/Elon was left out of this because he would leak this idea at 3am after a binge.

Post reply on HN