Live data from Hacker News

Claude Code daily benchmarks for degradation tracking

marginlab.ai

11–20 of 372 posts

Re: Claude Code daily benchmarks for degradation tracking

#13
post #6

I really like the idea, but a "±14.0% significance threshold" is meaningless here. The larger monthly scale should be the default, or you should get more samples.

Could you elaborate what you think the problems are? I guess they should be using some form of multiple comparison correction?

The daily scale is not statistically significant and is meaningless. You should lower the confidence interval by either increasing the scale or the evaluations.

Re: Claude Code daily benchmarks for degradation tracking

#14

Would love to see this idea expanded to ever alleged SoTA model currently in production. Any speculation as to why this degradation occurs?

Anecdote, I don't have any proof and it's just a feeling. But around afternoon in GMT+1 compared to the morning/midday, there seems to be a change in the quality of responses, which seems to line up with when the US wakes up. I consistently get (what feels like) worse responses in both Codex and Claude Code in the afternoon/night compared to morning/midday, so much that I usually give up then try the same prompt next morning and get better results. But I guess that might as well be about me being more tired in the night than morning too, as I said, haven't measured this.

Re: Claude Code daily benchmarks for degradation tracking

#17
post #10
post #5

This is probably entirely down to subtle changes to CC prompts/tools. I've been using CC more or less 8 hrs/day for the past 2 weeks, and if anything it feels like CC is getting better and better at actual tasks. Edit: Before you downvote, can you explain how the model could degrade WITHOUT changes to the prompts? Is your hypothesis that Opus 4.5, a huge static model, is somehow changing? Master system prompt changin…

That's why benchmarks are useful. We all suffer from the shortcomings of human perception.

Benchmarks shortcomings are no worse... they inevitably measure something that is only close to the thing you actually care about, not the thing you actually care about. It's entirely plausible that this decreased benchmark score is because Anthropic's initial prompting of the model was overtuned to the benchmark and as they're gaining more experience with real world use they are changing the prompt to do better at that and consequentially worse at the benchmark.

Re: Claude Code daily benchmarks for degradation tracking

#18

Why is this happening?

I have absolutely no insight knowledge, but I think it's not a bad assumption to have that, it's costly to run the models, when they release a new model they assume that cost and give per user more raw power, when they've captured the new users and wow factor, they start reducing costs by reducing the capacity they provide to users. Rinse and repeat.

Re: Claude Code daily benchmarks for degradation tracking

#19
There was a moment about a week ago where Claude went down for about an hour. And right after it came back up it was clear a lot of people had given up and were not using it.

It was probably 3x faster than usual. I got more done in the next hour with it than I do in half a day usually. It was definitely a bit of a glimpse into a potential future of “what if these things weren’t resource constrained and could just fly”.

Re: Claude Code daily benchmarks for degradation tracking

#20
post #12

Simply search user prompts for curse words and then measure hostility sentiment. User hostility rises as agents fail to meet expectations.

I uh might be skewing that as I generally just use a lot of curse words with Claude by default
Post reply on HN