Live data from Hacker News

Claude Code daily benchmarks for degradation tracking

marginlab.ai

1–10 of 372 posts

Re: Claude Code daily benchmarks for degradation tracking

#3
Very interesting. I would be curious to understand how granular these updates are being applied to CC + what might be causing things like this. I feel like I can notice a very small degradation but have compensated with more detailed prompts (which I think, perhaps naively, is offsetting this issue).

Re: Claude Code daily benchmarks for degradation tracking

#5
This is probably entirely down to subtle changes to CC prompts/tools.

I've been using CC more or less 8 hrs/day for the past 2 weeks, and if anything it feels like CC is getting better and better at actual tasks.

Edit: Before you downvote, can you explain how the model could degrade WITHOUT changes to the prompts? Is your hypothesis that Opus 4.5, a huge static model, is somehow changing? Master system prompt changing? Safety filters changing?

Re: Claude Code daily benchmarks for degradation tracking

#6

I really like the idea, but a "±14.0% significance threshold" is meaningless here. The larger monthly scale should be the default, or you should get more samples.

Could you elaborate what you think the problems are? I guess they should be using some form of multiple comparison correction?

Re: Claude Code daily benchmarks for degradation tracking

#7
post #5

This is probably entirely down to subtle changes to CC prompts/tools. I've been using CC more or less 8 hrs/day for the past 2 weeks, and if anything it feels like CC is getting better and better at actual tasks. Edit: Before you downvote, can you explain how the model could degrade WITHOUT changes to the prompts? Is your hypothesis that Opus 4.5, a huge static model, is somehow changing? Master system prompt changin…

I was going to ask, are all other variables accounted for? Are we really comparing apples to apples here? Still worth doing obviously, as it serves a good e2e evaluations, just for curiosity's sake.

Re: Claude Code daily benchmarks for degradation tracking

#9
post #5

This is probably entirely down to subtle changes to CC prompts/tools. I've been using CC more or less 8 hrs/day for the past 2 weeks, and if anything it feels like CC is getting better and better at actual tasks. Edit: Before you downvote, can you explain how the model could degrade WITHOUT changes to the prompts? Is your hypothesis that Opus 4.5, a huge static model, is somehow changing? Master system prompt changin…

Honest, good-faith question.

Is CC getting better, or are you getting better at using it? And how do you know the difference?

I'm an occasional user, and I can definitely see improvements in my prompts over the past couple of months.

Re: Claude Code daily benchmarks for degradation tracking

#10
post #5

This is probably entirely down to subtle changes to CC prompts/tools. I've been using CC more or less 8 hrs/day for the past 2 weeks, and if anything it feels like CC is getting better and better at actual tasks. Edit: Before you downvote, can you explain how the model could degrade WITHOUT changes to the prompts? Is your hypothesis that Opus 4.5, a huge static model, is somehow changing? Master system prompt changin…

That's why benchmarks are useful. We all suffer from the shortcomings of human perception.
Post reply on HN