Claude Code daily benchmarks for degradation tracking
11–20 of 372 posts
Re: Claude Code daily benchmarks for degradation tracking
#12Re: Claude Code daily benchmarks for degradation tracking
#13I really like the idea, but a "±14.0% significance threshold" is meaningless here. The larger monthly scale should be the default, or you should get more samples.
Could you elaborate what you think the problems are? I guess they should be using some form of multiple comparison correction?
Re: Claude Code daily benchmarks for degradation tracking
#14Would love to see this idea expanded to ever alleged SoTA model currently in production. Any speculation as to why this degradation occurs?
Re: Claude Code daily benchmarks for degradation tracking
#15Simply search user prompts for curse words and then measure hostility sentiment. User hostility rises as agents fail to meet expectations.
Re: Claude Code daily benchmarks for degradation tracking
#16Why is this happening?
Re: Claude Code daily benchmarks for degradation tracking
#17This is probably entirely down to subtle changes to CC prompts/tools. I've been using CC more or less 8 hrs/day for the past 2 weeks, and if anything it feels like CC is getting better and better at actual tasks. Edit: Before you downvote, can you explain how the model could degrade WITHOUT changes to the prompts? Is your hypothesis that Opus 4.5, a huge static model, is somehow changing? Master system prompt changin…
That's why benchmarks are useful. We all suffer from the shortcomings of human perception.
Re: Claude Code daily benchmarks for degradation tracking
#18Why is this happening?
Re: Claude Code daily benchmarks for degradation tracking
#19It was probably 3x faster than usual. I got more done in the next hour with it than I do in half a day usually. It was definitely a bit of a glimpse into a potential future of “what if these things weren’t resource constrained and could just fly”.
Re: Claude Code daily benchmarks for degradation tracking
#20Simply search user prompts for curse words and then measure hostility sentiment. User hostility rises as agents fail to meet expectations.