Live data from Hacker News

Claude Code daily benchmarks for degradation tracking

marginlab.ai

371–372 of 372 posts

Re: Claude Code daily benchmarks for degradation tracking

#371

Earlier quoted context omitted.

I object to your conclusion that "they have no durable principles": not sure how do you get to that from two different experiences documented with a single paragraph.

Because I can assess things via probability… without needing 100% certain proof either way?

This is becoming futile: this is not even about proof, but there not even being a full account of two cases you are basing your opinion on.

Obviously, you can derive any opinion you want out of that, but while I am used to terms like "probability" being misused like this, I've generally seen a higher standard at HN.

To each their own, though. Thank you for the discourse and have a good day.

Re: Claude Code daily benchmarks for degradation tracking

#372

Earlier quoted context omitted.

Catastrophic error accumulation can produce more profound effects than noise.

Just to make sure I got this right. They serve millions of requests a day & somehow catastrophic error accumulation is what is causing the 10% degradation & no one at Anthropic is noticing it. Is that the theory?

FYI something in that region happened last august/September. Some inference bug triggered worse performance on TPUs vs GPU.
Post reply on HN