Live data from Hacker News

Claude Code daily benchmarks for degradation tracking

marginlab.ai

81–90 of 372 posts

Re: Claude Code daily benchmarks for degradation tracking

#81
post #76
post #29

Why I do not believe this shows Anthropic serves folks a worse model: 1. The percentage drop is too low and oscillating, it goes up and down. 2. The baseline of Sonnet 4.5 (the obvious choice for when they have GPU busy for the next training) should be established to see Opus at some point goes Sonnet level. This was not done but likely we would see a much sharp decline in certain days / periods. The graph would look…

4. The graph starts January 8. Why January 8? Was that an outlier high point? IIRC, Opus 4.5 was released late november.

Or maybe, juste maybe, that's when they started testing…

Re: Claude Code daily benchmarks for degradation tracking

#82

[SWE-bench co-author here] It seems like they run this test on a subset of 50 tasks, and that they only run the test once per day. So a lot of the movement in accuracy could be attributed to that. I would run on 300 tasks and I'd run the test suite 5 or 10 times per day and average that score. Lots of variance in the score can come from random stuff like even Anthropic's servers being overloaded.

> Lots of variance in the score can come from random stuff like even Anthropic's servers being overloaded.

Are you suggesting result accuracy varies with server load?

Re: Claude Code daily benchmarks for degradation tracking

#83
post #29

Why I do not believe this shows Anthropic serves folks a worse model: 1. The percentage drop is too low and oscillating, it goes up and down. 2. The baseline of Sonnet 4.5 (the obvious choice for when they have GPU busy for the next training) should be established to see Opus at some point goes Sonnet level. This was not done but likely we would see a much sharp decline in certain days / periods. The graph would look…

> 1. The percentage drop is too low and oscillating, it goes up and down.

How do you define “too low”, they make sure to communicate about the statistical significance of their measurements, what's the point if people can just claim it's “too low” based on personal vibes…

Re: Claude Code daily benchmarks for degradation tracking

#85
post #79
post #66

Earlier quoted context omitted.

An operator at load capacity can either refuse requests, or move the knobs (quantization, thinking time) so requests process faster. Both of those things make customers unhappy, but only one is obvious.

This is intentional? I think delivering lower quality than what was advertised and benchmarked is borderline fraud, but YMMV.

There is no level of quality advertised, as far as I can see.

Re: Claude Code daily benchmarks for degradation tracking

#86
post #12

Simply search user prompts for curse words and then measure hostility sentiment. User hostility rises as agents fail to meet expectations.

Or there are global events that stress people out .. or their expectations change over time. Not that simple ;)

Re: Claude Code daily benchmarks for degradation tracking

#87
post #79
post #66

Earlier quoted context omitted.

An operator at load capacity can either refuse requests, or move the knobs (quantization, thinking time) so requests process faster. Both of those things make customers unhappy, but only one is obvious.

This is intentional? I think delivering lower quality than what was advertised and benchmarked is borderline fraud, but YMMV.

If there's no way to check, then how can you claim it's fraud? :)

Re: Claude Code daily benchmarks for degradation tracking

#88
post #79
post #66

Earlier quoted context omitted.

An operator at load capacity can either refuse requests, or move the knobs (quantization, thinking time) so requests process faster. Both of those things make customers unhappy, but only one is obvious.

This is intentional? I think delivering lower quality than what was advertised and benchmarked is borderline fraud, but YMMV.

Personally, I'd rather get queued up on a long wait time I mean not ridiculously long but I am ok waiting five minutes to get correct it at least more correct responses.

Sure, I'll take a cup of coffee while I wait (:

Re: Claude Code daily benchmarks for degradation tracking

#89
post #76

Earlier quoted context omitted.

4. The graph starts January 8. Why January 8? Was that an outlier high point? IIRC, Opus 4.5 was released late november.

Or maybe, juste maybe, that's when they started testing…

Wayback machine has nothing for this site before today, and article is "last updated Jan 29".

A benchmark like this ought to start fresh from when it is published.

I don't entirely doubt the degradation, but the choice of where they went back to feels a bit cherry-picked to demonstrate the value of the benchmark.

Re: Claude Code daily benchmarks for degradation tracking

#90
post #79
post #66

Earlier quoted context omitted.

An operator at load capacity can either refuse requests, or move the knobs (quantization, thinking time) so requests process faster. Both of those things make customers unhappy, but only one is obvious.

This is intentional? I think delivering lower quality than what was advertised and benchmarked is borderline fraud, but YMMV.

> I think delivering lower quality than what was advertised and benchmarked is borderline fraud

welcome to the Silicon Valley, I guess. everything from Google Search to Uber is fraud. Uber is a classic example of this playbook, even.

Post reply on HN