Live data from Hacker News

Claude Code daily benchmarks for degradation tracking

marginlab.ai

361–370 of 372 posts

Re: Claude Code daily benchmarks for degradation tracking

#361
post #315

Earlier quoted context omitted.

Anthropic just reduced the price of the team plan and refunded us on the prior invoice. YMMV

So they have no durable principles for deciding who or what to refund… doesnt that make them look even worse…?

Or they do, and two sentences from two different experiences don't tell a full story?

Re: Claude Code daily benchmarks for degradation tracking

#362

Earlier quoted context omitted.

I don't think so. There are other knobs they can tweak to reduce load that affect quality less than quantizing. Like trimming the conversation length without telling you, reducing reasoning effort, etc.

We never do anything that reduce model intelligence like that

You said "like that", ok but there may be some truth to reduced model intelligence. Also how AWS deployed Anthropic models for Amazons Kiro feel much dumber than those controlled entirely by Anthropic. Can't be just me

Re: Claude Code daily benchmarks for degradation tracking

#363

Earlier quoted context omitted.

If you don't do it that way then resizing the terminal corrupts what's on screen.

Counterpoint: Vim has existed for decades and does not use a bloated React rendering pipeline, and doesn't corrupt everything when it gets resized, and is much more full featured from a UI standpoint than Claude Code which is a textbox, and hits 60fps without breaking a sweat unlike Claude Code which drops frames constantly when typing small amounts of text.

Yes, I'm sure it's possible to do better with customized C, but vim took a lot longer to write. And again, fullscreen apps aren't the same as what Claude Code is doing, which is erasing and re-rendering much more than a single screenful of text.

Re: Claude Code daily benchmarks for degradation tracking

#364

Earlier quoted context omitted.

So they have no durable principles for deciding who or what to refund… doesnt that make them look even worse…?

Or they do, and two sentences from two different experiences don't tell a full story?

Okay “they do” based on what more compelling evidence?

Its not like the credibility of the two prior HN users are literally zero…

Re: Claude Code daily benchmarks for degradation tracking

#365

Earlier quoted context omitted.

Or they do, and two sentences from two different experiences don't tell a full story?

Okay “they do” based on what more compelling evidence? Its not like the credibility of the two prior HN users are literally zero…

I am saying there is no evidence either way: they had contrasting experiences and one GP established this means that company has no standardized policies. Maybe they do, maybe they don't — I don't think we can definitively conclude anything.

Re: Claude Code daily benchmarks for degradation tracking

#366

Earlier quoted context omitted.

Okay “they do” based on what more compelling evidence? Its not like the credibility of the two prior HN users are literally zero…

I am saying there is no evidence either way: they had contrasting experiences and one GP established this means that company has no standardized policies. Maybe they do, maybe they don't — I don't think we can definitively conclude anything.

So if you acknowledge the prior claims have more than literally zero credibility… then what’s the issue?

That I dont equally weigh them with all possible yet-to-be claimed things?

Re: Claude Code daily benchmarks for degradation tracking

#368
post #349
post #338

Earlier quoted context omitted.

Why not, can you expand? Asking because I’m considering Claude due to the sandbox feature.

FYI the sandbox feature is not fully baked and does not seem to be high priority. For example, for the last 3 weeks using the sandbox on Linux will almost-always litter your repo root with a bunch of write-protected trash files[0] - there are 2 PRs open to fix it, but Anthropic employees have so far entirely ignored both the issue and the PRs. Very frustrating, since models sometimes accidentally commit those files,…

Hmm, very good point indeed. So far it’s behaved, but I also admit I wasn’t crazy about the outputs it gave me. We’ll see, Anthropic should probably think about their reputation if these issues are common enough.

Re: Claude Code daily benchmarks for degradation tracking

#369

Earlier quoted context omitted.

I am saying there is no evidence either way: they had contrasting experiences and one GP established this means that company has no standardized policies. Maybe they do, maybe they don't — I don't think we can definitively conclude anything.

So if you acknowledge the prior claims have more than literally zero credibility… then what’s the issue? That I dont equally weigh them with all possible yet-to-be claimed things?

I object to your conclusion that "they have no durable principles": not sure how do you get to that from two different experiences documented with a single paragraph.

Re: Claude Code daily benchmarks for degradation tracking

#370

Earlier quoted context omitted.

So if you acknowledge the prior claims have more than literally zero credibility… then what’s the issue? That I dont equally weigh them with all possible yet-to-be claimed things?

I object to your conclusion that "they have no durable principles": not sure how do you get to that from two different experiences documented with a single paragraph.

Because I can assess things via probability… without needing 100% certain proof either way?
Post reply on HN