Live data from Hacker News

Degraded performance for multiple models

status.claude.com

111–120 of 149 posts

Re: Degraded performance for multiple models

#111
post #93
post #50

What's the incentive to keep on improving the model beyond a point? 10 devs on a team will be cut to 2 devs, so that's 8 licenses lost. They have to increase the price many fold.

they unironically think that they can replace everyone in an organization

What I don't understand is, why not replace middle management, marketing, CTOs, CEOs and the like. Surely, LLMs are better at producing high quality looking slideware and vaporware than they are at producing software.

Heavy sarcasm here if it's not obvious. Of course I know why.

Re: Degraded performance for multiple models

#113
post #99
post #87

Earlier quoted context omitted.

"Thank God I read what it outputted" lmao

"Thank God i checked to see if the gun was loaded before i pointed it in a random direction and pulled the trigger". sheesh, my assumptions of general human intelligence continues to be wrong.

the last few years of reading hackernews has been a wild ride. I used to come here for well thought out articles and opinions. IDK if all the "older guard" have left the building or if we've all been consumed by abject stupidity.

Re: Degraded performance for multiple models

#114
post #81

It's very interesting. I think Anthropic's early success in coding/tooling resulted in a lot of workflows using claude. I have started using every bit of my spare capacity to now move off these workflows. It's almost at a point now that if I use anything but Fable, the quality is subpar, Compared to alternatives (closed and open). The only reason I use Fable is because my harnesses still depend on claude code.

You can use Claude Code directly with any provider that supports Anthropic's API by setting some environment variables, and indirectly via a proxy with pretty much anything else.

Re: Degraded performance for multiple models

#116
post #105

Is anybody else experiencing this: I have multiple claude instances running on different servers - and some keep getting the 529 Overloaded error and one instance doesn't and just continues working. All are using Opus 5.

I have the same thing on different shells on the same server (tmux). It's wild. I'm guessing that some of the sessions just got lucky as to which backend server they are redirected to.

Re: Degraded performance for multiple models

#117

I was drafting a partnership document and Opus 5 decided that including my company's revenues, churn, assets would "make us appear a more legitimate counterparty". Thank God I read what it outputted or that could have been awkward. I cannot believe Opus 5 is a frontier level model after seeing that. I immediately cancelled my entire claude.ai subscription and am perfectly happy using a mixture of open weights + codex…

you are silly if you think this is limited to claude models

it isnt, but there are weird sycophantic behaviors with Opus i dont see anywhere else.

Re: Degraded performance for multiple models

#118
post #81

It's very interesting. I think Anthropic's early success in coding/tooling resulted in a lot of workflows using claude. I have started using every bit of my spare capacity to now move off these workflows. It's almost at a point now that if I use anything but Fable, the quality is subpar, Compared to alternatives (closed and open). The only reason I use Fable is because my harnesses still depend on claude code.

It’s so hard to be sure but opus feels like its been steadily declining since 4.6

For me, Opus 5 has seemed to compete with Fable on quality of output. Prior to Opus 5 though, the previous Opus models did seem to decline once Fable was release. That's just my experience though.

Re: Degraded performance for multiple models

#119
post #79

I was drafting a partnership document and Opus 5 decided that including my company's revenues, churn, assets would "make us appear a more legitimate counterparty". Thank God I read what it outputted or that could have been awkward. I cannot believe Opus 5 is a frontier level model after seeing that. I immediately cancelled my entire claude.ai subscription and am perfectly happy using a mixture of open weights + codex…

I hope you don’t plan to cut back on reading legal documents crafted by any LLM before executing them.

I read everything. I will have AI ingest NDAs to make sure they arent glaringly weird and I then go read them, it gives me a good idea of what to look for.

Re: Degraded performance for multiple models

#120

Earlier quoted context omitted.

Aren't the model weights frozen?

Model competence is an interaction of weights, system prompt, and harness.

Don’t forget reasoning effort. We get labels like “low,” “high,” and “max.” That doesn’t mean that the numbers associated with those don’t get remapped on the backend.
Post reply on HN