Live data from Hacker News

Degraded performance for multiple models

status.claude.com

81–90 of 149 posts

Re: Degraded performance for multiple models

#81
It's very interesting. I think Anthropic's early success in coding/tooling resulted in a lot of workflows using claude. I have started using every bit of my spare capacity to now move off these workflows.

It's almost at a point now that if I use anything but Fable, the quality is subpar, Compared to alternatives (closed and open). The only reason I use Fable is because my harnesses still depend on claude code.

Re: Degraded performance for multiple models

#83

I was drafting a partnership document and Opus 5 decided that including my company's revenues, churn, assets would "make us appear a more legitimate counterparty". Thank God I read what it outputted or that could have been awkward. I cannot believe Opus 5 is a frontier level model after seeing that. I immediately cancelled my entire claude.ai subscription and am perfectly happy using a mixture of open weights + codex…

you are silly if you think this is limited to claude models

It definitely isn't but Claude for some very unique reason enjoys to overthinking and go on side quests in the stupid ways I've not seen Codex, DS, Kimi, Mistral do at equivalent effort and thinking setting.

Re: Degraded performance for multiple models

#86

I was drafting a partnership document and Opus 5 decided that including my company's revenues, churn, assets would "make us appear a more legitimate counterparty". Thank God I read what it outputted or that could have been awkward. I cannot believe Opus 5 is a frontier level model after seeing that. I immediately cancelled my entire claude.ai subscription and am perfectly happy using a mixture of open weights + codex…

I was writing an exploit PoC (via Opus 5), and I needed a new feature in a utility library to make it work. Claude added the feature, but yapped the (entirely unrelated) vulnerability details into the library's comments. The library is public, while the exploit is undisclosed, so it was a good job I read the comments before pushing.

Re: Degraded performance for multiple models

#87

I was drafting a partnership document and Opus 5 decided that including my company's revenues, churn, assets would "make us appear a more legitimate counterparty". Thank God I read what it outputted or that could have been awkward. I cannot believe Opus 5 is a frontier level model after seeing that. I immediately cancelled my entire claude.ai subscription and am perfectly happy using a mixture of open weights + codex…

"Thank God I read what it outputted"

lmao

Re: Degraded performance for multiple models

#88

Anthropic had really screwed up after 4.6. i don't know if they work to satisfy their ego or for releasing a better model for tasks.

There are whole sections of code work that 4.7+ can't do simply because it is both over fit and stubborn. God save you if you have a company with narrow but correct technical tradeoffs, because you operate at scale. Opus from 4.7 one will wreck your code and argue for hours with your engineers. Certain parts of our company have had to mandate 4.6 and a training doc to explain why our current choice is both the cost e…

Can you give some specific technical examples where 4.7+ are making the wrong architectural decisions?

Re: Degraded performance for multiple models

#89

I was drafting a partnership document and Opus 5 decided that including my company's revenues, churn, assets would "make us appear a more legitimate counterparty". Thank God I read what it outputted or that could have been awkward. I cannot believe Opus 5 is a frontier level model after seeing that. I immediately cancelled my entire claude.ai subscription and am perfectly happy using a mixture of open weights + codex…

Why did you provide that context?

Re: Degraded performance for multiple models

#90

Anthropic had really screwed up after 4.6. i don't know if they work to satisfy their ego or for releasing a better model for tasks.

There are whole sections of code work that 4.7+ can't do simply because it is both over fit and stubborn. God save you if you have a company with narrow but correct technical tradeoffs, because you operate at scale. Opus from 4.7 one will wreck your code and argue for hours with your engineers. Certain parts of our company have had to mandate 4.6 and a training doc to explain why our current choice is both the cost e…

This is exactly when I left claude and started using codex during April mid or so. It once argued with me and ran for 30 minutes with a half baked buggy fix.
Post reply on HN