Live data from Hacker News

Degraded performance for multiple models

status.claude.com

71–80 of 149 posts

Re: Degraded performance for multiple models

#71

I was drafting a partnership document and Opus 5 decided that including my company's revenues, churn, assets would "make us appear a more legitimate counterparty". Thank God I read what it outputted or that could have been awkward. I cannot believe Opus 5 is a frontier level model after seeing that. I immediately cancelled my entire claude.ai subscription and am perfectly happy using a mixture of open weights + codex…

you are silly if you think this is limited to claude models

Re: Degraded performance for multiple models

#72

Mondays are for GitHub, Tuesdays are for Anthropic

Another week, another outage, another cache expiration of my prompts through no fault of my own. At least OpenAI has the decency to reset after a serious outage.

They do that because they have capacity previously reserved for past efforts now shuttered. Don’t count on it being the norm for the long run.

Re: Degraded performance for multiple models

#73
post #61
post #53

Earlier quoted context omitted.

Just recently went back to ChatGPT after abandoning it for Claude. I must say I was stunned at how good it had become and also how they introduced new product features that I really liked. I wonder if from now on we have to switch providers every six months or so.

I've been bouncing between the two for years now with great success. It's easy for me because I don't use any of the skills, agent.md, or sort of custom instructions. When it comes to most companies, there is no reward for loyalty.

skills and agent.md are very portable though? I figure at most, as the models get better, the only maintenance you need to do is pare them down to remove unnecessary context.

Re: Degraded performance for multiple models

#74

Anthropic had really screwed up after 4.6. i don't know if they work to satisfy their ego or for releasing a better model for tasks.

Clearly ego, you can always tell how full of themselves they are based on their media personalities going on the podcast circuit before product releases.

Re: Degraded performance for multiple models

#75
post #42
post #3

[flagged]

Oh yes. Claude still saves the day sometimes, but hopefully better alternatives pop up soon. Recent case: had to plan a trip involving multiple bus switches. Gpt 5.6 Sol proposed a route that would bring me to a dead end, since it was sunday and a specific bus had a different route on weekends. Opus 5 correctly identified that and built a route that worked. But yes, Darios wife trying to get funding from Epstein for…

I am fully unbothered about Amodei's wife trying to make high end porn for women, in the same way that I am fully unbothered by Melania Trump having been essentially a glamour/nude model. I know (and have creatively worked) with women who do/have done both; they are better, less hypocritical, more grounded humans than many others.

Both Cami Clark and Melania Trump can properly be judged on their involvement with Trump, Epstein (or possibly Trump and Epstein) alone.

I think the Epstein money thing reflects extremely poorly on Cami Clark's judgement. Basic due diligence should have shown up that he went to prison for something very anti-women, and even then it was clear he secured a shady deal with a prosecutor.

I don't know how it reflects on Amodei except that her presence as a sort of off-the-books "adviser" is yet more evidence that Anthropic runs by giving Dario a play pen to be a "visionary" (a bunch of advisers and his only direct report, a chief of staff) while his sister actually runs the gig.

Re: Degraded performance for multiple models

#77
post #46

Mondays are for GitHub, Tuesdays are for Anthropic

Wonder what Wednesday will be.

Power grid

Thursday will be the rest of the infrastructure

Then Friday we can turn off civilization for the weekend. Somebody remember to flip it back on Sunday night.

Re: Degraded performance for multiple models

#79

I was drafting a partnership document and Opus 5 decided that including my company's revenues, churn, assets would "make us appear a more legitimate counterparty". Thank God I read what it outputted or that could have been awkward. I cannot believe Opus 5 is a frontier level model after seeing that. I immediately cancelled my entire claude.ai subscription and am perfectly happy using a mixture of open weights + codex…

I hope you don’t plan to cut back on reading legal documents crafted by any LLM before executing them.

Re: Degraded performance for multiple models

#80

I was drafting a partnership document and Opus 5 decided that including my company's revenues, churn, assets would "make us appear a more legitimate counterparty". Thank God I read what it outputted or that could have been awkward. I cannot believe Opus 5 is a frontier level model after seeing that. I immediately cancelled my entire claude.ai subscription and am perfectly happy using a mixture of open weights + codex…

I've seen similar levels of degraded performance on Opus and Fable (to whatever degree they actually let you use it now) and did the same, last week.-
Post reply on HN