Live data from Hacker News

Degraded performance for multiple models

status.claude.com

61–70 of 149 posts

Re: Degraded performance for multiple models

#61
post #53

Anthropic had really screwed up after 4.6. i don't know if they work to satisfy their ego or for releasing a better model for tasks.

Just recently went back to ChatGPT after abandoning it for Claude. I must say I was stunned at how good it had become and also how they introduced new product features that I really liked. I wonder if from now on we have to switch providers every six months or so.

I've been bouncing between the two for years now with great success. It's easy for me because I don't use any of the skills, agent.md, or sort of custom instructions.

When it comes to most companies, there is no reward for loyalty.

Re: Degraded performance for multiple models

#64

Earlier quoted context omitted.

https://support.claude.com/en/articles/15910845-claude-code-...

Maybe I'm missing something here, but it sounds like limits were increased and now they're just going back to the levels they were at before?

You're not missing anything. That's correct.

Re: Degraded performance for multiple models

#65

ive found degraded performance on models larger than 4.7. i assume its model damage from overly self righteous post training resulting in false/feigned balance imported into any long running complex task. wish i was joking.

Aren't the model weights frozen?

I think there are other knobs that can be turned without retraining.

Re: Degraded performance for multiple models

#66

Anthropic had really screwed up after 4.6. i don't know if they work to satisfy their ego or for releasing a better model for tasks.

After using Fable more extensively, I've found that it often is lazy or lies or tries to take shortcuts. For a company so sanctimonious about alignment, they seem to be the ones doing the worst at it. Availability aside they've really made me appreciate OpenAI and cheer for other competitors in the marketplace even if I have mixed feelings about using Chinese models.

> I've found that it often is lazy or lies or tries to take shortcuts.

It's funny how Fable reflects the company that produced it.

Re: Degraded performance for multiple models

#67
I was drafting a partnership document and Opus 5 decided that including my company's revenues, churn, assets would "make us appear a more legitimate counterparty". Thank God I read what it outputted or that could have been awkward. I cannot believe Opus 5 is a frontier level model after seeing that. I immediately cancelled my entire claude.ai subscription and am perfectly happy using a mixture of open weights + codex.

Re: Degraded performance for multiple models

#69
post #34
post #27

Despite the years-long moaning on HN about AWS US East being a single point of failure, we've sold our souls to yet another unstable monolith.

LLMs for coding are new. There are lots of alternatives, and there's a burgeoning open source compliment. We'll be fine no matter how Anthropic fares.

Yeah, compared to AWS the lock-in effect is tiny. I'm sure there are highly prioritized plans to "improve" on this.

I guess they would need to control/"own" more of their customers data in proprietary formats. Not markdown/source code in English with agents running on customers' machines.

Something cloud/web-based, "preferably".

Post reply on HN