Live data from Hacker News

A postmortem of three recent issues

anthropic.com

71–80 of 120 posts

Re: A postmortem of three recent issues

#71
post #14

And yet no offers of credits to make things right for the users, for what was essentially degraded performance of what you paid for. I know I'll probably get push back on this, but it left a sour taste in my mouth when I paid for a $200 sub that felt like it was less useful than ChatGPT Plus ($20) at times. Or to summarize: [south park "we're sorry" gif]

I’m pretty certain if you check the ToS that Anthropic doesn’t guarantee a level of response quality, and explicitly even says there is zero guarantee, even for paid plans.

So to be fair, you are getting exactly what you paid for - a non-deterministic set of generated responses of varying quality and accuracy.

Re: A postmortem of three recent issues

#72
post #58

I don't believe for one second that response quality dropped because of an infrastructural change and remained degraded, unnoticed, for weeks. This simply does not pass the sniff test.

Can you provide any proof of what you’re saying? Any examples that would bear out what you’re asserting? Anything at all?

“I refuse to believe what the people who would know the best said, for no real reason except that it doesn’t feel right” isn’t exactly the level of considered response we’re hoping for here on HN. :)

Re: A postmortem of three recent issues

#73
post #58

I don't believe for one second that response quality dropped because of an infrastructural change and remained degraded, unnoticed, for weeks. This simply does not pass the sniff test.

Can you provide any proof of what you’re saying? Any examples that would bear out what you’re asserting? Anything at all? “I refuse to believe what the people who would know the best said, for no real reason except that it doesn’t feel right” isn’t exactly the level of considered response we’re hoping for here on HN. :)

Have you used these tools at all? It's incredibly obvious. It was obvious for weeks during August, where several people posted about degradation on r/ClaudeAI ...

There's a thousand and one reasons why a company valued in the billions, with the eyes of the world watching, would not be completely honest in their public response.

Re: A postmortem of three recent issues

#74
post #40

Earlier quoted context omitted.

(Anthropic employee, speaking in a personal capacity) > I’m pretty surprised that Anthropic can directly impact the infra for AWS Bedrock as this article suggests. We don't directly manage AWS Bedrock deployments today, those are managed by AWS. > I can’t imagine the average person equates hitting this button with forfeiting their privacy. We specify > Submitting this report will send the entire current conversation…

Sounds fine to me. I'm assuming it wasn't obvious to readers that there was a confirmation message that appears when thumbs down is clicked.

Yes, I don't use Claude so I wasn't aware. I'm glad to hear it sounds like it is conspicuous.

Re: A postmortem of three recent issues

#75
post #38

I’m pretty surprised that Anthropic can directly impact the infra for AWS Bedrock as this article suggests. That goes against AWSs commitments. I’m sure the same is true for Google Vertex but I haven’t digged in there from a compliance perspective before. > Our own privacy practices also created challenges in investigating reports. Our internal privacy and security controls limit how and when engineers can access use…

> This is pretty concerning. I can’t imagine the average person equates hitting this button with forfeiting their privacy. When you click "thumbs down" you get the message "Submitting this report will send the entire current conversation to Anthropic for future improvements to our models." before you submit the report, I'd consider that pretty explicit.

Great to hear. I'm not a Claude user and the article did not make it seem that way.

Re: A postmortem of three recent issues

#76
post #40

I’m pretty surprised that Anthropic can directly impact the infra for AWS Bedrock as this article suggests. That goes against AWSs commitments. I’m sure the same is true for Google Vertex but I haven’t digged in there from a compliance perspective before. > Our own privacy practices also created challenges in investigating reports. Our internal privacy and security controls limit how and when engineers can access use…

(Anthropic employee, speaking in a personal capacity) > I’m pretty surprised that Anthropic can directly impact the infra for AWS Bedrock as this article suggests. We don't directly manage AWS Bedrock deployments today, those are managed by AWS. > I can’t imagine the average person equates hitting this button with forfeiting their privacy. We specify > Submitting this report will send the entire current conversation…

> We don't directly manage AWS Bedrock deployments today, those are managed by AWS.

That was my understanding before this article. But the article is pretty clear that these were "infrastructure bugs" and the one related to AWS Bedrock specifically says it was because "requests were misrouted to servers". If Anthropic doesn't manage the AWS Bedrock deployments, how could it be impacting the load balancer?

Re: A postmortem of three recent issues

#77

With all due respect to the Anthropic team, I think the Claude status page[1] warrants an internal code red for quality. There were 50 incidents in July, 40 incidents in August, and 21 so far in September. I have worked in places where we started approaching half these numbers and they always resulted in a hard pivot to focusing on uptime and quality. Despite this I'm still a paying customer because Claude is a fanta…

This is always why you should put as few incidents on status page as possible. People's opinion will drop and then the negative effect will fade over time. But if you have a status page then it's incontrovertible proof. Better to lie. They'll forget.

e.g. S3 has many times encountered increased error rate but doesn't report. No one says anything about S3.

People will say many things, but their behaviour is to reward the lie. Every growth hack startup guy knows this already.

Re: A postmortem of three recent issues

#78

With all due respect to the Anthropic team, I think the Claude status page[1] warrants an internal code red for quality. There were 50 incidents in July, 40 incidents in August, and 21 so far in September. I have worked in places where we started approaching half these numbers and they always resulted in a hard pivot to focusing on uptime and quality. Despite this I'm still a paying customer because Claude is a fanta…

What makes it even worse is the status page doesn't capture all smaller incidents. This is the same for all providers. If they actually provided real time graphs of token latency, failed requests, token/s etc I think they'd be pretty horrific. If you trust this OpenRouter data the uptime record of these APIs is... not good to say the least: https://openrouter.ai/openai/gpt-5/uptime It's clear to me that every provide…

Artificial Analysis also monitor LLM provider APIs independently "based on 8 measurements each day at different times" you can see the degradation as opus 4.1 came online https://artificialanalysis.ai/providers/anthropic#end-to-end...

Re: A postmortem of three recent issues

#79
post #14

And yet no offers of credits to make things right for the users, for what was essentially degraded performance of what you paid for. I know I'll probably get push back on this, but it left a sour taste in my mouth when I paid for a $200 sub that felt like it was less useful than ChatGPT Plus ($20) at times. Or to summarize: [south park "we're sorry" gif]

Hey, they need that $200 to postpone their inevitable bankruptcy

Re: A postmortem of three recent issues

#80
post #18

Earlier quoted context omitted.

I'm not sure if you can claim these were "less prevalent than anecdotal online reports". From their article: > Approximately 30% of Claude Code users had at least one message routed to the wrong server type, resulting in degraded responses. > However, some users were affected more severely, as our routing is "sticky". This meant that once a request was served by the incorrect server, subsequent follow-ups were likely…

I don't know about you but my feed is filled with people claiming that they are surely quantizating the model, Anthropic is purposefully degrading things to save money, etc etc. 70% of users were not impacted. 30% had at least one message degraded. One message is basically nothing. I would have appreciated if they had released the full distribution of impact though.

Routing bug was sticky, "one message is basically nothing" is not what was happening - if you were affected, you were more likely to be affected even more.
Post reply on HN