Live data from Hacker News

A postmortem of three recent issues

anthropic.com

41–50 of 120 posts

Re: A postmortem of three recent issues

#41
post #27
post #23

Earlier quoted context omitted.

That 30% is of ALL users, not users who made a request, important to note the weasel wording there. How many users forget they have a sub? How many get a sub through work and don't use it often? I'd bet a large number tbh based on other subscription services.

(I work at Anthropic) It's 30% of all CC users that made a request during that period. We've updated the post to be clearer.

Thanks for the correction and updating the post.

I typically read corporate posts as cynically as possible, since it's so common to word things in any way to make the company look better.

Glad to see an outlier!

Re: A postmortem of three recent issues

#42

The value of figuring out how to make their LLM serving deterministic might help them track this down. There was a recent paper about how the received wisdom that kept assigning it to floating point associativity actually overlooked the real reasons for non-determinism [1]. [1] https://thinkingmachines.ai/blog/defeating-nondeterminism-in...

It has a big impact on performance to do determinism. Which leaves using another model to essentially IQ test their models with reporting and alerting.

Re: A postmortem of three recent issues

#43
With all due respect to the Anthropic team, I think the Claude status page[1] warrants an internal code red for quality. There were 50 incidents in July, 40 incidents in August, and 21 so far in September. I have worked in places where we started approaching half these numbers and they always resulted in a hard pivot to focusing on uptime and quality.

Despite this I'm still a paying customer because Claude is a fantastic product and I get a lot of value from it. After trying the API it became a no brainer to buy a 20x Max membership. The amount of stuff I've gotten done with Claude has been awesome.

The last several weeks have strongly made me question my subscription. I appreciate the openness of this post, but as a customer I'm not happy.

I don't trust that these issues are all discovered and resolved yet, especially the load balancing ones. At least anecdotally I notice that around 12 ET (9AM pacific) my Claude Code sessions noticeably drop in quality. Again, I hope the team is able to continue finding and fixing these issues. Even running local models on my own machine at home I run into complicated bugs all the time — I won't pretend these are easy problems, they are difficult to find and fix.

[1] https://status.anthropic.com/history

Re: A postmortem of three recent issues

#44
post #18

Earlier quoted context omitted.

I'm not sure if you can claim these were "less prevalent than anecdotal online reports". From their article: > Approximately 30% of Claude Code users had at least one message routed to the wrong server type, resulting in degraded responses. > However, some users were affected more severely, as our routing is "sticky". This meant that once a request was served by the incorrect server, subsequent follow-ups were likely…

I don't know about you but my feed is filled with people claiming that they are surely quantizating the model, Anthropic is purposefully degrading things to save money, etc etc. 70% of users were not impacted. 30% had at least one message degraded. One message is basically nothing. I would have appreciated if they had released the full distribution of impact though.

> Anthropic is purposefully degrading things to save money

Regardless of whether it’s to save money, it’s purposefully inaccurate:

“When Claude generates text, it calculates probabilities for each possible next word, then randomly chooses a sample from this probability distribution.”

I think the reason for this is that if you were to always choose the highest probable next word, you may actually always end up with the wrong answer and/or get stuck in a loop.

They could sandbag their quality or rate limit, and I know they will rate limit because I’ve seen it. But, this is a race. It’s not like Microsoft being able to take in the money for years because people will keep buying Windows. AI companies can try to offer cheap service to government and college students, but brand loyalty is less important than selecting the smarter AI to help you.

Re: A postmortem of three recent issues

#45
post #40

I’m pretty surprised that Anthropic can directly impact the infra for AWS Bedrock as this article suggests. That goes against AWSs commitments. I’m sure the same is true for Google Vertex but I haven’t digged in there from a compliance perspective before. > Our own privacy practices also created challenges in investigating reports. Our internal privacy and security controls limit how and when engineers can access use…

(Anthropic employee, speaking in a personal capacity) > I’m pretty surprised that Anthropic can directly impact the infra for AWS Bedrock as this article suggests. We don't directly manage AWS Bedrock deployments today, those are managed by AWS. > I can’t imagine the average person equates hitting this button with forfeiting their privacy. We specify > Submitting this report will send the entire current conversation…

"have a human take a look at this conversation (from {time} to {time})"

Re: A postmortem of three recent issues

#46

> We don't typically share this level of technical detail about our infrastructure, but the scope and complexity of these issues justified a more comprehensive explanation. Layered in aggrandizing. You host a service, people give you money.

No, what that statement means is "we know that if we just say 'we weren't downgrading performance to save money', you won't believe us, so here is a deep dive on the actual reason it happened"

they're big, and we expect proper behavior out of them when they mess up. that includes public details.

Re: A postmortem of three recent issues

#48
post #40

I’m pretty surprised that Anthropic can directly impact the infra for AWS Bedrock as this article suggests. That goes against AWSs commitments. I’m sure the same is true for Google Vertex but I haven’t digged in there from a compliance perspective before. > Our own privacy practices also created challenges in investigating reports. Our internal privacy and security controls limit how and when engineers can access use…

(Anthropic employee, speaking in a personal capacity) > I’m pretty surprised that Anthropic can directly impact the infra for AWS Bedrock as this article suggests. We don't directly manage AWS Bedrock deployments today, those are managed by AWS. > I can’t imagine the average person equates hitting this button with forfeiting their privacy. We specify > Submitting this report will send the entire current conversation…

Sounds fine to me. I'm assuming it wasn't obvious to readers that there was a confirmation message that appears when thumbs down is clicked.

Re: A postmortem of three recent issues

#49

With all due respect to the Anthropic team, I think the Claude status page[1] warrants an internal code red for quality. There were 50 incidents in July, 40 incidents in August, and 21 so far in September. I have worked in places where we started approaching half these numbers and they always resulted in a hard pivot to focusing on uptime and quality. Despite this I'm still a paying customer because Claude is a fanta…

I've become extremely nervous about these sudden declines in quality. Thankfully I don't have a production product using AI (yet), but in my own development experience - the model becoming dramatically dumber suddenly is very difficult to work around.

At this point, I'd be surprised if the different vendors on openrouter weren't abusing their trust by silently dropping context/changing quantization levels/reducing experts - or other mischievous means of delivering the same model at lower compute.

Re: A postmortem of three recent issues

#50

With all due respect to the Anthropic team, I think the Claude status page[1] warrants an internal code red for quality. There were 50 incidents in July, 40 incidents in August, and 21 so far in September. I have worked in places where we started approaching half these numbers and they always resulted in a hard pivot to focusing on uptime and quality. Despite this I'm still a paying customer because Claude is a fanta…

I don’t know whether they are better or worse than others. One for sure, a lot of companies lie on their status pages. I encounter outages frequently which are not reported on their status pages. Nowadays, I’m more surprised when they self report some problems. Personally, I didn’t have serious problems with Claude so far, but it’s possible that I was just lucky. In my perspective, it just seems that they are reporting outages in a more faithful way. But that can be completely coincidental.
Post reply on HN