Live data from Hacker News

Elevated errors across many models

status.claude.com

121–130 of 170 posts

Re: Elevated errors across many models

#121
post #96
post #58

Hello, I'm one of the engineers who worked on the incident. We have mitigated the incident as of 14:43 PT / 22:43 UTC. Sorry for the trouble.

Also an engineer on this incident. This was a network routing misconfiguration - an overlapping route advertisement caused traffic to some of our inference backends to be blackholed. Detection took longer than we’d like (about 75 minutes from impact to identification), and some of our normal mitigation paths didn’t work as expected during the incident. The bad route has been removed and service is restored. We’re doi…

I don't know if you guys do write ups, but cloudflare's write ups on outages is in my eyes the gold standard the entire industry should follow.

Re: Elevated errors across many models

#122
post #58

Hello, I'm one of the engineers who worked on the incident. We have mitigated the incident as of 14:43 PT / 22:43 UTC. Sorry for the trouble.

Any chance you guys could do write ups on these incidents similar to how CloudFlare does? For all the heat some people give them, I trust CloudFlare more with my websites than a lot of other companies because of their dedication to transparency.

We're considering this!

Re: Elevated errors across many models

#123
post #80

Earlier quoted context omitted.

Thank you! Opening an incident as soon as user impact begins is one of those instincts you develop after handling major incidents for years as an SRE at Google, and now at Anthropic. I was also fortunate to be using Claude at that exact moment (for personal reasons), which meant I could immediately see the severity of the outage.

It's important for companies to use their own products.

Unless using your own dogfood prevents you from fixing it if it breaks

https://www.theguardian.com/technology/2021/oct/05/facebook-...

I have a memory that Slack fell into this trap too (I could be wrong)

Re: Elevated errors across many models

#124

I’m imagining a steampunk dystopia in 50 years: “all world production stopped, LLM hosting went down. The market is in free-fall. Sam, are you there?” Man that cracks me up.

Claude code cut me off a few days ago and I _seriously_ had no idea what to do. I’ve been coding for 33 years and I suddenly felt like anything I did manually would be an order of magnitude slower than it had to be.

you can’t say things like this on HN these days :)

Re: Elevated errors across many models

#125

Props to them for actually updating their status page as issues are happening rather than hours later. I was working with claude code and hit an API error, checked the status page and sure enough there was an outage. This should be a given for any service that others rely on, but sadly this is seldom the case.

"There's a problem and we already know about it" is so much better than "there's a problem and we don't know about it and/or are hoping it will magically go away and that we won't be embarrassed".

"If we admit to it we may have to compensate per SLAs, so dishonesty it is!"

Re: Elevated errors across many models

#127

Props to them for actually updating their status page as issues are happening rather than hours later. I was working with claude code and hit an API error, checked the status page and sure enough there was an outage. This should be a given for any service that others rely on, but sadly this is seldom the case.

Indeed! I checked their status page within 2 minutes of having issues and it was updated to show they had detected it.

Re: Elevated errors across many models

#128

I’m imagining a steampunk dystopia in 50 years: “all world production stopped, LLM hosting went down. The market is in free-fall. Sam, are you there?” Man that cracks me up.

"We vibe coded the problem into existence but now the LLM is down we can't vibe fix it"

Re: Elevated errors across many models

#129

I’m imagining a steampunk dystopia in 50 years: “all world production stopped, LLM hosting went down. The market is in free-fall. Sam, are you there?” Man that cracks me up.

The nice thing is unlike Cloudflare or AWS you can actually host good LLMs locally. I see a future where a non-trivial percentage of devs have an expensive workstation that runs all of the AI locally.

I'd imagine at some point the companies will just... stop publishing any open models precisely to stop that and keep people paying the subscription.

Re: Elevated errors across many models

#130
post #80

Props to them for actually updating their status page as issues are happening rather than hours later. I was working with claude code and hit an API error, checked the status page and sure enough there was an outage. This should be a given for any service that others rely on, but sadly this is seldom the case.

Thank you! Opening an incident as soon as user impact begins is one of those instincts you develop after handling major incidents for years as an SRE at Google, and now at Anthropic. I was also fortunate to be using Claude at that exact moment (for personal reasons), which meant I could immediately see the severity of the outage.

Sweet. Hopefully it is more than instinct but a codified at Anthropic. I.e. a graduate engineer with little experience can assess and raise incident if needed.
Post reply on HN