Hello, I'm one of the engineers who worked on the incident. We have mitigated the incident as of 14:43 PT / 22:43 UTC. Sorry for the trouble.
Also an engineer on this incident. This was a network routing misconfiguration - an overlapping route advertisement caused traffic to some of our inference backends to be blackholed. Detection took longer than we’d like (about 75 minutes from impact to identification), and some of our normal mitigation paths didn’t work as expected during the incident. The bad route has been removed and service is restored. We’re doi…
Elevated errors across many models
121–130 of 170 posts
Re: Elevated errors across many models
#122Hello, I'm one of the engineers who worked on the incident. We have mitigated the incident as of 14:43 PT / 22:43 UTC. Sorry for the trouble.
Any chance you guys could do write ups on these incidents similar to how CloudFlare does? For all the heat some people give them, I trust CloudFlare more with my websites than a lot of other companies because of their dedication to transparency.
Re: Elevated errors across many models
#123Earlier quoted context omitted.
Thank you! Opening an incident as soon as user impact begins is one of those instincts you develop after handling major incidents for years as an SRE at Google, and now at Anthropic. I was also fortunate to be using Claude at that exact moment (for personal reasons), which meant I could immediately see the severity of the outage.
It's important for companies to use their own products.
https://www.theguardian.com/technology/2021/oct/05/facebook-...
I have a memory that Slack fell into this trap too (I could be wrong)
Re: Elevated errors across many models
#124I’m imagining a steampunk dystopia in 50 years: “all world production stopped, LLM hosting went down. The market is in free-fall. Sam, are you there?” Man that cracks me up.
Claude code cut me off a few days ago and I _seriously_ had no idea what to do. I’ve been coding for 33 years and I suddenly felt like anything I did manually would be an order of magnitude slower than it had to be.
Re: Elevated errors across many models
#125Props to them for actually updating their status page as issues are happening rather than hours later. I was working with claude code and hit an API error, checked the status page and sure enough there was an outage. This should be a given for any service that others rely on, but sadly this is seldom the case.
"There's a problem and we already know about it" is so much better than "there's a problem and we don't know about it and/or are hoping it will magically go away and that we won't be embarrassed".
Re: Elevated errors across many models
#126Re: Elevated errors across many models
#127Props to them for actually updating their status page as issues are happening rather than hours later. I was working with claude code and hit an API error, checked the status page and sure enough there was an outage. This should be a given for any service that others rely on, but sadly this is seldom the case.
Re: Elevated errors across many models
#128I’m imagining a steampunk dystopia in 50 years: “all world production stopped, LLM hosting went down. The market is in free-fall. Sam, are you there?” Man that cracks me up.
Re: Elevated errors across many models
#129I’m imagining a steampunk dystopia in 50 years: “all world production stopped, LLM hosting went down. The market is in free-fall. Sam, are you there?” Man that cracks me up.
The nice thing is unlike Cloudflare or AWS you can actually host good LLMs locally. I see a future where a non-trivial percentage of devs have an expensive workstation that runs all of the AI locally.
Re: Elevated errors across many models
#130Props to them for actually updating their status page as issues are happening rather than hours later. I was working with claude code and hit an API error, checked the status page and sure enough there was an outage. This should be a given for any service that others rely on, but sadly this is seldom the case.
Thank you! Opening an incident as soon as user impact begins is one of those instincts you develop after handling major incidents for years as an SRE at Google, and now at Anthropic. I was also fortunate to be using Claude at that exact moment (for personal reasons), which meant I could immediately see the severity of the outage.