Live data from Hacker News

Elevated errors across many models

status.claude.com

141–150 of 170 posts

Re: Elevated errors across many models

#142
post #122

Earlier quoted context omitted.

Any chance you guys could do write ups on these incidents similar to how CloudFlare does? For all the heat some people give them, I trust CloudFlare more with my websites than a lot of other companies because of their dedication to transparency.

We're considering this!

I already love the product, and I think it would be great to see. Even if its not as "quickly" as CloudFlares (they post ASAP its insane) I would still be happy to see postmortem threads. We all learn industry wide from them.

Re: Elevated errors across many models

#143
post #96

Earlier quoted context omitted.

Also an engineer on this incident. This was a network routing misconfiguration - an overlapping route advertisement caused traffic to some of our inference backends to be blackholed. Detection took longer than we’d like (about 75 minutes from impact to identification), and some of our normal mitigation paths didn’t work as expected during the incident. The bad route has been removed and service is restored. We’re doi…

I don't know if you guys do write ups, but cloudflare's write ups on outages is in my eyes the gold standard the entire industry should follow.

A big reason for that is it comes from the CEO. Other providers have a team and then at least 2 to 3 layers of management above them and a dotted line legal counsel. So the goal posts randomly shift from "more information" to "no information" over time based on the relationships of that entire chain, the customer heat of the moment, and personality.

Underneath a public statement they all have extremely detailed post-mortems. But how much goes public is 100% random from the customer's perspective. There's no Monday Morning QB'ing the CEO, but there absolutely is "Day-Shift SRE Leader Phil"

Re: Elevated errors across many models

#144

Earlier quoted context omitted.

Claude code cut me off a few days ago and I _seriously_ had no idea what to do. I’ve been coding for 33 years and I suddenly felt like anything I did manually would be an order of magnitude slower than it had to be.

You can always use gemini cli, its pretty good

Nah, I've now resolved to never use it, it'd a total time waster. When it works it's decent, but it does not work like half of time for me.

Re: Elevated errors across many models

#145
post #96

Earlier quoted context omitted.

Also an engineer on this incident. This was a network routing misconfiguration - an overlapping route advertisement caused traffic to some of our inference backends to be blackholed. Detection took longer than we’d like (about 75 minutes from impact to identification), and some of our normal mitigation paths didn’t work as expected during the incident. The bad route has been removed and service is restored. We’re doi…

I don't know if you guys do write ups, but cloudflare's write ups on outages is in my eyes the gold standard the entire industry should follow.

Cloudflare deploys stuff on Fridays, and it directly affected shopify, one of their major ecommerce customers. Until they fix their internal processes all writeups should be seen as purely marketing material.

Re: Elevated errors across many models

#146

Earlier quoted context omitted.

It's important for companies to use their own products.

Unless using your own dogfood prevents you from fixing it if it breaks https://www.theguardian.com/technology/2021/oct/05/facebook-... I have a memory that Slack fell into this trap too (I could be wrong)

Now I’m imagining the folks at Slack gritting their teeth and using MS Teams

Re: Elevated errors across many models

#147
post #96

Earlier quoted context omitted.

Also an engineer on this incident. This was a network routing misconfiguration - an overlapping route advertisement caused traffic to some of our inference backends to be blackholed. Detection took longer than we’d like (about 75 minutes from impact to identification), and some of our normal mitigation paths didn’t work as expected during the incident. The bad route has been removed and service is restored. We’re doi…

Was this a typo situation or a bad process thing ? Back when I did website QA Automation I'd manually check the website at the end of my day. Nothing extensive, just looking at the homepage for piece of mind. Once a senior engineer decided to bypass all of our QA, deploy and took down prod. Fun times.

In these times, it could be "the AI did it".

Re: Elevated errors across many models

#148
post #32

Earlier quoted context omitted.

I use it as much as my brain can handle and I never exceed my Max plan quota.

Just a warning for those not on the max plan; if you pay by the token or have the lower tier plans you can easily blow through $100s or cap your plan in under an hour. The rates for paying by the token are insane and the scaling from pro to max is also pretty crazy. They made pro have many times more value than paying per token and then they made max again have 25x more tokens than pro on the $200 plan. It’s a bit li…

Yeah well, wait til they take it away

Re: Elevated errors across many models

#149

Earlier quoted context omitted.

It's important for companies to use their own products.

Unless using your own dogfood prevents you from fixing it if it breaks https://www.theguardian.com/technology/2021/oct/05/facebook-... I have a memory that Slack fell into this trap too (I could be wrong)

Facebook notoriously had to cut open the doors to one of their data centers.

Google SRE still keeps IRC available in case of an emergency.

Re: Elevated errors across many models

#150

I’m imagining a steampunk dystopia in 50 years: “all world production stopped, LLM hosting went down. The market is in free-fall. Sam, are you there?” Man that cracks me up.

Claude code cut me off a few days ago and I _seriously_ had no idea what to do. I’ve been coding for 33 years and I suddenly felt like anything I did manually would be an order of magnitude slower than it had to be.

[deleted]
Post reply on HN