Live data from Hacker News

Elevated errors across many models

status.claude.com

101–110 of 170 posts

Re: Elevated errors across many models

#101

In the Claude.ai chat, this was announced to me as "You have reached the messages quota for your account. It will reset in 2 hours, or you can upgrade now" Either I have perfect timing for reaching my quota limits, or some product monetization manager deserves a raise.

i ran into the same thing, i thought it was just timing

Re: Elevated errors across many models

#102
post #96
post #58

Hello, I'm one of the engineers who worked on the incident. We have mitigated the incident as of 14:43 PT / 22:43 UTC. Sorry for the trouble.

Also an engineer on this incident. This was a network routing misconfiguration - an overlapping route advertisement caused traffic to some of our inference backends to be blackholed. Detection took longer than we’d like (about 75 minutes from impact to identification), and some of our normal mitigation paths didn’t work as expected during the incident. The bad route has been removed and service is restored. We’re doi…

Was this a typo situation or a bad process thing ?

Back when I did website QA Automation I'd manually check the website at the end of my day. Nothing extensive, just looking at the homepage for piece of mind.

Once a senior engineer decided to bypass all of our QA, deploy and took down prod. Fun times.

Re: Elevated errors across many models

#104
post #80

Props to them for actually updating their status page as issues are happening rather than hours later. I was working with claude code and hit an API error, checked the status page and sure enough there was an outage. This should be a given for any service that others rely on, but sadly this is seldom the case.

Thank you! Opening an incident as soon as user impact begins is one of those instincts you develop after handling major incidents for years as an SRE at Google, and now at Anthropic. I was also fortunate to be using Claude at that exact moment (for personal reasons), which meant I could immediately see the severity of the outage.

It's important for companies to use their own products.

Re: Elevated errors across many models

#105

I’m imagining a steampunk dystopia in 50 years: “all world production stopped, LLM hosting went down. The market is in free-fall. Sam, are you there?” Man that cracks me up.

“A lone coder, trained in the direct manipulation of symbols—an elegant weapon from a more civilized age—-is now all that stands between humanity and darkness.” etc

Re: Elevated errors across many models

#106
post #96

Earlier quoted context omitted.

Also an engineer on this incident. This was a network routing misconfiguration - an overlapping route advertisement caused traffic to some of our inference backends to be blackholed. Detection took longer than we’d like (about 75 minutes from impact to identification), and some of our normal mitigation paths didn’t work as expected during the incident. The bad route has been removed and service is restored. We’re doi…

Was this a typo situation or a bad process thing ? Back when I did website QA Automation I'd manually check the website at the end of my day. Nothing extensive, just looking at the homepage for piece of mind. Once a senior engineer decided to bypass all of our QA, deploy and took down prod. Fun times.

[flagged]

Re: Elevated errors across many models

#108
post #77

Earlier quoted context omitted.

Ah, you need to buy into this dystopia wholesale. The internet is also down because the LLMs fucked up the BGP routing table, which congress agreed (at the time) should run through the LLM interface. Imagination, either the first or last thing to die in 2075.

Congress administrating BGP? Now we’re talking dystopia!

“Hey folks, did you know in 100 years you can’t just call the town doc? Nah, you need to go get a referral. No, for real. Yeah, yeah, that is in fact a compound fracture. I can’t treat it without a referral. Congress made the rules.”

Is it so different?

Re: Elevated errors across many models

#109
post #96
post #58

Hello, I'm one of the engineers who worked on the incident. We have mitigated the incident as of 14:43 PT / 22:43 UTC. Sorry for the trouble.

Also an engineer on this incident. This was a network routing misconfiguration - an overlapping route advertisement caused traffic to some of our inference backends to be blackholed. Detection took longer than we’d like (about 75 minutes from impact to identification), and some of our normal mitigation paths didn’t work as expected during the incident. The bad route has been removed and service is restored. We’re doi…

Trying to understand what this means.

Did the bad route cause an overload? Was there a code error on that route that wasn’t spotted? Was it a code issue or an instance that broke?

Re: Elevated errors across many models

#110
post #67

Earlier quoted context omitted.

What's the best you can do hosting an LLM locally for under $X dollars. Let's say $5000. Is there a reference guide online for this? Is there a straight answer or does it depend? I've looked at Nvidia spark and high end professional GPUs but they all seem to have serious drawbacks.

https://www.reddit.com/r/LocalLLaMA/

That's nice, thank you, I've joined and will follow. They don't seem to have a wiki or about page that synthesizes the current state of the art though.
Post reply on HN