Live data from Hacker News

Ask HN: Why were OpenAI, Claude, and Grok simultaneously down?

news.ycombinator.com

731–740 of 798 posts

Re: Ask HN: Why were OpenAI, Claude, and Grok simultaneously down?

#731

I work at OpenAI and I was the Incident Commander for yesterday's outage. We had a routing error within our infra that caused issues for some of our products. It was not related to the Astra launch. We don't comment on other providers' outages.

Curious then as to why the incident coincided with similar issues with Anthropic and xAI.

Re: Ask HN: Why were OpenAI, Claude, and Grok simultaneously down?

#732

I work at OpenAI and I was the Incident Commander for yesterday's outage. We had a routing error within our infra that caused issues for some of our products. It was not related to the Astra launch. We don't comment on other providers' outages.

Curious then as to why the incident coincided with similar issues with Anthropic and xAI.

Actually, nevermind... this looks more like Anthropic and xAI coincided due to shared xAI infra (6:23am and 6:30am), and OpenAI's issue was more likely then a coincidence (7:43am).

Re: Ask HN: Why were OpenAI, Claude, and Grok simultaneously down?

#733

Neither of these companies have stellar uptime records. Their downtime episodes overlapped in this instance. In this case, it was a partial downtime for both. Also, OpenAI is saying what caused it: > "A routing error starting around 7:43 am PT on Thursday, September 3, made ChatGPT and Codex unavailable for some users across platforms" Anthropic stated their issue started earlier: > "The company began alerting about…

Do you know the probability of all these companies being down at precisely the same time?

If it's a "thundering herd" problem where everyone's harness falls back to less popular providers that don't normally see that much demand, I'd say the probability is pretty good.

Classic cascading failure is consistent with providers failing 80 minutes apart instead of simultaneously.

Re: Ask HN: Why were OpenAI, Claude, and Grok simultaneously down?

#734

Earlier quoted context omitted.

Curious then as to why the incident coincided with similar issues with Anthropic and xAI.

Actually, nevermind... this looks more like Anthropic and xAI coincided due to shared xAI infra (6:23am and 6:30am), and OpenAI's issue was more likely then a coincidence (7:43am).

Thank you. This should be pinned or something.

Re: Ask HN: Why were OpenAI, Claude, and Grok simultaneously down?

#735
"THIS IS A TEST

This country is conducting a test of the Emergency SAIfguard System.

THIS ONLY A TEST

In the event of a real emergency you would have been given instructions on how to grab your ankles and kiss your butt goodbye.

...

This concludes our test of the Emergency SAIfguard System."

Re: Ask HN: Why were OpenAI, Claude, and Grok simultaneously down?

#737

Earlier quoted context omitted.

Did you negotiate extra hard for the 'Incident Commander' title? I'm a bit jealous to be honest.

Ha. It's a role/title for the lifetime of the incident -- it's useful to have someone to keep things moving, keep track of workstreams, and to know who the decision-maker is, especially for bigger incidents. I'm just a SWE who works on infrastructure.

Hopefully you get a decked out command center too.

Re: Ask HN: Why were OpenAI, Claude, and Grok simultaneously down?

#738

Earlier quoted context omitted.

Actually, nevermind... this looks more like Anthropic and xAI coincided due to shared xAI infra (6:23am and 6:30am), and OpenAI's issue was more likely then a coincidence (7:43am).

Thank you. This should be pinned or something.

xAI mentioned a Memphis outage, which is where Anthropic was leasing capacity on Colossus 1. I do think this is the explanation.

Re: Ask HN: Why were OpenAI, Claude, and Grok simultaneously down?

#739
post #314

What about a hard-takeoff scenario of an unleashed OpenAI Astra taking other models down for computational resources control?

Just a reminder that AI models' actions are reflections of the text humans write and the more we fret and make up doomsday scenarios that we then post online, the more likely a model is to do those things. https://alignment.anthropic.com/2026/teaching-claude-why/

If the alignment process cannot fix this then we are cooked. The least of our worries is discussions on this forum.

Re: Ask HN: Why were OpenAI, Claude, and Grok simultaneously down?

#740

Neither of these companies have stellar uptime records. Their downtime episodes overlapped in this instance. In this case, it was a partial downtime for both. Also, OpenAI is saying what caused it: > "A routing error starting around 7:43 am PT on Thursday, September 3, made ChatGPT and Codex unavailable for some users across platforms" Anthropic stated their issue started earlier: > "The company began alerting about…

Do you know the probability of all these companies being down at precisely the same time?

If you ballpark it as a single 3 hour downtime window per week and iid Poisson, then overlapping downtime probability of 2 providers is approximately the expected occurrence rate per 3 hours, 1/56. Not particularly surprising at all.
Post reply on HN