Live data from Hacker News

Ask HN: Why were OpenAI, Claude, and Grok simultaneously down?

news.ycombinator.com

671–680 of 780 posts

Re: Ask HN: Why were OpenAI, Claude, and Grok simultaneously down?

#671

Earlier quoted context omitted.

OpenAI and Anthropic have their systems in dozens of data centers, including using compute from the major cloud providers. Are you implying that all of these data centers (and many of their employees) are involved in helping the the US government secretly tap every single AI conversation by routing them through some unknown network/device? Or did a couple of companies with poor uptime records happen to have overlappi…

What makes you think they'd need to touch every datacenter? All of these endpoints use existing providers with decades-long history at this point, and network monitoring is already a proven 'feature' of the agencies they'd need to co-exist with over their lifetimes. If anything, Occam's Razor would point to a common denominator with all of them, given it wasn't network-wide, as far as i know.

Explain the system in which you could capture all of these chats with no knowledge of anyone in these data centers. How are they routed to this NSA system or through some NSA device when these companies' compute are spread over hundreds of data centers?

> All of these endpoints use existing providers with decades-long history at this point

That is just factually inaccurate. Their data centers aren't old and they lease a lot of compute from companies that didn't exist 5 years ago.

Re: Ask HN: Why were OpenAI, Claude, and Grok simultaneously down?

#673

Neither of these companies have stellar uptime records. Their downtime episodes overlapped in this instance. In this case, it was a partial downtime for both. Also, OpenAI is saying what caused it: > "A routing error starting around 7:43 am PT on Thursday, September 3, made ChatGPT and Codex unavailable for some users across platforms" Anthropic stated their issue started earlier: > "The company began alerting about…

[deleted]

Re: Ask HN: Why were OpenAI, Claude, and Grok simultaneously down?

#674

Earlier quoted context omitted.

What does having a "well trusted TLS cert" enable for them in this case, exactly? Having a magical cert doesn't mean you can just intercept everything.

On the contrary, it lets you MITM encrypted communications by swapping the website's original certificate for the "well trusted TLS cert"

No, it doesn't. HSTS and other methods prevent this from happening.

Re: Ask HN: Why were OpenAI, Claude, and Grok simultaneously down?

#675

It’s probably the thing that everyone thinks it is. OpenAI, Anthropic and SpaceXAI are all routed through something that we’re not supposed to know exists and that thing had a whoopsie.

If it does exist, why would it work this way, and not the obvious way of streaming logs… which would not cause an outage if it failed.

Re: Ask HN: Why were OpenAI, Claude, and Grok simultaneously down?

#676
post #639

The OpenAI outage lasted only 15 minutes and when it happened everyone started to use the other models which created super heavy load for them. This then cascaded into them all being down. Does this really need an explanation?

it wasn't timed like a cascade, and it relies on the premise that every single frontier provider is working so efficiently that they spend exactly what they need to provide for their exact market with perfect margins. I do not believe personally that 1) they can forecast their load that perfectly 2) they chose to remain that inflexible in a world where they are at each others' throats and a single meme can cause burs…

Or it could simply mean that they're operating at the limit of the capacity they were able to purchase and do not have headroom to handle load spikes. From what I understand, that's the situation Anthropic is in. And since Anthropic is now leasing a large portion of xAI's datacenter capacity, it's plausible that an Anthropic load spike could cause issues for xAI as well.

Your comment makes it sound like they can just push a button and spin up more capacity -- but at this scale and in this GPU-constrained environment, that's not really how it works.

Re: Ask HN: Why were OpenAI, Claude, and Grok simultaneously down?

#677

Earlier quoted context omitted.

Could you expand on this more? It's not clear to me what this is implying.

https://en.wikipedia.org/wiki/Room_641A

Traffic doesn’t need to be routed through it though, just tee’d to it,

Re: Ask HN: Why were OpenAI, Claude, and Grok simultaneously down?

#679
post #650

It could be as simple as a new model (astra) was released which takes more resources combined with a surge in usage due to novelty took down OpenAI. Meanwhile everyone at big companies have the ability to switch models and moved to Anthropic pushing it too over the edge.

People keep saying this "everyone can switch" thing, but it's not my experience at $VeryBigCorp. We don't have an Anthropic contract at all. Is this different in other places? The bigger and more bureaucratic an org is, the less I would expect it to have contracts with all the providers. Curious about others' experiences.

At my company (not massive but not tiny either - I think its about 8000 global employees) we get a choice between pretty much all available Google, OpenAI, Anthropic, XAI models.

Using an agent-agnostic harness like Pi switching is trivial - I run into occasional disconnects and slowdown and switch quite easily.

Re: Ask HN: Why were OpenAI, Claude, and Grok simultaneously down?

#680

Earlier quoted context omitted.

Could you expand on this more? It's not clear to me what this is implying.

https://en.wikipedia.org/wiki/Room_641A

Worth mentioning that this room takes a split from the main feed and is not in the path of traffic. Whatever is in this room could go down and it would not cause an outage.
Post reply on HN