Live data from Hacker News

Ask HN: Why were OpenAI, Claude, and Grok simultaneously down?

news.ycombinator.com

701–710 of 781 posts

Re: Ask HN: Why were OpenAI, Claude, and Grok simultaneously down?

#702

Neither of these companies have stellar uptime records. Their downtime episodes overlapped in this instance. In this case, it was a partial downtime for both. Also, OpenAI is saying what caused it: > "A routing error starting around 7:43 am PT on Thursday, September 3, made ChatGPT and Codex unavailable for some users across platforms" Anthropic stated their issue started earlier: > "The company began alerting about…

In my experience I've seen plenty of failures caused by user behaviour in these type of cases.

Biggest competitor goes down and all of a sudden you have a lot more traffic...

Re: Ask HN: Why were OpenAI, Claude, and Grok simultaneously down?

#703
post #683

Earlier quoted context omitted.

Worth mentioning that this room takes a split from the main feed and is not in the path of traffic. Whatever is in this room could go down and it would not cause an outage.

And if the hypothetical splitter is the thing that breaks? Then it would break main traffic too.

Interestingly, fiber signals can be split passively[1] (without a transducer), which should be extremely reliable. No idea what technology this kind of application would use though.

[1] https://en.wikipedia.org/wiki/Fiber-optic_splitter

Re: Ask HN: Why were OpenAI, Claude, and Grok simultaneously down?

#704

Earlier quoted context omitted.

What makes you think they'd need to touch every datacenter? All of these endpoints use existing providers with decades-long history at this point, and network monitoring is already a proven 'feature' of the agencies they'd need to co-exist with over their lifetimes. If anything, Occam's Razor would point to a common denominator with all of them, given it wasn't network-wide, as far as i know.

Explain the system in which you could capture all of these chats with no knowledge of anyone in these data centers. How are they routed to this NSA system or through some NSA device when these companies' compute are spread over hundreds of data centers? > All of these endpoints use existing providers with decades-long history at this point That is just factually inaccurate. Their data centers aren't old and they leas…

>Explain the system in which you could capture all of these chats with no knowledge of anyone in these data centers.

They all transit the same wires as all other traffic. Copy them at any regional bottleneck. https://en.wikipedia.org/wiki/Room_641A. Additionally, i'd admit that maybe someone(s) at these companies knows. But if we think there isn't any person who would agree to do this then I think we're being naive.

> Their data centers aren't old and they lease a lot of compute...

Again, they transit the same wires as everyone else. Here i'll also add that these companies have been actively courting government relationships (and Anthropic attempting to repair damaged ones), why would they stand on principles here and not any of the other many frontlines they've visibly acquiesced?

I just think it's easier to re-route their traffic than, as you say, touch every single datacenter and its employees in some way.

Re: Ask HN: Why were OpenAI, Claude, and Grok simultaneously down?

#705

It’s probably the thing that everyone thinks it is. OpenAI, Anthropic and SpaceXAI are all routed through something that we’re not supposed to know exists and that thing had a whoopsie.

it's also possible one has an outage, routing extraordinary traffic to the other(s), with cascading failures in quick succession.

Re: Ask HN: Why were OpenAI, Claude, and Grok simultaneously down?

#706

Earlier quoted context omitted.

"AI Traffic" is just traffic. If the infrastructure exists to monitor/buffer traffic (it does) then this can be monitored as well. Whether this hiccup was due to them hitting their limits briefly (or turning it on, or etc) who knows.

I bet it's even negligible traffic compared to e.g. Netflix or YouTube. Even a 'huge' context window is nothing compared to the random library of a basic Web page.

Not once you de-dup.

Re: Ask HN: Why were OpenAI, Claude, and Grok simultaneously down?

#707

I work at OpenAI and I was the Incident Commander for yesterday's outage. We had a routing error within our infra that caused issues for some of our products. It was not related to the Astra launch. We don't comment on other providers' outages.

Straight from the turkey's beak.

Re: Ask HN: Why were OpenAI, Claude, and Grok simultaneously down?

#708

Earlier quoted context omitted.

On the contrary, it lets you MITM encrypted communications by swapping the website's original certificate for the "well trusted TLS cert"

No, it doesn't. HSTS and other methods prevent this from happening.

HSTS doesn't protect you from this at all. It only requires HTTPS, which a spoofed-but-trusted cert passes just fine.

No mainstream browser (or any browser?) is doing cert pinning.

What "other methods" are there that are deployed and actually in use?

Re: Ask HN: Why were OpenAI, Claude, and Grok simultaneously down?

#709

Neither of these companies have stellar uptime records. Their downtime episodes overlapped in this instance. In this case, it was a partial downtime for both. Also, OpenAI is saying what caused it: > "A routing error starting around 7:43 am PT on Thursday, September 3, made ChatGPT and Codex unavailable for some users across platforms" Anthropic stated their issue started earlier: > "The company began alerting about…

[deleted]

Re: Ask HN: Why were OpenAI, Claude, and Grok simultaneously down?

#710

It’s probably the thing that everyone thinks it is. OpenAI, Anthropic and SpaceXAI are all routed through something that we’re not supposed to know exists and that thing had a whoopsie.

Wouldn't be surprised: Snowden's revelations 10 plus years ago already showed how the NSA was injected into the data centers of Social Media, it's only logical that they would now demand to be injected into the biggest, most information providing data stream of the planet of the present: LLM services.

[deleted]
Post reply on HN