Live data from Hacker News

Ask HN: Why were OpenAI, Claude, and Grok simultaneously down?

news.ycombinator.com

721–730 of 776 posts

Re: Ask HN: Why were OpenAI, Claude, and Grok simultaneously down?

#721

I work at OpenAI and I was the Incident Commander for yesterday's outage. We had a routing error within our infra that caused issues for some of our products. It was not related to the Astra launch. We don't comment on other providers' outages.

Did you negotiate extra hard for the 'Incident Commander' title? I'm a bit jealous to be honest.

Ha. It's a role/title for the lifetime of the incident -- it's useful to have someone to keep things moving, keep track of workstreams, and to know who the decision-maker is, especially for bigger incidents. I'm just a SWE who works on infrastructure.

Re: Ask HN: Why were OpenAI, Claude, and Grok simultaneously down?

#723

I thought the consensus on here yesterday was that it was likely caused by cascading failures. OpenAI had an issue during their GPT-6 rollout, taking down their service. This caused a lot of OpenAI users to push their requests (or a larger share of their requests) to Claude and/or Grok, which pushed their load high enough to cause outages. We used to experience similar effects when I worked at a CDN. If one CDN would…

[dead]

Re: Ask HN: Why were OpenAI, Claude, and Grok simultaneously down?

#724
post #683

Earlier quoted context omitted.

And if the hypothetical splitter is the thing that breaks? Then it would break main traffic too.

Interestingly, fiber signals can be split passively[1] (without a transducer), which should be extremely reliable. No idea what technology this kind of application would use though. [1] https://en.wikipedia.org/wiki/Fiber-optic_splitter

We don't know the architecture of this hypothetical spy splitter tho in the current case. It can be less covert and more of a complicated config, etc.

Re: Ask HN: Why were OpenAI, Claude, and Grok simultaneously down?

#725

It’s probably the thing that everyone thinks it is. OpenAI, Anthropic and SpaceXAI are all routed through something that we’re not supposed to know exists and that thing had a whoopsie.

Very likely yes. I wouldn't be surprised if they were hosted from the same datacenters even. There has been a story every few weeks about how Musk has sublet X.ai capacity for one company or another.

This whole thing makes me thing about a passage in Dune where they mentioned the Spacing Guild transported entire fleets of ships in isolated compartments and leaving said compartments was a capital offense. This way, entire militaries of mortal enemies were shipped to battlefield, with nothing but bulkheads separating each other.

Re: Ask HN: Why were OpenAI, Claude, and Grok simultaneously down?

#727

It’s probably the thing that everyone thinks it is. OpenAI, Anthropic and SpaceXAI are all routed through something that we’re not supposed to know exists and that thing had a whoopsie.

OpenAI and Anthropic have their systems in dozens of data centers, including using compute from the major cloud providers. Are you implying that all of these data centers (and many of their employees) are involved in helping the the US government secretly tap every single AI conversation by routing them through some unknown network/device? Or did a couple of companies with poor uptime records happen to have overlappi…

Normally there are only a handful of employees on the payroll at each major company that exposes the US or US government to risk.

It is not often the Executives or Legal even know, but sometimes they did. AT&T bent over backwards to help.

This is standard behavior by the CIA and NSA, and has been for a long time.

https://www.propublica.org/article/nsa-documents-suggest-clo...

https://www.nytimes.com/2015/08/16/us/politics/att-helped-ns...

https://www.theguardian.com/world/2014/mar/19/us-tech-giants...

https://theintercept.com/2018/06/25/att-internet-nsa-spy-hub...

Re: Ask HN: Why were OpenAI, Claude, and Grok simultaneously down?

#728

It’s probably the thing that everyone thinks it is. OpenAI, Anthropic and SpaceXAI are all routed through something that we’re not supposed to know exists and that thing had a whoopsie.

Wouldn't be surprised: Snowden's revelations 10 plus years ago already showed how the NSA was injected into the data centers of Social Media, it's only logical that they would now demand to be injected into the biggest, most information providing data stream of the planet of the present: LLM services.

I read Nowhere to Hide recently, really worth it if you can get past Greenwald sticking himself in the middle (start halfway through).

The stuff in there is horrifying, and incredibly cute compared to what's possible now. The bottleneck back then would have been analysis, trivial now.

Everyone in the world, especially our leaders, sit under a colossal, omniscient blackmail machine. I don't believe democracy can exist under these conditions.

Re: Ask HN: Why were OpenAI, Claude, and Grok simultaneously down?

#729
post #677

Earlier quoted context omitted.

https://en.wikipedia.org/wiki/Room_641A

Traffic doesn’t need to be routed through it though, just tee’d to it,

In the context of arbitrary line-level traffic, sure. But we're talking about robust reverse proxies here if this _is_ what's going on, again, not a fiber tap inside a closet. In that case, there's no such thing as tee'ing to it.

Re: Ask HN: Why were OpenAI, Claude, and Grok simultaneously down?

#730

It’s probably the thing that everyone thinks it is. OpenAI, Anthropic and SpaceXAI are all routed through something that we’re not supposed to know exists and that thing had a whoopsie.

The old PRISM servers got overloaded
Post reply on HN