Live data from Hacker News

Ask HN: Why were OpenAI, Claude, and Grok simultaneously down?

news.ycombinator.com

771–780 of 794 posts

Re: Ask HN: Why were OpenAI, Claude, and Grok simultaneously down?

#771

I work at OpenAI and I was the Incident Commander for yesterday's outage. We had a routing error within our infra that caused issues for some of our products. It was not related to the Astra launch. We don't comment on other providers' outages.

Did you negotiate extra hard for the 'Incident Commander' title? I'm a bit jealous to be honest.

Note: Incident Command is almost certain an allusion (or implementation) of the Incident Command System [1]

It is a common system in all kinds of emergency response scenarios, including local emergency services (fire/police/ambulance) and it scales all the way to massive disasters.

It is especially useful to clarify command structures when multiple response entities need to coordinate. That is true even within organizations like public companies, where the reporting structures may be distinct.

1. https://en.wikipedia.org/wiki/Incident_Command_System

Re: Ask HN: Why were OpenAI, Claude, and Grok simultaneously down?

#772

Earlier quoted context omitted.

OpenAI and Anthropic have their systems in dozens of data centers, including using compute from the major cloud providers. Are you implying that all of these data centers (and many of their employees) are involved in helping the the US government secretly tap every single AI conversation by routing them through some unknown network/device? Or did a couple of companies with poor uptime records happen to have overlappi…

Ok, then where's the report? Don't they usually release a retrospective report after outages?

It happened yesterday.

Re: Ask HN: Why were OpenAI, Claude, and Grok simultaneously down?

#773

Earlier quoted context omitted.

Explain the system in which you could capture all of these chats with no knowledge of anyone in these data centers. How are they routed to this NSA system or through some NSA device when these companies' compute are spread over hundreds of data centers? > All of these endpoints use existing providers with decades-long history at this point That is just factually inaccurate. Their data centers aren't old and they leas…

>Explain the system in which you could capture all of these chats with no knowledge of anyone in these data centers. They all transit the same wires as all other traffic. Copy them at any regional bottleneck. https://en.wikipedia.org/wiki/Room_641A . Additionally, i'd admit that maybe someone(s) at these companies knows. But if we think there isn't any person who would agree to do this then I think we're being naive.…

> They all transit the same wires as all other traffic. Copy them at any regional bottleneck.

Oh, is that all?

> Again, they transit the same wires as everyone else

Which shared wires does Google's traffic go across? Do they share that network with others or do they not function as a Tier 1 network? How about Amazon? And the NSA is doing deep packet inspection of all the traffic going to these places to pull out the data they are interested in? What systems allow them to do this at the scale that would be necessary to accomplish this task? Anthropic, OpenAI, etc have their models hosted and provided by these and other Internet giants. So you now have to be able to somehow also get access to every "regional bottleneck" that those places have. Google has 40+ regions alone, as does Amazon. And each of these regions don't just have a single point of inbound/outbound traffic, so now you're over 200-300 points of interception that you would need just to grab the data you're suggesting they are.

And that's just the start of it. When I use Vertex or the Bedrock, my AI API requests from my instances doesn't leave their networks. So where does that traffic get intercepted?

> I just think it's easier to re-route their traffic than, as you say, touch every single datacenter and its employees in some way.

That would be really, really loud and really obvious. Messing with routes like that is very easily detectable.

Or, maybe, they don't care about this traffic at all and doing all of the huge amount of work necessary to accomplish what you're proposing isn't worth 0.1% of the effort it would require.

If the government needs the chat messages, the companies are already saving them. They can just request them through a variety of legal means. None of this vast conspiracy nonsense is needed. Just a couple lawyers and a willing judge. The NSA stopped their collection of phone metadata because they can just request what they need from the telecoms directly.

Re: Ask HN: Why were OpenAI, Claude, and Grok simultaneously down?

#774

Earlier quoted context omitted.

OpenAI and Anthropic have their systems in dozens of data centers, including using compute from the major cloud providers. Are you implying that all of these data centers (and many of their employees) are involved in helping the the US government secretly tap every single AI conversation by routing them through some unknown network/device? Or did a couple of companies with poor uptime records happen to have overlappi…

I don't think you're a sysadmin - because what you're saying really doesn't matter. It can still all fail at a single point.

I'm confused by your statement. Are you suggesting that these companies that have invested tens of billions in their networks and datacenters decided to create a shared single point of failure?

Re: Ask HN: Why were OpenAI, Claude, and Grok simultaneously down?

#775

Neither of these companies have stellar uptime records. Their downtime episodes overlapped in this instance. In this case, it was a partial downtime for both. Also, OpenAI is saying what caused it: > "A routing error starting around 7:43 am PT on Thursday, September 3, made ChatGPT and Codex unavailable for some users across platforms" Anthropic stated their issue started earlier: > "The company began alerting about…

Do you know the probability of all these companies being down at precisely the same time?

"All of these" is two. OpenAI had a router issue. Anthropic had a separate issue. Anthropic uses a lot of SpaceX compute, so an Anthropic issue and a SpaceX issue can be one in the same, as was likely the case this time.

And they weren't down at precisely the same time. Anthropic's issue started ~1 hour before OpenAI's.

Re: Ask HN: Why were OpenAI, Claude, and Grok simultaneously down?

#776
post #334

Cloudflare, Azure, AWS, and Google Cloud all have a similar uptick in reported errors around 7:30. I suspect an outage on Cloudflare or another load-bearing service cascaded through all the major cloud providers. https://downdetector.com/status/cloudflare/ https://downdetector.com/status/windows-azure/ https://downdetector.com/status/aws-amazon-web-services/ https://downdetector.com/status/google-cloud/

The internet is not supposed to work like this. The network was designed for robustness and fault tolerance, which allows it to reroute data if parts of the network fail. Why are we all depending on one entity for it all to work? Makes me mad.

What year did the internet routing around failure work?

Re: Ask HN: Why were OpenAI, Claude, and Grok simultaneously down?

#777
post #760

Earlier quoted context omitted.

Why would you default to that explanation? That’s not at all a reasonable default, and I say that as something pretty paranoid with regards to US surveillance

Looking at history, it's much more reasonable to assume there's surveillance, since there are whole branches of the government, and departments, that exist for surveillance in the goal of "national security", which this easily falls into. See Marissa Mayer explaining that it's not an option to refuse [1]. I assume this is just the same ole' Room 641A [2]. [1] https://www.cnet.com/tech/services-and-software/yahoo-repo…

We already know there is surveillance, but why would you default to that explanation for a downtime? As said by others monitoring the traffic isn’t done by routing through the NSA monitoring system, it’s not a bottleneck that would take down all those services. If you look at more details such as the timing it’s even less likely to be the case, but even as a default it’s not a reasonable explanation given the symptoms

Re: Ask HN: Why were OpenAI, Claude, and Grok simultaneously down?

#779
post #764

Earlier quoted context omitted.

I read Nowhere to Hide recently, really worth it if you can get past Greenwald sticking himself in the middle (start halfway through). The stuff in there is horrifying, and incredibly cute compared to what's possible now. The bottleneck back then would have been analysis, trivial now. Everyone in the world, especially our leaders, sit under a colossal, omniscient blackmail machine. I don't believe democracy can exist…

So following this to the obvious conclusion, the NSA is responsible for the closure of the Straight of America, high tarrifs, dropping employment, and high gas prices?

As part of a military industrial complex which spans the Five Eyes and Israel, yes.

Re: Ask HN: Why were OpenAI, Claude, and Grok simultaneously down?

#780
post #585
post #562

Earlier quoted context omitted.

I did, at least until Fable 5.1. Grok’s models are excellent, amazing price-performance and speed too.

All you have to worry about is whether or not it’ll output kiddie porn or racist vitriol

SOMETHING SOMETHING HITLER SOMETHING GROK
Post reply on HN