Live data from Hacker News

Update about the October 4th outage

engineering.fb.com

121–130 of 239 posts

Re: Update about the October 4th outage

#121

Earlier quoted context omitted.

We almost went down the ‘this is a subterfuge to delete whistleblower evidence’ rabbit hole.

I saw a couple of people clearly guessing something along these lines but none of them seemed to be claiming that it was actually happening, more like “isn’t it convenient that…”

The timing was uncanny. I still don't see a reason why it couldn't have been an intrusion/rogue employee? Like someone had access to a system to push router firmware updates or something?

Re: Update about the October 4th outage

#122
post #106
post #18

Earlier quoted context omitted.

BGP has to converge to a single routing table. You are effectively asking is why is there a single routing table for the internet. To put in simple terms having a single routing table is what it makes it the internet we can share, otherwise it would just be a bunch of independent networks.

> BGP has to converge to a single routing table. It certainly does not. If I peer with you, neither of us (generally) announce that route to our other peers, but often announce to our customers. There are many routes that are not visible to everyone, and there is no single routing table for the internet. Each BGP speaker ends up with their own routing table, although there are a lot of similarities.

I was trying to give simple explanation for someone who said they don't know networking.

BGP in this context implicitly meant external BGP . Yes, no single router necessarily sees all the routes, but all routers combined generally see the internet as a single network of networks was my point.

Convergence in this context is how most ASN will resolve on where/how to route a specific ASN traffic.

It is hard to peer with someone and not trust the routing table they publish, that is why Pakistan could by mistake block YouTube for everyone few years back.

In this case if you peered with Facebook, and they published incorrect routing for their AS you would accept it.

This doesn't mean FB couldn't have used multiple ASNs did some rolling updates etc, however without knowing what exactly fb screwed up for five hours it is hard to say what they could done differently.

Re: Update about the October 4th outage

#123
post #87

It just occurred to me to wonder if Facebook has a Twitter account and if they used it to update people about the outage. It turns out they do, and they did, which makes sense. Boy, it must have been galling to have to use a competing communication network to tell people that your network is down. It looks like Zuckerberg doesn't have a personal Twitter though, nor does Jack Dorsey have a public Facebook page (or the…

> It looks like Zuckerberg doesn't have a personal Twitter though He does: https://twitter.com/finkd

Ah, I saw that one, but it wasn't verified so I figured it was an imposter. It has only a handful of tweets from 2009 and 1 from 2012, but it could really be him, I suppose.

Re: Update about the October 4th outage

#124

On a side note: when I browse to that page in Firefox (92.0.1) from HN I can't go back to HN - the back arrow is disabled. What gives?

Do you have the facebook container extension? That closes the current tab, opens a new tab with a container, then goes to the facebook link. Reopening the last closed tab works for me, although I haven't noticed this before since I always open links in a new tab.

Re: Update about the October 4th outage

#125
post #60

Earlier quoted context omitted.

There seems to be nothing uncertain about the immediate cause of the issue - Facebook revoked all of their BGP routes, and all of their IP addresses couldn't receive packets until they were restored.

The didn't revoke all their routes, FWIW, just a lot of them (including the anycast DNS routes)

What I don't understand is why, when a route is revoked, if there is no other route announced the routing table gets updated? It seems like either it's a black hole or it still works and there was a BGP error (or the route works but the resources aren't present, so traffic would be dropped). What's the reason for designing the system to revoke routes when no new route is announced?

It strikes me it's like DNS when you get a SERVFAIL, why not try the prior IP address. The similarity in the design here suggests there may be common reasoning??

Re: Update about the October 4th outage

#126
post #87

It just occurred to me to wonder if Facebook has a Twitter account and if they used it to update people about the outage. It turns out they do, and they did, which makes sense. Boy, it must have been galling to have to use a competing communication network to tell people that your network is down. It looks like Zuckerberg doesn't have a personal Twitter though, nor does Jack Dorsey have a public Facebook page (or the…

> It looks like Zuckerberg doesn't have a personal Twitter though He does: https://twitter.com/finkd

His LinkedIn photo used to be this really awkward laptop camera photo of roughly this face: (-_-)

It was amazing. I’m sad he remove it.

Re: Update about the October 4th outage

#127
post #41

Earlier quoted context omitted.

- It was government intervention - Facebook was hacked - They did it on purpose to bury the whistleblower story - No one could access Facebook offices - They had to cut open servers with angle grinders - Disgruntled employees changed DNS records - Lots of made up numbers for how much money Facebook/the rest of the economy was losing (or gaining) They probably rushed out this blog post just to dispel some of these rum…

> - It was government intervention There was an Indian opposition Member of Parliament blaming the current government that it blocked FB due to some protests being held in the Capital.

Not to defend politicians looking to gain political mileage from everything, this is not a far-fetched claim. Current Indian government has blocked websites and shutdown internet access wholesale to towns and even whole states citing many reasons, and multiple times in recent past. Latest was internet shutdown in multiple Rajasthan towns apparently to prevent exam takers from cheating on a test.

Re: Update about the October 4th outage

#129
post #87

It just occurred to me to wonder if Facebook has a Twitter account and if they used it to update people about the outage. It turns out they do, and they did, which makes sense. Boy, it must have been galling to have to use a competing communication network to tell people that your network is down. It looks like Zuckerberg doesn't have a personal Twitter though, nor does Jack Dorsey have a public Facebook page (or the…

I’m not sure they are competing though. They serve different purposes and co-exist pretty well together.

Re: Update about the October 4th outage

#130

Earlier quoted context omitted.

I saw a couple of people clearly guessing something along these lines but none of them seemed to be claiming that it was actually happening, more like “isn’t it convenient that…”

The timing was uncanny. I still don't see a reason why it couldn't have been an intrusion/rogue employee? Like someone had access to a system to push router firmware updates or something?

A rogue employee would have been very easy to detect and that employee would have known this. The core network infrastructure involved is extremely sensitive and is quite unlikely to be accessible in a break in.

Also IIRC an employee was posting on Reddit saying the incident started shortly after a network update was posted this morning.

When you know more about the tech and systems involved a mistake seems infinitely more likely than sabotage.

Post reply on HN