Live data from Hacker News

Update about the October 4th outage

engineering.fb.com

211–220 of 239 posts

Re: Update about the October 4th outage

#212
post #131

Gotta love how painfully vague this is. Sounds like a PR piece for investors, not an engineering blog piece.

It's less vague than you realize. It points out that the problem was within Facebook's network between its datacenters. This not only suggests it's related to express backbone, but also suggests that the DNS BGP withdrawal which Cloudflare observed was not the primary issue. It's not a full root cause analysis, to be sure, and leaves many open questions, but I definitely wouldn't describe it as painfully vague.

A point of distinction - there is no "DNS BGP withdrawl".

DNS is related to BGP only that without the right BGP routes in the routers, no packets can get to the facebook networks and thus the facbook DNS servers.

That their DNS servers were taken out was a side affect of the root issue - they withdrew all the routes to their networks from the rest of the Internet.

Not picking on you - but there has been a lot of confusion around DNS that is mostly a red herring and people should just drop it from the conversation. Everything on facebook networks disappeared, not just DNS. The main issue is they effectively took a pair of scissors to every one of their internet connections - d'oh!

Re: Update about the October 4th outage

#213

So this is pure conspiracy theory, but to me this could be a security issue. What if something deep in the core of your infrastructure is compromised? Everything at risk? Id ask my best engineer, hed suggest to shut it down, and the best way to do that is to literally pull the plug on what makes you public. Tell everyone we accidentally messed up a BGP and thats it. But yeah, likely not.

Many have pointed out that a couple of weeks ago Facebook had a paper out on how they had implemented a fancy new automated system to manage their BGP routes.

Whoops! Never attribute to malice that which can more easily be explained by stupidity and all that.

Re: Update about the October 4th outage

#214
post #69

Earlier quoted context omitted.

A New York Times reporter said employees had trouble getting in, and I don't see a retraction. So I believe that is true: https://twitter.com/sheeraf/status/1445099150316503057

Do we hold journalists posting on Twitter to the same standards as writing in newspapers and their websites? Genuine question

Ha - journalists don't seem to have any standards for their websites so why would twitter be any different?

Re: Update about the October 4th outage

#215

Earlier quoted context omitted.

I saw a couple of people clearly guessing something along these lines but none of them seemed to be claiming that it was actually happening, more like “isn’t it convenient that…”

The timing was uncanny. I still don't see a reason why it couldn't have been an intrusion/rogue employee? Like someone had access to a system to push router firmware updates or something?

Many have pointed out a couple of weeks ago Facebook released a paper talking about a new system to automate the management of the their BGP routing.

Seems like the new system having an unanticipated flaw is a far more likely scenario than a malicious actor.

More boring - but usually the boring stuff is the far more reasonable.

Re: Update about the October 4th outage

#216
post #163
post #161

Earlier quoted context omitted.

On the off-chance this isn't sarcasm, Facebook's routing shenanigans slowed down the entire internet . Not to mention that they're a publicly traded company, and one which has gone out of its way to assume an infrastructure role. They don't have a right to privacy here, and we are all owed an explanation.

I don't want an explanation nor do I care, Facebook could disappear tomorrow like all the other networks before it and it wouldn't make a dent in my day.

Whether we like it or not, all three platforms are relied upon by hundreds of millions of people and businesses everyday for communication.

I'm sure the world would quickly adapt by re-adopting these things called "websites" and "email" but in the meantime, it's highly self-centered to think this "didn't matter".

Re: Update about the October 4th outage

#217
post #41

Earlier quoted context omitted.

I didn’t see any disinformation, just initial reports that it was DNS which were later explained to be caused by BGP.

- It was government intervention - Facebook was hacked - They did it on purpose to bury the whistleblower story - No one could access Facebook offices - They had to cut open servers with angle grinders - Disgruntled employees changed DNS records - Lots of made up numbers for how much money Facebook/the rest of the economy was losing (or gaining) They probably rushed out this blog post just to dispel some of these rum…

You say "believed as fact" where I say "speculated because we had nothing to go on".

Sure, my friends and I wondered if it was a malicious insider. It takes surprisingly few people in an organization to cause chaos.

Knowing IoT, it isn't unbelievable that badge readers could be offline.

Knowing division of duties, it isn't hard to believe that the network engineers, domain admins and datacenter ops people may have hustled to a DC to get things back online.

Never did I see a large number of people take anything as fact that didn't seem to be substantiated.

Re: Update about the October 4th outage

#219
post #102

Earlier quoted context omitted.

For a status page to be actually independent, it needs to have all it's requirements hosted on other infrastructure. fb.com authoritative DNS is the same as facebook.com, so it's going down when (FB) DNS goes down (and DNS is going down when BGP is broken, apparently). It looks like the status page is hosted on CloudFront though, so it got part of the way. (Of course, the other question is if it was updatable / updat…

Pardon me if it's a stupid question, but out of curiosity: Is there any way to keep DNS up in case BGP goes down for any reason? Like a fallback nameserver hosted elsewhere/not affected by Facebook's ASs? Is it technically impossible or did Facebook just assume something like yesterday would never happen and kept things simple instead of complicating things?

BGP didn't "go down" - they erroneously removed all routes between the Internet and several facebook internal networks via BGP. BGP was the instrument of their destruction, but not the source. Someone or something told BGP to do that; whatever that was is the cause of the issue.

At least one of those networks they accidentally removed also happened to contain the DNS servers; DNS being unavailable was a symptom - but not part of the root problem. Any focus on DNS at this point is a red herring.

Think of routes as street directions - they tell routers where to ship packets. If you erase all your addresses and directions to them from the outside world at at large, then there literally is no way for network packets to get from the global Internet to Facebooks networks (where I imagine the DNS servers were up and probably twiddling their thumbs wondering where everyone went).

An easier way to think of it - they essentially took a pair of scissors and cut the cable connections to the Internet - which is why it was so catastrophic.

They only way to mitigate that is to have an identical infrastructure managed by different tooling so a bad configuration setting from one environment wouldn't pollute the second in the same way. Not exactly an easy thing to do and might cause more other problems than it's worth. And you would have to do that for all services, not just DNS. Let's say Facebook used Cloudflare for their DNS. Great - DNS can resolve your request for fb.com to the IP address of the facebook datacenter - there still is no path for your packets to get to that facebook datacenter because they accidentally purged the routes to their networks.

It's easier to just not cut your connection to the Internet :) I'm sure there are all kinds of internal discussions picking this incident apart and formulating ways to either prevent it, or more realistically - have improved procedures to speed recovery when it inevitably happens again. BGP is not known for its inherent robustness or security. But since it's at the core of the Internet, any changes to it would have to be done on a massive internet-wide scale in perfect unison or the "cure" would be a lot worse than the current problems with it.

Murphy was indeed an optimist! (search "Murphy's Law" for those unfamiliar with the idiom)

Re: Update about the October 4th outage

#220
post #165

The mobile whatsapp app should notify that the whatsapp servers are down and not allow you to just send messages that won't arrive for six hours

The app is designed under the assumption that Facebook servers are never down. If you can't reach the servers, the problem is assumed to be client-side, in which case they have decided the best UI is to keep retrying (not unreasonably in a mobile context). The only way to disambiguate "no internet service" (extremely common) with "Facebook dropped off the internet" (black-swan rare) is to ping some other, third party…

>The app is designed under the assumption that Facebook servers are never down.

Which was and is a lame assumption. Stuff happens. SMTP wouldn't even be phased by this; it would just pick up where it left off.

I've seen far too many applications fail in bizzare ways because people make unrealistic assumptions like "X will ALWAYS be there". Sure it's highly unlikely, but when you have multiple things making the same dumb assumptions, on the inevitable day when multiple things that need X and X is suddenly no longer there then you start to get cascading effects of Y that relied on something that relied on X not being there when it is assumed that it would always be there so now Y fails, and then something dependent in the same way on Y unexpectedly fails and so on.

One should never assume that anything will "always" be available. That's an incredibly unrealistic assumption; and the more interconnected things become, the chances of these really nasty dependency chains/cascade failures skyrocket - leading to far worse outages and longer recovery times.

Post reply on HN