Live data from Hacker News

Update about the October 4th outage

engineering.fb.com

171–180 of 239 posts

Re: Update about the October 4th outage

#171
post #163
post #161

Earlier quoted context omitted.

On the off-chance this isn't sarcasm, Facebook's routing shenanigans slowed down the entire internet . Not to mention that they're a publicly traded company, and one which has gone out of its way to assume an infrastructure role. They don't have a right to privacy here, and we are all owed an explanation.

I don't want an explanation nor do I care, Facebook could disappear tomorrow like all the other networks before it and it wouldn't make a dent in my day.

Doe you honestly believe that Facebook and its subsidiaries don't have a major impact on the world?

Re: Update about the October 4th outage

#172
post #41

Earlier quoted context omitted.

I didn’t see any disinformation, just initial reports that it was DNS which were later explained to be caused by BGP.

- It was government intervention - Facebook was hacked - They did it on purpose to bury the whistleblower story - No one could access Facebook offices - They had to cut open servers with angle grinders - Disgruntled employees changed DNS records - Lots of made up numbers for how much money Facebook/the rest of the economy was losing (or gaining) They probably rushed out this blog post just to dispel some of these rum…

It could still be any/some of these. Unlikely, but possible. Taking down BGP would just be a cover-up of something bigger. If this was a more honest company I'd call myself a tinfoiler, but since we're dealing with Facebook... ¯\_(ツ)_/¯

Re: Update about the October 4th outage

#173
post #165

The mobile whatsapp app should notify that the whatsapp servers are down and not allow you to just send messages that won't arrive for six hours

The app is designed under the assumption that Facebook servers are never down. If you can't reach the servers, the problem is assumed to be client-side, in which case they have decided the best UI is to keep retrying (not unreasonably in a mobile context). The only way to disambiguate "no internet service" (extremely common) with "Facebook dropped off the internet" (black-swan rare) is to ping some other, third party…

Good points. Then it should just tell the user that they appear to have no internet service.

Re: Update about the October 4th outage

#174
post #123

Earlier quoted context omitted.

> It looks like Zuckerberg doesn't have a personal Twitter though He does: https://twitter.com/finkd

Ah, I saw that one, but it wasn't verified so I figured it was an imposter. It has only a handful of tweets from 2009 and 1 from 2012, but it could really be him, I suppose.

Yeah, that's kinda sus.

Re: Update about the October 4th outage

#175

I worked with a network engineer who misconfigured a router that was connecting a bank to it's DR site. The engineer had to drive across town to manually patch into the router to fix it. DR downtime was about an hour, but the bank fired him anyway. Given that Zuck lost a substantial amount of money, I wonder if the engineer faced any ramifications. Sidenote: I asked the bank infrastructure team why the DR site was in…

"I worked with a network engineer who misconfigured a router that was connecting a bank to it's DR site. The engineer had to drive across town to manually patch into the router to fix it. DR downtime was about an hour, but the bank fired him anyway." so prod wasn't down and he fixed it in a hour and they fired the guy who knew how to fix such things so quickly. Idiot manager at the bank.

Re: Update about the October 4th outage

#176
post #161

Earlier quoted context omitted.

It’s not reasonable to demand any details at all, it’s nice of them to notify people of what went wrong but it really is none of our business.

On the off-chance this isn't sarcasm, Facebook's routing shenanigans slowed down the entire internet . Not to mention that they're a publicly traded company, and one which has gone out of its way to assume an infrastructure role. They don't have a right to privacy here, and we are all owed an explanation.

Or else? You will angrily stamp your foot? Start an e-petition?

Re: Update about the October 4th outage

#177
post #102

Earlier quoted context omitted.

For a status page to be actually independent, it needs to have all it's requirements hosted on other infrastructure. fb.com authoritative DNS is the same as facebook.com, so it's going down when (FB) DNS goes down (and DNS is going down when BGP is broken, apparently). It looks like the status page is hosted on CloudFront though, so it got part of the way. (Of course, the other question is if it was updatable / updat…

Pardon me if it's a stupid question, but out of curiosity: Is there any way to keep DNS up in case BGP goes down for any reason? Like a fallback nameserver hosted elsewhere/not affected by Facebook's ASs? Is it technically impossible or did Facebook just assume something like yesterday would never happen and kept things simple instead of complicating things?

It’s definitely technically possible to have secondary’s on a separate network that do zone axfr from the primary. That’s not to imply it’s trivial / easy at FB’s scale (query volume) or topology complexity (as in GSLB).

Re: Update about the October 4th outage

#178
post #85

Earlier quoted context omitted.

I work in video games. It amazing how wrong people can be and how confident they are about being right. Even sometimes fighting _me_ about things _I_ designed and built. It’s quite sobering; taught me not to believe all the speculation I read.

Just out of curiosity, how often do people claim that your random number generator is broken? And then when you ask why they think that, it's an anecdote about some time when they had incredibly bad luck?

it's more common for people to tell me the tickrate/framerate of my server or talk about how "[I] moved away from amazon to cheaper bare metal servers for cost savings" and such.

Re: Update about the October 4th outage

#179

I worked with a network engineer who misconfigured a router that was connecting a bank to it's DR site. The engineer had to drive across town to manually patch into the router to fix it. DR downtime was about an hour, but the bank fired him anyway. Given that Zuck lost a substantial amount of money, I wonder if the engineer faced any ramifications. Sidenote: I asked the bank infrastructure team why the DR site was in…

>DR downtime was about an hour, but the bank fired him anyway

The US, not even once.

The guy should have had "reload in 10", an outage window and config review. There must be more to this story than it being a firable offence for causing a P2 outage for an hour.

Re: Update about the October 4th outage

#180

It would be interesting to estimate what dollar value can be ascribed to a x-hour FB outage, both in terms of lost ad revenue for FB itself as well missed conversions/revenue for businesses running ads on FB/IG.

Don’t forget WhatsApp users switching to Signal and possibly never returning
Post reply on HN