Live data from Hacker News

Facebook-owned sites were down

facebook.com

681–690 of 1001 posts

Re: Facebook-owned sites were down

#681

It seems it has caused DNS servers crash for one of biggest Czechia's internet provider - Vodafone. Can be unrelated but I doubt it ( https://twitter.com/BlazejKrajnak/status/1445063232486531099 ). Think of it - half the country doesn't have internet because of this crash, that's terrifying. (Switching DNS servers obviously works but that's not something the general population will do)

Same in the UK, I've just experienced external DNS outage on BT!

Re: Facebook-owned sites were down

#682

There's still no connectivity to Facebook's DNS servers: > traceroute a.ns.facebook.com traceroute to a.ns.facebook.com (129.134.30.12), 30 hops max, 60 byte packets 1 dsldevice.attlocal.net (192.168.1.254) 0.484 ms 0.474 ms 0.422 ms 2 107-131-124-1.lightspeed.sntcca.sbcglobal.net (107.131.124.1) 1.592 ms 1.657 ms 1.607 ms 3 71.148.149.196 (71.148.149.196) 1.676 ms 1.697 ms 1.705 ms 4 12.242.105.110 (12.242.105.110)…

My suspicion is that since a lot of internal comms runs through the FB domain and since everyone is still WFH, then its probably a massive issue just to get people talking to each other to solve the problem.

I don’t know how true it is but a few reports claim employees can’t get into the building with their badges.

Re: Facebook-owned sites were down

#684
post #445

It seems it has caused DNS servers crash for one of biggest Czechia's internet provider - Vodafone. Can be unrelated but I doubt it ( https://twitter.com/BlazejKrajnak/status/1445063232486531099 ). Think of it - half the country doesn't have internet because of this crash, that's terrifying. (Switching DNS servers obviously works but that's not something the general population will do)

If only the news reporting was not as stupid as "internet is not working at UPC", instead of DNS resolvers at UPC crashed, here's what you can do... Anyway, I didn't even notice since I run knot-resolver at home. I wonder what it will be like connecting Facebook back to the internet, thundering herd and everything...

I suspect that the DNS aspect will be fine. The middle DNS servers only need one valid response to cache it for $TTL, but they can't cache SERVFAIL.

Re: Facebook-owned sites were down

#686
post #314

Reddit r/Sysadmin user that claims to be on the "Recovery Team" for this ongoing issue: > As many of you know, DNS for FB services has been affected and this is likely a symptom of the actual issue, and that's that BGP peering with Facebook peering routers has gone down, very likely due to a configuration change that went into effect shortly before the outages happened (started roughly 1540 UTC). There are people now…

> the people with physical access is separate from the people with knowledge of [...] Welcome to the brave new world of troubleshooting. This will seriously bite us one day.

folks with physical access are also denied. source - https://twitter.com/YourAnonOne/status/1445100431181598723

Re: Facebook-owned sites were down

#687
post #653

Earlier quoted context omitted.

Yeah the patch to fix BGP to reach the DNS is sent by email to @facebook.com. Ooops no DNS to resolve the MX records to send the patch to fix the BGP routers.

Seriously? Is that how it works?

No, the backbone of the internet is not maintained with patches sent in emails.

Re: Facebook-owned sites were down

#688

Earlier quoted context omitted.

My suspicion is that since a lot of internal comms runs through the FB domain and since everyone is still WFH, then its probably a massive issue just to get people talking to each other to solve the problem.

You mean the same problem as when GMail goes down and Googlers can't reach each other? I guess good decentralized public communication services could solve those issues for everybody.

I think the issue there is that in exchange for solving the "one fat finger = outage" problem, you lose the ability to update the server fleet quickly or consistently.

Re: Facebook-owned sites were down

#689

Earlier quoted context omitted.

I can't fathom how they didn't plan for this. In any business of size, you have to change configuration remotely on a regular basis, and can easily lock yourself out on a regular basis. Every single system has a local user with a random password that we can hand out for just this kind of circumstance...

> I can't fathom how they didn't plan for this Maybe because they were planning for a million other possible things to go wrong, likely with higher probability than this. And busy with each day's pressing matters.

[deleted]

Re: Facebook-owned sites were down

#690

Earlier quoted context omitted.

Unlikely, PagerDuty was invented for this kind of thing

Oh I'm sure everyone knows whats wrong, but how am I supposed to send an email, find a coworkers phone number, get the crisis team on video chat etc etc if all of those connections rely on the facebook domain existing?

Hence the suggestion for PagerDuty. It handles all this, because responders set their notification methods (phone, SMS, e-mail, and app) in their profiles, so that when in trouble nobody has to ask those questions and just add a person as a responder to the incident.
Post reply on HN