Live data from Hacker News

Facebook-owned sites were down

facebook.com

791–800 of 1001 posts

Re: Facebook-owned sites were down

#791

There's still no connectivity to Facebook's DNS servers: > traceroute a.ns.facebook.com traceroute to a.ns.facebook.com (129.134.30.12), 30 hops max, 60 byte packets 1 dsldevice.attlocal.net (192.168.1.254) 0.484 ms 0.474 ms 0.422 ms 2 107-131-124-1.lightspeed.sntcca.sbcglobal.net (107.131.124.1) 1.592 ms 1.657 ms 1.607 ms 3 71.148.149.196 (71.148.149.196) 1.676 ms 1.697 ms 1.705 ms 4 12.242.105.110 (12.242.105.110)…

Looks like they misconfigured a web interface that they can't reach anymore now that they're off the net. "anyone have a Cisco console cable lying around?"

The only one they have is serial and the company's one usb-to-serial converter is missing.

Re: Facebook-owned sites were down

#792

If it is an DNS error, why is the .onion site also offline? - https://en.wikipedia.org/wiki/Facebook_onion_address - facebookwkhpilnemxj7asaniu7vnjjbiltxjqhye3mhbshg7kx5tfyd.onion

My guess is that the FB backend also required DNS. The .Onion site isn't backed by a backend built on a onion native stack (is that a thing?)

Re: Facebook-owned sites were down

#794

The timing of this is so rich in irony I can't help but wonder if there is an element of internal sabotage. How many FB employees hate FB right now? The latest expose of FB is both effective and truly awful. I can't imagine feeling good about a FB job. And it's gotten worse! Now they look like they can't even keep their websites up.

Can we really ever know? There are million of $ at stake!

Re: Facebook-owned sites were down

#795

The media coverage and lots of the comments don't make sense to me. FB would not be so stupid and put all of their crucial DNS servers into a single autonomous system (which is now offline due to BGP issues). They operate literally dozens of datacenters around the world, and are surely not using a single AS for them - why not put secondary Nameservers there? Can someone make a sense of this?

Sounds like automation deployed a configuration update to most of Facebook's peering routers simultaneously. Something similar brought down Google in 2019.

Re: Facebook-owned sites were down

#796
post #705

Earlier quoted context omitted.

You mean the same problem as when GMail goes down and Googlers can't reach each other? I guess good decentralized public communication services could solve those issues for everybody.

I can assure you that Google has a procedure in place for that.

Yup, they make a new chat app if the previous one is down.

Re: Facebook-owned sites were down

#797

Earlier quoted context omitted.

My suspicion is that since a lot of internal comms runs through the FB domain and since everyone is still WFH, then its probably a massive issue just to get people talking to each other to solve the problem.

You mean the same problem as when GMail goes down and Googlers can't reach each other? I guess good decentralized public communication services could solve those issues for everybody.

Word is that the last time Google had a failure involving a cyclical dependency they had to rip open a safe. It contained the backup password to the system that stored the safe combination.

Re: Facebook-owned sites were down

#799

Reddit r/Sysadmin user that claims to be on the "Recovery Team" for this ongoing issue: > As many of you know, DNS for FB services has been affected and this is likely a symptom of the actual issue, and that's that BGP peering with Facebook peering routers has gone down, very likely due to a configuration change that went into effect shortly before the outages happened (started roughly 1540 UTC). There are people now…

Just imagine the amount of stress on this people, hope the money really worth it.

This is a one off event, not a chronic stress trigger. I find them envigorating personally, as long as everybody concerned understands that this is not good in the long run, and that you are not going to write your best code this way.

Re: Facebook-owned sites were down

#800

Earlier quoted context omitted.

My suspicion is that since a lot of internal comms runs through the FB domain and since everyone is still WFH, then its probably a massive issue just to get people talking to each other to solve the problem.

I don’t know how true it is but a few reports claim employees can’t get into the building with their badges.

I guess they didn't have an "emergency ingress" plan.
Post reply on HN