Live data from Hacker News

Facebook-owned sites were down

facebook.com

471–480 of 1001 posts

Re: Facebook-owned sites were down

#471

Earlier quoted context omitted.

Just imagine the amount of stress on this people, hope the money really worth it.

It shouldn't be too stressful. Well-managed companies blame processes rather than people, and have systems set up to communicate rapidly when large-scale events occur. It can be sort of exciting, but it's not like there is one person typing at a keyboard with a hundred managers breathing down their neck. These resolutions are collaborative, shared efforts.

As one of the major responders to an incident analogous to this at a different fang... you're high, its still hella stressful.

Re: Facebook-owned sites were down

#473
post #349

Earlier quoted context omitted.

I like how FB decided to send "ramenporn" as their spokesperson.

A particular facet I love of the internet era is journalists reporting serious events while having to use the completely absurd usernames... "A Facebook engineer in the response team, ramenporn..."

[deleted]

Re: Facebook-owned sites were down

#474
post #263
post #242

Unsurprisingly, Oculus is down as well, as are most services for the VR headset. So that's 4 major properties right now.

Can you not use an Oculus headset if FB servers are down? That’s absurd.

The Rift headsets probably still work fine, but the Quest headsets require a FB connection to work.

Re: Facebook-owned sites were down

#475
post #380

Earlier quoted context omitted.

Wondering how Facebook communicates now internally - most of their work streams likely depend on Facebooks systems which are all down. Can engineers and security teams even access prod systems anymore? Like, would "Bastion" hosts be reachable? Wonder if they use Signal and Slack now?

There are various non-FB fallback measures, including IRC as a last-ditch method. The IRC fallback is usually tested once a year for each engineer.

Good planning! Now, where does the IRC server live, and is it currently routable from the internet?

While normally I know the advice is "Don't plan for mistakes not to happen, it's impossible, murphy's law, plan for efficient recovery for mistakes"... when it comes to "literally our entire infrastructure is no longer routable from the internet", I'm not sure there's a great alternative to "don't let that happen. ever." And yet, here facebook is.

Re: Facebook-owned sites were down

#476

Reddit r/Sysadmin user that claims to be on the "Recovery Team" for this ongoing issue: > As many of you know, DNS for FB services has been affected and this is likely a symptom of the actual issue, and that's that BGP peering with Facebook peering routers has gone down, very likely due to a configuration change that went into effect shortly before the outages happened (started roughly 1540 UTC). There are people now…

He just deleted all his updates.

user:

https://old.reddit.com/user/ramenporn

some messages:

* This is a global outage for all FB-related services/infra (source: I'm currently on the recovery/investigation team).

* Will try to provide any important/interesting bits as I see them. There is a ton of stuff flying around right now and like 7 separate discussion channels and video calls.

* Update 1440 UTC: \

    As many of you know, DNS for FB services has been affected and this is likely a symptom of the actual issue, and that's that BGP peering with Facebook peering routers has gone down, very likely due to a configuration change that went into effect shortly before the outages happened (started roughly 1540 UTC).

    There are people now trying to gain access to the peering routers to implement fixes, but the people with physical access is separate from the people with knowledge of how to actually authenticate to the systems and people who know what to actually do, so there is now a logistical challenge with getting all that knowledge unified.

    Part of this is also due to lower staffing in data centers due to pandemic measures.

Re: Facebook-owned sites were down

#477
post #194

Earlier quoted context omitted.

I can confirm, HN, GitHub and Slack are very slow for me as well. Google is very fast, on the other hand.

All running their DNS on AWS. My guess is that AWS is seeing a massive flood of failed and retried DNS requests for facebook properties, similar to what jgrahamc mentions here for Cloudflare: https://twitter.com/jgrahamc/status/1445066136547217413

Is there a "Kessler syndrome" analogue for the internet, where failures beget failures until it's just an impenetrable cloud of fail, forever?

Re: Facebook-owned sites were down

#478

Reddit r/Sysadmin user that claims to be on the "Recovery Team" for this ongoing issue: > As many of you know, DNS for FB services has been affected and this is likely a symptom of the actual issue, and that's that BGP peering with Facebook peering routers has gone down, very likely due to a configuration change that went into effect shortly before the outages happened (started roughly 1540 UTC). There are people now…

Archived version: https://archive.is/QvdmH

Re: Facebook-owned sites were down

#480
post #169

Is it just me or HN also feels kinda laggy?

Yep. I am the developer of HN client HACK for iOS and Android and a bunch of users emailed me asking if my app was broken. Looks like something bigger is afoot.

Best HN client app ever. Thanks for the great work!
Post reply on HN