Earlier quoted context omitted.
Just imagine the amount of stress on this people, hope the money really worth it.
It shouldn't be too stressful. Well-managed companies blame processes rather than people, and have systems set up to communicate rapidly when large-scale events occur. It can be sort of exciting, but it's not like there is one person typing at a keyboard with a hundred managers breathing down their neck. These resolutions are collaborative, shared efforts.
Facebook-owned sites were down
471–480 of 1001 posts
Re: Facebook-owned sites were down
#472Re: Facebook-owned sites were down
#473Earlier quoted context omitted.
I like how FB decided to send "ramenporn" as their spokesperson.
A particular facet I love of the internet era is journalists reporting serious events while having to use the completely absurd usernames... "A Facebook engineer in the response team, ramenporn..."
Re: Facebook-owned sites were down
#474Unsurprisingly, Oculus is down as well, as are most services for the VR headset. So that's 4 major properties right now.
Can you not use an Oculus headset if FB servers are down? That’s absurd.
Re: Facebook-owned sites were down
#475Earlier quoted context omitted.
Wondering how Facebook communicates now internally - most of their work streams likely depend on Facebooks systems which are all down. Can engineers and security teams even access prod systems anymore? Like, would "Bastion" hosts be reachable? Wonder if they use Signal and Slack now?
There are various non-FB fallback measures, including IRC as a last-ditch method. The IRC fallback is usually tested once a year for each engineer.
While normally I know the advice is "Don't plan for mistakes not to happen, it's impossible, murphy's law, plan for efficient recovery for mistakes"... when it comes to "literally our entire infrastructure is no longer routable from the internet", I'm not sure there's a great alternative to "don't let that happen. ever." And yet, here facebook is.
Re: Facebook-owned sites were down
#476Reddit r/Sysadmin user that claims to be on the "Recovery Team" for this ongoing issue: > As many of you know, DNS for FB services has been affected and this is likely a symptom of the actual issue, and that's that BGP peering with Facebook peering routers has gone down, very likely due to a configuration change that went into effect shortly before the outages happened (started roughly 1540 UTC). There are people now…
user:
https://old.reddit.com/user/ramenporn
some messages:
* This is a global outage for all FB-related services/infra (source: I'm currently on the recovery/investigation team).
* Will try to provide any important/interesting bits as I see them. There is a ton of stuff flying around right now and like 7 separate discussion channels and video calls.
* Update 1440 UTC: \
As many of you know, DNS for FB services has been affected and this is likely a symptom of the actual issue, and that's that BGP peering with Facebook peering routers has gone down, very likely due to a configuration change that went into effect shortly before the outages happened (started roughly 1540 UTC).
There are people now trying to gain access to the peering routers to implement fixes, but the people with physical access is separate from the people with knowledge of how to actually authenticate to the systems and people who know what to actually do, so there is now a logistical challenge with getting all that knowledge unified.
Part of this is also due to lower staffing in data centers due to pandemic measures.Re: Facebook-owned sites were down
#477Earlier quoted context omitted.
I can confirm, HN, GitHub and Slack are very slow for me as well. Google is very fast, on the other hand.
All running their DNS on AWS. My guess is that AWS is seeing a massive flood of failed and retried DNS requests for facebook properties, similar to what jgrahamc mentions here for Cloudflare: https://twitter.com/jgrahamc/status/1445066136547217413
Re: Facebook-owned sites were down
#478Reddit r/Sysadmin user that claims to be on the "Recovery Team" for this ongoing issue: > As many of you know, DNS for FB services has been affected and this is likely a symptom of the actual issue, and that's that BGP peering with Facebook peering routers has gone down, very likely due to a configuration change that went into effect shortly before the outages happened (started roughly 1540 UTC). There are people now…