[1] https://www.infoworld.com/article/2648947/youtube-outage-und...
Facebook-owned sites were down
591–600 of 1001 posts
Re: Facebook-owned sites were down
#592Earlier quoted context omitted.
Just imagine the amount of stress on this people, hope the money really worth it.
It shouldn't be too stressful. Well-managed companies blame processes rather than people, and have systems set up to communicate rapidly when large-scale events occur. It can be sort of exciting, but it's not like there is one person typing at a keyboard with a hundred managers breathing down their neck. These resolutions are collaborative, shared efforts.
As someone who formerly did Ops for many many years... this is not accurate. Even in a well organized company there are usually stakeholders at every level on IM calls so that they don't need to play "telephone" for status. For an incident of this size, it wouldn't be unusual to have C-level executives on the call.
While those managers are mostly just quietly listening in on mute if they know what's good (e.g. don't distract the people doing the work to fix your problem), their mere presence can make the entire situation more tense and stressful for the person banging keyboards. If they decide to be chatty or belligerent, it makes everything 100x worse.
I don't envy the SREs at Facebook today. Godspeed fellow Ops homies.
Re: Facebook-owned sites were down
#593Is it just me or HN also feels kinda laggy?
Probably people flooding in to see if anyone knows why things are down. Even Google speed test was down, presumably from too many people testing if it's their internet that's at issue.
Re: Facebook-owned sites were down
#594Earlier quoted context omitted.
There are various non-FB fallback measures, including IRC as a last-ditch method. The IRC fallback is usually tested once a year for each engineer.
I just heard from a contact that the fallback/backup IRC is also down.
Re: Facebook-owned sites were down
#595Re: Facebook-owned sites were down
#596Earlier quoted context omitted.
> the people with physical access is separate from the people with knowledge of [...] Welcome to the brave new world of troubleshooting. This will seriously bite us one day.
This is why so many teams fight back against the audit findings: "The information systems office did not enforce logical access to the system in accordance with role-based access policies." Invariably, you want your best people to have full access to all systems.
If you're a mature larger company, that's the team leads in your networking area on the team that deal with that service area (BGP routing, or routers in general).
Most likely Facebook et. al. management never understood this could happen because it's "never been a problem before".
Re: Facebook-owned sites were down
#597Is it just me or HN also feels kinda laggy?
General tip: If HN is being laggy and you're determined you want to waste some time here, open it in a private window. HN works extremely quickly if it doesn't know who you are.
Re: Facebook-owned sites were down
#598> This must be incredibly stressful so for your sake I hope you sort it out quickly... but for the world's sake, I hope you fail and make the problem worse before jumping ship followed by every other engineer, leaving it to Zuckerberg to fix himself. But I still hope it's not too stressful for you!
Re: Facebook-owned sites were down
#599Re: Facebook-owned sites were down
#600Reddit r/Sysadmin user that claims to be on the "Recovery Team" for this ongoing issue: > As many of you know, DNS for FB services has been affected and this is likely a symptom of the actual issue, and that's that BGP peering with Facebook peering routers has gone down, very likely due to a configuration change that went into effect shortly before the outages happened (started roughly 1540 UTC). There are people now…
"I believe the original change was 'automatic' (as in configuration done via a web interface). However, now that connection to the outside world is down, remote access to those tools don't exist anymore, so the emergency procedure is to gain physical access to the peering routers and do all the configuration locally." Hmm, could be a UI/UX bug then :)