Live data from Hacker News

Facebook-owned sites were down

facebook.com

391–400 of 1001 posts

Re: Facebook-owned sites were down

#392
post #314

Reddit r/Sysadmin user that claims to be on the "Recovery Team" for this ongoing issue: > As many of you know, DNS for FB services has been affected and this is likely a symptom of the actual issue, and that's that BGP peering with Facebook peering routers has gone down, very likely due to a configuration change that went into effect shortly before the outages happened (started roughly 1540 UTC). There are people now…

> the people with physical access is separate from the people with knowledge of [...] Welcome to the brave new world of troubleshooting. This will seriously bite us one day.

[deleted]

Re: Facebook-owned sites were down

#393
It seems it has caused DNS servers crash for one of biggest Czechia's internet provider - Vodafone. Can be unrelated but I doubt it (https://twitter.com/BlazejKrajnak/status/1445063232486531099).

Think of it - half the country doesn't have internet because of this crash, that's terrifying. (Switching DNS servers obviously works but that's not something the general population will do)

Re: Facebook-owned sites were down

#394
post #314

Reddit r/Sysadmin user that claims to be on the "Recovery Team" for this ongoing issue: > As many of you know, DNS for FB services has been affected and this is likely a symptom of the actual issue, and that's that BGP peering with Facebook peering routers has gone down, very likely due to a configuration change that went into effect shortly before the outages happened (started roughly 1540 UTC). There are people now…

> the people with physical access is separate from the people with knowledge of [...] Welcome to the brave new world of troubleshooting. This will seriously bite us one day.

Telecommunication satellite communication issues might seriously shut down whole regions if they occur.

Re: Facebook-owned sites were down

#395
Somebody just had their very own "onosecond".

https://www.youtube.com/watch?v=X6NJkWbM1xk

The video is one that Tom Scott published in June 2020 about the worst typo he ever made in one of his prior jobs, and while the Facebook mistake is almost certainly not going to be anything irrecoverable like this one, you can bet that Facebook pride themselves on being available all the time.

Re: Facebook-owned sites were down

#396

Okay, let me tell you the difference between Facebook and everyone else, we don't crash EVER! If those servers are down for even a day, our entire reputation is irreversibly destroyed! Users are fickle, Friendster has proved that. Even a few people leaving would reverberate through the entire userbase. The users are interconnected, that is the whole point. College kids are online because their friends are online, and…

Is this a quote?

Re: Facebook-owned sites were down

#397

Reddit r/Sysadmin user that claims to be on the "Recovery Team" for this ongoing issue: > As many of you know, DNS for FB services has been affected and this is likely a symptom of the actual issue, and that's that BGP peering with Facebook peering routers has gone down, very likely due to a configuration change that went into effect shortly before the outages happened (started roughly 1540 UTC). There are people now…

Just imagine the amount of stress on this people, hope the money really worth it.

The stress for me usually goes away once the incident is fully escalated and there's a team with me working on the issue. I imagine that happened quite quick in this case...

Re: Facebook-owned sites were down

#398
post #377

Earlier quoted context omitted.

It shouldn't be too stressful. Well-managed companies blame processes rather than people, and have systems set up to communicate rapidly when large-scale events occur. It can be sort of exciting, but it's not like there is one person typing at a keyboard with a hundred managers breathing down their neck. These resolutions are collaborative, shared efforts.

> It shouldn't be too stressful. (...) it's not like there is one person typing at a keyboard with a hundred managers breathing down their neck Earlier comment mentioned that there is a bottleneck, and that people who are physically able to solve the issue are few and that they need to be informed what to do; being one of these people sounds pretty stressful to me. "but the people with physical access is separate (..…

Sure, but that's what conference calls are for.

Most big tech companies automatically start a call for every large scale incident, and adjacent teams are expected to have a representative call in and contribute to identifying/remediating the issue.

None of the people with physical access are individually responsible, and they should have a deep bench of advice and context to draw from.

Re: Facebook-owned sites were down

#399
Reading the thread, I'm surprised at the number of nearly identical "How much do we have to pay to keep it down? xD" posts I'm seeing, often from throwaway accounts. Some accounts with multiple near-identical posts within the same minute.

Could this be a coordinated smear in HN comments?

Re: Facebook-owned sites were down

#400
post #380

Reddit r/Sysadmin user that claims to be on the "Recovery Team" for this ongoing issue: > As many of you know, DNS for FB services has been affected and this is likely a symptom of the actual issue, and that's that BGP peering with Facebook peering routers has gone down, very likely due to a configuration change that went into effect shortly before the outages happened (started roughly 1540 UTC). There are people now…

Wondering how Facebook communicates now internally - most of their work streams likely depend on Facebooks systems which are all down. Can engineers and security teams even access prod systems anymore? Like, would "Bastion" hosts be reachable? Wonder if they use Signal and Slack now?

FB uses a separate IRC instance for these kinds of issues, at least when I used to work there
Post reply on HN