More details about the October 4 outage
31–40 of 306 posts
Re: More details about the October 4 outage
#32Interesting bit on recovery w.r.t. the electrical grid > flipping our services back on all at once could potentially cause a new round of crashes due to a surge in traffic. Individual data centers were reporting dips in power usage in the range of tens of megawatts, and suddenly reversing such a dip in power consumption could put everything from electrical systems ... I wish there was a bit more detail in here. What'…
Brownouts is probably the most proximate concern - a sudden increase in demand will draw down the system frequency in the vicinity, and if there aren't generation units close enough or with enough dispatchable capacity there's a small chance they would trip a protective breaker. A person I know on the power grid side said at one data center there were step functions when FB went down and then when it came up, equal t…
Re: More details about the October 4 outage
#33> One of the jobs performed by our smaller facilities is to respond to DNS queries. DNS is the address book of the internet, enabling the simple web names we type into browsers to be translated into specific server IP addresses. What is the target audience of this post? It is too technical for non-technical people, but also it is dumbed down to try to include people that does not know how the internet works. I feel l…
Re: More details about the October 4 outage
#34Google had a comparable outage several years ago. https://status.cloud.google.com/incident/cloud-networking/19... This event left a lot of scar tissue across all of Technical Infrastructure, and the next few months were not a fun time (e.g. a mandatory training where leadership read out emails from customers telling us how we let them down and lost their trust). I'd be curious to see what systemic changes happen at F…
It was a global backbone isolation, caused by configuration changes (as they all are...). It was detected fairly early on, but recovery was difficult because internal tools / debugging workflows were also impacted, and even after the problem was identified, it still took time to back out the change.
"But wait, a global backbone isolation? Google wasn't totally down," you might say. That's because Google has two (primary) backbones (B2 and B4), and only B4 was isolated, so traffic spilled over onto B2 (which has much less capacity), causing heavy congestion.
Re: More details about the October 4 outage
#35 the total loss of DNS broke many of the internal tools we’d normally use to investigate and resolve outages like this.
this took time, because these facilities are designed with high levels of physical and system security in mind. They’re hard to get into, and once you’re inside, the hardware and routers are designed to be difficult to modify even when you have physical access to them.
Sounds like it was the perfect storm.Re: More details about the October 4 outage
#36https://twitter.com/cullend/status/1445156376934862848?t=P5u...
Re: More details about the October 4 outage
#37We want ramenporn
Re: More details about the October 4 outage
#38It's a no apologies messages: "We failed, our processes failed, our recovery process only partially worked, we celebrate failure. Our investors were not happy, our users were not happy, some people probably ended in physically dangerous situations due to WhatsApp being unavailable but it's ok. We believe a tradeoff like this is worth it." - Your engineering team.
Re: More details about the October 4 outage
#39Apparently they had to bring in the angle grinder to get access to the server room. https://twitter.com/cullend/status/1445156376934862848?t=P5u...
> the team dispatched to the Facebook site had issues getting in because of physical security but did not need to use a saw/ grinder.
Re: More details about the October 4 outage
#40Apparently they had to bring in the angle grinder to get access to the server room. https://twitter.com/cullend/status/1445156376934862848?t=P5u...