Live data from Hacker News

Update about the October 4th outage

engineering.fb.com

201–210 of 239 posts

Re: Update about the October 4th outage

#201

The badge story only shows how people are looking for "efficiency" where it doesn't matter, with predictable results. The badge system should be local to the building. There are few actual reasons (sure, besides "efficiency") of why badge control should be centralized. Even less reasons for it to be a subdomain of fb. Another option would be to keep the system but make it failsafe (but it seems the newer generation d…

The issue I would imagine is that during this outage some people needed badge access that had previously been revoked due to covid. All of the caching doesn't help if your source of truth is offline.

Re: Update about the October 4th outage

#202

Earlier quoted context omitted.

Both posts have essentially the same info - the fb one just didn't include an explainer on how the internet works.

The Cloudflare post includes graphs depicting how the issue looked downstream and a rough initial timeline for the incident. The Facebook post says basically nothing more than that there was a networking mistake. I hope we get something more substantial and informative than that over the next couple of weeks, but it doesn't seem (at least from my searching) that Facebook is in the business of publicly posting in-dept…

How tech savvy are people who pay Cloudflare money?

vs

How tech savvy are people that Facebook profits from?

Gotta target your audience, every communication is PR...

Re: Update about the October 4th outage

#203
post #18

Earlier quoted context omitted.

BGP has to converge to a single routing table. You are effectively asking is why is there a single routing table for the internet. To put in simple terms having a single routing table is what it makes it the internet we can share, otherwise it would just be a bunch of independent networks.

> You are effectively asking is why is there a single routing table for the internet. No, they're asking why a single set of routers is in charge of announcing BGP routes for all of facebook. If you have multiple ASes, with independent configuration sources and independent routers broadcasting them, it's a lot harder to break everything at once.

Facebook definitely has multiple ASes and likely rolling BGP update strategies already. This isn't the first high profile BGP screwup after all.

However if even one set of routers were misconfigured and it was announcing incorrect routes for all their ASes as result of the issue then their peers will not typically drop that set alone automatically.

BGP doesn't have a Paxos / Raft style smart consensus algorithms, it runs on trust. Either their peers had to trust what FB published or they won't be peering with FB ASes in the first place.

That's what I was meant when I said it comes down to one network web of trust, and if there is a breakdown of that redundancy cannot typically help

Re: Update about the October 4th outage

#204
post #197

Earlier quoted context omitted.

> Facebook's routing shenanigans slowed down the entire internet This is Hacker News, so the distinction between network performance, server performance and application performance should matter. "The Internet" did not slow down. "The Internet" infact probably had more available capacity as a result of Facebook's outage, as all those bits of outrage and cats ceased to be transferred for the duration. Some application…

Perhaps you're unaware that billions of devices attempting to resolve Facebook's unresolvable domains effectively DDOS-ed the DNS system? It most certainly did slow down big chunks of internet which otherwise had nothing to do with Facebook. https://www.theverge.com/2021/10/4/22709123/facebook-outage-...

> Some applications may have seen increased load and suffered due to server resourcing constraints, caused by applications like the above failing to fail gracefully, and instead polling more aggressively.

I had Cloudflare's woes in mind when I wrote that.

Re: Update about the October 4th outage

#206

I worked with a network engineer who misconfigured a router that was connecting a bank to it's DR site. The engineer had to drive across town to manually patch into the router to fix it. DR downtime was about an hour, but the bank fired him anyway. Given that Zuck lost a substantial amount of money, I wonder if the engineer faced any ramifications. Sidenote: I asked the bank infrastructure team why the DR site was in…

"I worked with a network engineer who misconfigured a router that was connecting a bank to it's DR site. The engineer had to drive across town to manually patch into the router to fix it. DR downtime was about an hour, but the bank fired him anyway." so prod wasn't down and he fixed it in a hour and they fired the guy who knew how to fix such things so quickly. Idiot manager at the bank.

I agree it was very heavy handed, but I suspect there was more at play (not the first mistake, and some regulatory reporting that may have looked bad for higher ups)

Re: Update about the October 4th outage

#207

So their actual deployment process is quite rigorous and should have a tight blast radius. After lots of emulated and canary testing, their deployments are phased out over weeks. I don't see how a bad push could have done what happened yesterday. I found a paper that describes the process in detail. See page 10-11: https://web.archive.org/web/20211005034928/https://research.... Phase Specification P1 Small number of…

If they have such a rigorous release process, what could have caused all of the dns records to get wiped?

"Are you sure you want to remove ALL routes to AS32934? Type YES to confirm."

Hey what is our internal BGP called again? AS32934?

"Yeah"

"OOK."

Re: Update about the October 4th outage

#208

I worked with a network engineer who misconfigured a router that was connecting a bank to it's DR site. The engineer had to drive across town to manually patch into the router to fix it. DR downtime was about an hour, but the bank fired him anyway. Given that Zuck lost a substantial amount of money, I wonder if the engineer faced any ramifications. Sidenote: I asked the bank infrastructure team why the DR site was in…

"I worked with a network engineer who misconfigured a router that was connecting a bank to it's DR site. The engineer had to drive across town to manually patch into the router to fix it. DR downtime was about an hour, but the bank fired him anyway." so prod wasn't down and he fixed it in a hour and they fired the guy who knew how to fix such things so quickly. Idiot manager at the bank.

Had a DBA once who was playing around with database projects in visual studio and he managed to hose the production database in the course of it. This caused our entire system to go down.

Prostrate, he came before the COO expecting to be canned with much malice. The COO just asked if he learned his lesson and said all is forgiven.

Re: Update about the October 4th outage

#209
post #41

Earlier quoted context omitted.

I didn’t see any disinformation, just initial reports that it was DNS which were later explained to be caused by BGP.

- It was government intervention - Facebook was hacked - They did it on purpose to bury the whistleblower story - No one could access Facebook offices - They had to cut open servers with angle grinders - Disgruntled employees changed DNS records - Lots of made up numbers for how much money Facebook/the rest of the economy was losing (or gaining) They probably rushed out this blog post just to dispel some of these rum…

"They had to cut open servers with angle grinders"

I want to believe.

Re: Update about the October 4th outage

#210

Earlier quoted context omitted.

I don't think this is the case. Wasn't TechLead fired for SEV events?

Where did you hear that from? He doesn’t even say that in his video, he says he was fired for having side income on YouTube.

Sorry, I was thinking about the engineer who got PIPd and committed suicide.
Post reply on HN