The badge story only shows how people are looking for "efficiency" where it doesn't matter, with predictable results. The badge system should be local to the building. There are few actual reasons (sure, besides "efficiency") of why badge control should be centralized. Even less reasons for it to be a subdomain of fb. Another option would be to keep the system but make it failsafe (but it seems the newer generation d…
Update about the October 4th outage
201–210 of 239 posts
Re: Update about the October 4th outage
#202Earlier quoted context omitted.
Both posts have essentially the same info - the fb one just didn't include an explainer on how the internet works.
The Cloudflare post includes graphs depicting how the issue looked downstream and a rough initial timeline for the incident. The Facebook post says basically nothing more than that there was a networking mistake. I hope we get something more substantial and informative than that over the next couple of weeks, but it doesn't seem (at least from my searching) that Facebook is in the business of publicly posting in-dept…
vs
How tech savvy are people that Facebook profits from?
Gotta target your audience, every communication is PR...
Re: Update about the October 4th outage
#203Earlier quoted context omitted.
BGP has to converge to a single routing table. You are effectively asking is why is there a single routing table for the internet. To put in simple terms having a single routing table is what it makes it the internet we can share, otherwise it would just be a bunch of independent networks.
> You are effectively asking is why is there a single routing table for the internet. No, they're asking why a single set of routers is in charge of announcing BGP routes for all of facebook. If you have multiple ASes, with independent configuration sources and independent routers broadcasting them, it's a lot harder to break everything at once.
However if even one set of routers were misconfigured and it was announcing incorrect routes for all their ASes as result of the issue then their peers will not typically drop that set alone automatically.
BGP doesn't have a Paxos / Raft style smart consensus algorithms, it runs on trust. Either their peers had to trust what FB published or they won't be peering with FB ASes in the first place.
That's what I was meant when I said it comes down to one network web of trust, and if there is a breakdown of that redundancy cannot typically help
Re: Update about the October 4th outage
#204Earlier quoted context omitted.
> Facebook's routing shenanigans slowed down the entire internet This is Hacker News, so the distinction between network performance, server performance and application performance should matter. "The Internet" did not slow down. "The Internet" infact probably had more available capacity as a result of Facebook's outage, as all those bits of outrage and cats ceased to be transferred for the duration. Some application…
Perhaps you're unaware that billions of devices attempting to resolve Facebook's unresolvable domains effectively DDOS-ed the DNS system? It most certainly did slow down big chunks of internet which otherwise had nothing to do with Facebook. https://www.theverge.com/2021/10/4/22709123/facebook-outage-...
I had Cloudflare's woes in mind when I wrote that.
Re: Update about the October 4th outage
#205May the outage be longer. And May Mark be removed as its snakehead.
Re: Update about the October 4th outage
#206I worked with a network engineer who misconfigured a router that was connecting a bank to it's DR site. The engineer had to drive across town to manually patch into the router to fix it. DR downtime was about an hour, but the bank fired him anyway. Given that Zuck lost a substantial amount of money, I wonder if the engineer faced any ramifications. Sidenote: I asked the bank infrastructure team why the DR site was in…
"I worked with a network engineer who misconfigured a router that was connecting a bank to it's DR site. The engineer had to drive across town to manually patch into the router to fix it. DR downtime was about an hour, but the bank fired him anyway." so prod wasn't down and he fixed it in a hour and they fired the guy who knew how to fix such things so quickly. Idiot manager at the bank.
Re: Update about the October 4th outage
#207So their actual deployment process is quite rigorous and should have a tight blast radius. After lots of emulated and canary testing, their deployments are phased out over weeks. I don't see how a bad push could have done what happened yesterday. I found a paper that describes the process in detail. See page 10-11: https://web.archive.org/web/20211005034928/https://research.... Phase Specification P1 Small number of…
If they have such a rigorous release process, what could have caused all of the dns records to get wiped?
Hey what is our internal BGP called again? AS32934?
"Yeah"
"OOK."
Re: Update about the October 4th outage
#208I worked with a network engineer who misconfigured a router that was connecting a bank to it's DR site. The engineer had to drive across town to manually patch into the router to fix it. DR downtime was about an hour, but the bank fired him anyway. Given that Zuck lost a substantial amount of money, I wonder if the engineer faced any ramifications. Sidenote: I asked the bank infrastructure team why the DR site was in…
"I worked with a network engineer who misconfigured a router that was connecting a bank to it's DR site. The engineer had to drive across town to manually patch into the router to fix it. DR downtime was about an hour, but the bank fired him anyway." so prod wasn't down and he fixed it in a hour and they fired the guy who knew how to fix such things so quickly. Idiot manager at the bank.
Prostrate, he came before the COO expecting to be canned with much malice. The COO just asked if he learned his lesson and said all is forgiven.
Re: Update about the October 4th outage
#209Earlier quoted context omitted.
I didn’t see any disinformation, just initial reports that it was DNS which were later explained to be caused by BGP.
- It was government intervention - Facebook was hacked - They did it on purpose to bury the whistleblower story - No one could access Facebook offices - They had to cut open servers with angle grinders - Disgruntled employees changed DNS records - Lots of made up numbers for how much money Facebook/the rest of the economy was losing (or gaining) They probably rushed out this blog post just to dispel some of these rum…
I want to believe.
Re: Update about the October 4th outage
#210Earlier quoted context omitted.
I don't think this is the case. Wasn't TechLead fired for SEV events?
Where did you hear that from? He doesn’t even say that in his video, he says he was fired for having side income on YouTube.