Has Facebook ever publicly written about the major Intel chip SSE failures they ran into a few years ago?
How Facebook deals with PCIe faults to keep its data centers running reliably
21–30 of 53 posts
Re: How Facebook deals with PCIe faults to keep its data centers running reliably
#22Has Facebook ever publicly written about the major Intel chip SSE failures they ran into a few years ago?
Re: How Facebook deals with PCIe faults to keep its data centers running reliably
#23Re: How Facebook deals with PCIe faults to keep its data centers running reliably
#24Earlier quoted context omitted.
I believe this is called "technical marketing" Does Facebook still want to launch an own AWS clone?
Not sure if they did who would end up being the bigger asshole. My opinion of AWS is already at an all time low, but my opinion of Facebook has always been low too. At least it would introduce some competition for AWS.
Re: How Facebook deals with PCIe faults to keep its data centers running reliably
#25Earlier quoted context omitted.
Not sure if they did who would end up being the bigger asshole. My opinion of AWS is already at an all time low, but my opinion of Facebook has always been low too. At least it would introduce some competition for AWS.
Some competition for AWS? Seems like Azure and GCP are competing pretty hard don't you think?
Azure is a minefield that's only worth it when on a .NET stack.
Re: How Facebook deals with PCIe faults to keep its data centers running reliably
#26Sucks that I have to worry while reading their technical blog post that it's pretty likely that Facebook is tracking the fact that I'm reading their article.
Re: How Facebook deals with PCIe faults to keep its data centers running reliably
#27This is the most important bit for someone reimplementing this...
Never let automation 'run wild' - always have a maximum number of machines per second it can act on, and a maximum percentage of the fleet unhealthy to allow it to continue taking things out of service.
Re: How Facebook deals with PCIe faults to keep its data centers running reliably
#28Earlier quoted context omitted.
Some competition for AWS? Seems like Azure and GCP are competing pretty hard don't you think?
I know many people who won't dare touch GCP with a ten foot pole simply because they're afraid of Google's random banning AI wiping their digital lives. Azure is a minefield that's only worth it when on a .NET stack.
Re: How Facebook deals with PCIe faults to keep its data centers running reliably
#29> also important to rate-limit remediations and repairs as a safety net to prevent bugs in the code from mass draining and unprovisioning, which can result in service outages if not handled properly. This is the most important bit for someone reimplementing this... Never let automation 'run wild' - always have a maximum number of machines per second it can act on, and a maximum percentage of the fleet unhealthy to al…
Re: How Facebook deals with PCIe faults to keep its data centers running reliably
#30Sucks that I have to worry while reading their technical blog post that it's pretty likely that Facebook is tracking the fact that I'm reading their article.