Lessons Learned from Twenty Years of Site Reliability Engineering
1–10 of 128 posts
Re: Lessons Learned from Twenty Years of Site Reliability Engineering
#2Re: Lessons Learned from Twenty Years of Site Reliability Engineering
#3Re: Lessons Learned from Twenty Years of Site Reliability Engineering
#4"We once narrowly missed a major outage because the engineer who submitted the would-be-triggering change unplugged their desktop computer before the change could propagate". Sorry, what?
Re: Lessons Learned from Twenty Years of Site Reliability Engineering
#5"We once narrowly missed a major outage because the engineer who submitted the would-be-triggering change unplugged their desktop computer before the change could propagate". Sorry, what?
Re: Lessons Learned from Twenty Years of Site Reliability Engineering
#6"We once narrowly missed a major outage because the engineer who submitted the would-be-triggering change unplugged their desktop computer before the change could propagate". Sorry, what?
Re: Lessons Learned from Twenty Years of Site Reliability Engineering
#7> COMMUNICATION CHANNELS! AND BACKUP CHANNELS!! AND BACKUPS FOR THOSE BACKUP CHANNELS!!!
Yes! At Netflix, when we picked vendors for systems that we used during an outage, we always had to make sure they were not on AWS. At reddit we had a server in the office with a backup IRC server in case the main one we used was unavailable.
And I think Google has a backup IRC server on AWS, but that might just apocryphal.
Always good to make sure you have a backup side channel that has as little to do with your infrastructure as possible.
Re: Lessons Learned from Twenty Years of Site Reliability Engineering
#8"We once narrowly missed a major outage because the engineer who submitted the would-be-triggering change unplugged their desktop computer before the change could propagate". Sorry, what?
Re: Lessons Learned from Twenty Years of Site Reliability Engineering
#9"We once narrowly missed a major outage because the engineer who submitted the would-be-triggering change unplugged their desktop computer before the change could propagate". Sorry, what?
I do think Google should be more open about their internal outages. This one in particular was very famous internally.
Re: Lessons Learned from Twenty Years of Site Reliability Engineering
#10This is a great writeup, and very broadly applicable. I don't see any "this would only apply at Google" in here. > COMMUNICATION CHANNELS! AND BACKUP CHANNELS!! AND BACKUPS FOR THOSE BACKUP CHANNELS!!! Yes! At Netflix, when we picked vendors for systems that we used during an outage, we always had to make sure they were not on AWS. At reddit we had a server in the office with a backup IRC server in case the main one…
> very broadly applicable. I don't see any "this would only apply at Google" in here.
One thing I haven’t even heard of anyone else doing was production panic rooms - a secure room with backup vpn to prod