Postmortem of database outage of January 31
about.gitlab.com
Postmortem of database outage of January 31
1–10 of 269 posts
Re: Postmortem of database outage of January 31
#2http://serverfault.com/questions/587102/monday-morning-mista...
Re: Postmortem of database outage of January 31
#3I'm glad it all worked out in the end!
Re: Postmortem of database outage of January 31
#4Reading this, the thing that stuck out to me was how remarkably lucky they were to have the two snapshots. The one from 6 hours earlier was there seemingly by chance, as an engineer had created it for unrelated reasons. And for both the 6- and 24-hour snapshots, it seems just lucky that neither had any breaking changes made to them by pre-production code (they _were_ dev/staging snapshots, after all). I'm glad it all…
Re: Postmortem of database outage of January 31
#5Re: Postmortem of database outage of January 31
#6Reading this, the thing that stuck out to me was how remarkably lucky they were to have the two snapshots. The one from 6 hours earlier was there seemingly by chance, as an engineer had created it for unrelated reasons. And for both the 6- and 24-hour snapshots, it seems just lucky that neither had any breaking changes made to them by pre-production code (they _were_ dev/staging snapshots, after all). I'm glad it all…
We too are glad we had those snapshots. And while it was the worst thing that ever happend at GitLab it is humbling to know that it could have been worse.
Re: Postmortem of database outage of January 31
#7Re: Postmortem of database outage of January 31
#8definitely monitor your replication lag--or at least disk usage on the master--with this approach (in case wal starts piling up there).
Re: Postmortem of database outage of January 31
#9Re: Postmortem of database outage of January 31
#10I could feel the sweat drops just from reading this.
I'd bet every one of us has experienced the panicked Ctrl+C of Death at some point or another.