Which major outage ? According to the Slack uptime, there was barely 1.5 hour of outage :-) P.S. Yes I know the uptime is decide by committee, and doesn't reflect reality. I am just being cynical.
1.5 hours for a tool like Slack is major. Lots of productivity lost (or gained depending on how you view Slack) and thus $ impacts at companies that heavily rely on Slack for internal/team comms
Slack’s Incident on 2-22-22
141–150 of 183 posts
Re: Slack’s Incident on 2-22-22
#142Earlier quoted context omitted.
The complaint isn't about the particular other order, but the fact that the order is ambiguous. In this case that doesn't matter, but often it does. Americans memorize inches and yards, and often also memorize centimeters and meters, and working with either is fine, but we're not so often faced with numbers where it might be inches or centimeters and we have to figure out which (and when we are, it's sometimes a pain…
I instinctively felt that the ambiguity ought to not truly exist, but I wanted to find some backing. According to this: https://en.wikipedia.org/wiki/Date_format_by_country … all the major predominantly English speaking counties will use mostly hyphens in the dd-mm-yyyy format. So although there is ambiguity, it’s easily resolved by picking that as the default mentally and only back tracking on failure. Now in the mo…
Picking one default and back-tracking on failure really isn't that comforting nor the constant reminder that the date you thought it was might be something else.
Re: Slack’s Incident on 2-22-22
#143Earlier quoted context omitted.
The fact that the current top comment thread is quibbling about the date format in the title seems to agree with this assessment, if there was anything real to complain about that’s what we’d be seeing, instead we get bikeshedding on the date format in the title of a post.
[adds "bikeshedding" to vocabulary.]
Re: Slack’s Incident on 2-22-22
#144Great post mortem, I love reading these. Its pretty neat that the tech industry is relatively transparent about these situations -- we all benefit from learning about them.
I am really happy you think so. But no. We are really not transparent about them. At all.
Re: Slack’s Incident on 2-22-22
#145Just my personal take, I think this is a really well-written incident postmortem. It's specific, extensive, candid, and dare I say, entertaining? Many incident reports are fully lacking in any meaningful detail, or wholly unapologetic. I actually enjoyed learning tidbits about the author, in particular their mention of https://how.complexsystems.fail/ . Reading this boosted my confidence in Slack's teams, which shoul…
Re: Slack’s Incident on 2-22-22
#146Tidbit: 2-22-22 was also when Russia invaded Ukraine And Joe Bidens statement about the invasion was on 2-22-22 2:22pm on the dot. I could not figure out the significance of this more than 11:11 was when WW1 ended, but it's probably something else.
Nope, you’re off by two days: https://en.wikipedia.org/wiki/2022_Russian_invasion_of_Ukrai...
Then Biden did a speech against on 2:22 22-2-22
The actual war started in 2014
Re: Slack’s Incident on 2-22-22
#147Tidbit: 2-22-22 was also when Russia invaded Ukraine And Joe Bidens statement about the invasion was on 2-22-22 2:22pm on the dot. I could not figure out the significance of this more than 11:11 was when WW1 ended, but it's probably something else.
Nope, you’re off by two days: https://en.wikipedia.org/wiki/2022_Russian_invasion_of_Ukrai...
https://edition.cnn.com/europe/live-news/ukraine-russia-news...
Re: Slack’s Incident on 2-22-22
#148Even the largest Slack instance probably has under 100,000 users and less than 1000 peak messages per second. That feels to me like it could be served by a single master DB. It feels like that would be a better way to shard.
Major downside is the major difference in shard sizes so some management/migration might be needed but it seems doable to me.
Certainly it feels naively that scaling should be easy due to the way slack instances are independent (unlike say Twitter).
Re: Slack’s Incident on 2-22-22
#149I am confused by the query that had the problem. Specifically, I am confused by why the sharding is done by user id. Even the largest Slack instance probably has under 100,000 users and less than 1000 peak messages per second. That feels to me like it could be served by a single master DB. It feels like that would be a better way to shard. Major downside is the major difference in shard sizes so some management/migra…
Re: Slack’s Incident on 2-22-22
#150Earlier quoted context omitted.
which only makes it confusing for the rest of the world
It's not that confusing if you know there are only 12 months. If it happened to occur on 2-4-22, then it would be pretty confusing.