Live data from Hacker News

Slack is down

status.slack.com

81–90 of 840 posts

Re: Slack is down

#81

Could it be the obvious? Everyone signing on / loading slack clients at the same time?

then wouldn't this happen every monday morning?

A lot of organizations essentially took the last two weeks off from work, which is long enough for a 10-day autoscale window to spin down servers, and then get confronted by a load spike that wasn't pre-spun for.

Re: Slack is down

#82
post #57

Could it be the obvious? Everyone signing on / loading slack clients at the same time?

Nice of you to ignore the majority of the world's population who have been up and working long before America woke up.

Slack makes ~61% of its revenue from US customers which only has 4 time zones compared to the remainder of their revenue being spread out across ~20 time zones. It's not an unreasonable hypothesis.

See page 12 of the document (which is page 14 of the PDF) https://d18rn0p25nwr6d.cloudfront.net/CIK-0001764925/70df834...

Re: Slack is down

#83
post #29
post #19

Earlier quoted context omitted.

If you're asking genuinely then I can tell you my experience when I was part of a SaaS shop, though the times have changed a lot and "my metric is not necessarily your metric". But it was roughly "one large impact a month, for six months", with large caveats that upper management for whatever company had to be working with the product during that month. Large companies don't care if X service went out during the nigh…

> But migrating everything is _so painful_ This is a key point is the popularity amongst VCs in investing in B2B SaaS. I take their (and your) word for it. But honestly, I don't actually understand this. Why is migration so hard?

It's not, honestly.

I worked at a huge corp which employed many non-technical people and had global offices and all sorts of contractors / full-time employees at various levels.

But they invested in organizational / business software early (think SAP-type stuff) and so at any moment every employee / hired hand is accounted for, has a number, has a position in the org chart, has access to a wiki-type platform where they can be trained and informed of any changes to the workplace software suite and guided through any migrations.

I've seen the company migrate off Slack onto Microsoft Teams. I've seen the company migrate to MS Sharepoint from Box. I've seen them migrate everyone onto a platform called SuccessFactors (still don't really know what it does, it's for tracking your career progress I think).

There's work involved. Someone has to write a guide to get users to sign up with the new service (even with SSO linked to your corporate account it's not easy for most non-technical people). In the case of Slack, any hooks or bots created for any teams need to be turned off and people need to be informed well in advance of the shift and multiple times. Optional in-person trainings need to be provided. Some employees may have issues with the change (a missing feature in the new software that they rely on) and they need a forum (or even just someone to contact) where they can lay out these issues and get them resolved.

But that's not that bad honestly. If moving from a platform gives you serious gains in uptime, it's not that bad. I think Slack's downtime problems are not so bad that most people will move yet, but they may soon get there.

Re: Slack is down

#86

Earlier quoted context omitted.

I'm a Carl. I'm also looking for a coworker who was trying to contact me. If it's about last saturday, I promise nothing really happened between me and her, but I'm sure she already told you.

Do you guys not have e-mail? looks through Inbox of 850 new aws, batch job and logging messages oh yea, that's right..

Don't you guys have e-mail filters?

"Hey, our site has been down for 2 hours, why aren't you doing anything"

Looks at 850 unread messages in ops-notifications folder

ooh yeah, that's right..

Re: Slack is down

#87
I'll concede that it's possible to not know what the problem is by now, but I won't concede that this should not be called an "outage" at this point.

Re: Slack is down

#88
post #29
post #19

Earlier quoted context omitted.

If you're asking genuinely then I can tell you my experience when I was part of a SaaS shop, though the times have changed a lot and "my metric is not necessarily your metric". But it was roughly "one large impact a month, for six months", with large caveats that upper management for whatever company had to be working with the product during that month. Large companies don't care if X service went out during the nigh…

> But migrating everything is _so painful_ This is a key point is the popularity amongst VCs in investing in B2B SaaS. I take their (and your) word for it. But honestly, I don't actually understand this. Why is migration so hard?

Many reasons, almost none of them technical. Off the top of my head, a few:

* Getting out of the Enterprise Contract, or waiting for the year to end. * Training people on new software. * Loss of productivity. (1) Learning a new UI, processes, workflows -- both individually and organizationally. A feature or concept in "Tool A" may exist in a completely different form in "Tool B". Or not exist, and then people need to adapt to and work around the missing feature. (2) Missing out on needed information due to the above. Ultimately, software exists to move and transform data, and when you change the software people have to adjust. Sometimes that doesn't go great. "Oh, I didn't realize I needed to check this checkbox".

Another way to say this is "organizational inertia", which is a fancy term that means "it's hard for people to adjust to change".

And you might think developers and other technical people would have an easier time of it. They (we) do, but not to the extent you may expect. I've been on the front lines of a handful of migrations that affected only the IT staff, and it was a long and arduous process each time.

Re: Slack is down

#89
When I was at Uber, we noticed that most incidents are directly caused by human actions that modify the state of the system. Therefore, a large "backlog" of human actions that modify the system state have a much higher chance of causing an incident.

My bet is that this incident is caused by a big release after a post-holiday "code freeze".

Re: Slack is down

#90
post #77

Earlier quoted context omitted.

I say this every time Slack is down, but they just seem so shady to me. Nobody can connect right now, and their status site says "100% uptime in the last quarter". Maybe it's close to 100%, but it ain't 100%. I think we should push for a metric where "up" means 100% of people that want to use the service are able to use the service. If 1% of users can't send messages, then that should count as a full-blown outage and…

You'll just end up with no SLA or pay a hefty amount to use services because that's an impossible standard to support for any service of a size like this.

Isn't this the problem? Companies like Slack set SLA's that they only meet by lying about their uptime. It's as good as having no SLA, except you're likely paying a premium based on the SLA they set.
Post reply on HN