Live data from Hacker News

Slack is down

status.slack.com

181–190 of 840 posts

Re: Slack is down

#181
Now an outage -------------

We're continuing to investigate connection issues for customers, and have upgraded the incident on our side to reflect an outage in service. All hands are on deck on our end to further investigate. We'll be back in a half hour to keep you posted.

Jan 4, 8:20 AM PST

Re: Slack is down

#182
post #87

I'll concede that it's possible to not know what the problem is by now, but I won't concede that this should not be called an "outage" at this point.

I initially misread this as saying that you won't concede it shouldn't be called an outrage.

Re: Slack is down

#183
post #170

Looks like Slack just updated their status page to show a complete outage, not just an incident for "Messaging" and "Connections."

IMO it's really poor Slack took 1 hour to update this to an outage, given the impact this seemingly had right from the off.

It's also extremely bad that we're 1 hour in, and they are still "investigating", with no more details than that.

Re: Slack is down

#185
post #127

Earlier quoted context omitted.

If something went awry, and it caused more pain because Slack was down, how would you feel? If you’re missing comms/observability then waiting to deploy seems prudent.

So the answer is... people should stop working?

I can't speak for you, but I can:

  * work on code
  * update JIRA
  * complete required trainings
  * work on my peer reviews (Workday/Okta are up)
  * review tech specs
Deploying is actually a very small part of my job.

Re: Slack is down

#188
Atleast their status page works ¯\_(ツ)_/¯ (Looking at you AWS).

I am really looking forward to a better competitor taking over their market share, I presume things will only get worse after Salesforce acquisition.

Re: Slack is down

#189
post #89

When I was at Uber, we noticed that most incidents are directly caused by human actions that modify the state of the system. Therefore, a large "backlog" of human actions that modify the system state have a much higher chance of causing an incident. My bet is that this incident is caused by a big release after a post-holiday "code freeze".

This is very likely a broken release. The timing lines up with pacific time too well.

They declared the issue at 7:14AM PST. How long is their deploy process?

That sounds pretty early to think somebody on the west coast did something, other than maybe acknowledge the pages and declare the incident.

Re: Slack is down

#190

Earlier quoted context omitted.

A lot of organizations essentially took the last two weeks off from work, which is long enough for a 10-day autoscale window to spin down servers, and then get confronted by a load spike that wasn't pre-spun for.

I would be shocked if Slack operations wasn’t aware of this return to work spike and didn’t pre-scale in anticipation.

That doesn’t mean they chose the right number to scale to.

See for example, Amazon Prime day:

https://www.cnbc.com/2018/07/19/amazon-internal-documents-wh...

Post reply on HN