Live data from Hacker News

Slack – Degraded service affecting multiple features

status.slack.com

71–80 of 246 posts

Re: Slack – Degraded service affecting multiple features

#72

In other news: an unexplained productivity spike has been recorded today across the tech industry.

Or not. I'm remote and I cannot ask questions nor coordinate action to solve live production problems due to this outage. Slack is becoming a SPOF for many organizations, especially distributed.

Do phones exist?

Re: Slack – Degraded service affecting multiple features

#74
Can someone explain to me why any company should outsource something as critical as internal communication to a company that is 5 years old? The video messaging left aside, is it really that better than time-tested solutions like email or just running an IRC server?

Re: Slack – Degraded service affecting multiple features

#75

Earlier quoted context omitted.

You're implying we don't update the status board when there are errors?

The recent outage was very very delayed. I stand by my comment. Also don't get me wrong, you do a great service. It's just a pet peeve that it seems invariably status pages are a lie.

Let me quote from the previous HN discussion

”"”sauldcosta 4 days ago [-]

We use downdetector.com because status pages tend to take up to an hour or so to update, if they ever do. reply

jgrahamc 4 days ago [-]

1042 UTC First alert of global traffic problem 1057 UTC Internal group chat room up and running 1102 UTC Status page updated So, first alert to status page was 20 minutes.”””

In those 20m we had repeatedly checked your status page, realised it was our issue and started pulling engineers to deal with it as per procedure. People are on call, it's highly disruptive.

Surely you knew within those 20m that something was up?

In the end we realised it wasn't our issue because we checked Twitter.

Edit - added quotes

Re: Slack – Degraded service affecting multiple features

#76
post #31

There are reasons why email is async and supposed to be decentralized, with priority based fallback options. (The priority in mx records). Most email servers try to deliver to fallbacks, and even if that fails, will try to deliver for _days_. For all people who think slack can replace email, think about these safeguards.

Isn't it better in nearly every situation to know something has failed immediately, rather than get a delivery failure notice a couple of days later?

Re: Slack – Degraded service affecting multiple features

#77
post #74

Can someone explain to me why any company should outsource something as critical as internal communication to a company that is 5 years old? The video messaging left aside, is it really that better than time-tested solutions like email or just running an IRC server?

no, it's not. It's just a company based whatsapp group chat you can use while looking busy

Re: Slack – Degraded service affecting multiple features

#79

Earlier quoted context omitted.

The recent outage was very very delayed. I stand by my comment. Also don't get me wrong, you do a great service. It's just a pet peeve that it seems invariably status pages are a lie.

Let me quote from the previous HN discussion ”"”sauldcosta 4 days ago [-] We use downdetector.com because status pages tend to take up to an hour or so to update, if they ever do. reply jgrahamc 4 days ago [-] 1042 UTC First alert of global traffic problem 1057 UTC Internal group chat room up and running 1102 UTC Status page updated So, first alert to status page was 20 minutes.””” In those 20m we had repeatedly chec…

Yep. I tend to agree that we could have gone faster. The difficulty is that getting clear information out fast when you are dealing with a difficult situation is hard.

But, I guess we could have put some status up quicker.

Re: Slack – Degraded service affecting multiple features

#80

Earlier quoted context omitted.

Or not. I'm remote and I cannot ask questions nor coordinate action to solve live production problems due to this outage. Slack is becoming a SPOF for many organizations, especially distributed.

That is rather unfortunate if there is only a single channel that an organization can use in a production outage situation.

Well, there are other means of communication but they either:

* do not fit well for a quick message interchange like you can have during an outage (email)

* don't have the whole company (or at least the whole tech team) on it (whatsapp, telegram, you name it)

* make compliated to share big chunks of text/logs (phone call)

It probably makes sense creating another secondary - usually dormant - communication hub like a WhatsApp group for emergencies but people tend to misuse it (like using it to ping for some minor incident when all the other more suited communication tools are working perfectly).

Anyway it wasn't a big outage on our systems, just a minor hiccup, but it's a lesson learnt nonetheless.

Post reply on HN