Live data from Hacker News

Slack outage: Connectivity issues affecting all workspaces

status.slack.com

231–240 of 278 posts

Re: Slack outage: Connectivity issues affecting all workspaces

#231
post #207

Earlier quoted context omitted.

It's not about assigning blame, it's about sharing lessons learned with the broader community and being transparent and honest with paying customers about issues that may have significant impact on downstream productivity.

It’s not about assigning blame for the company writing the post-mortem. But it’s definitely about assigning blame for most people reading the post-mortem. Very few people read post-mortems for the sake of learning how to be better at release engineering and ops.

If I pay for your service, and you are transparent about mistakes and flaws, I will be more forgiving about mistakes and flaws in the future, and appreciate the work you do to fix them.

If I pay for your service, and the only communication is, "We know there is a problem, and we'll let you know when it's fixed", I may assume you are not equipped to thoroughly explain the problem, and therefore not well equipped to solve it.

The blame is already assigned. The users already know there is a problem. A post-mortem likely has a positive effect for the readers attitude toward the handling of the issue.

Re: Slack outage: Connectivity issues affecting all workspaces

#232
post #138
post #33

You know how much of the community uses one messaging system when 15 minutes after it going down, it has over 40 points on the front page! This says a lot about how it's a single point of failure in modern company comms. It's even worrying to think about how some users probably have production-dependent (dare I postulate it) workflows in Slack that get crippled by its outage... ITT: Chat about decentralisation that w…

Sure, it's worrying but worth it for me personally. I might go to jail due to this (seriously) but at least people won't die. For me that's the threshold.

You can't leave us hanging like that. How could a Slack failure possibly send you to jail?

Re: Slack outage: Connectivity issues affecting all workspaces

#233

Maybe unrelated, but my AWS-hosted websockets-using app had an outage starting at the same time. Also a third-party API provider we use for handling inbound phone calls. So this smells like a wider outage than just Slack.

When I was in Moscow a few weeks back, Slack wouldn't work. Exact same behaviour - it loaded up the gui, loaded up previous conversations, but then wouldn't work past there.

Russia blocks a lot of AWS IPs, when I did a full VPN out to a server in Germany slack came good.

Re: Slack outage: Connectivity issues affecting all workspaces

#234
Seems to be fixed... with zero info on their status page about what went wrong or otherwise.

>We're happy to report that workspaces should be able to connect again, as we've isolated the problem. Some folks may need to refresh (Ctrl + R or Cmd + R). If you're still experiencing issues, please drop us a line

Hilariously, their "uptime in the last 30 days" still shows 100%.

Re: Slack outage: Connectivity issues affecting all workspaces

#235

Earlier quoted context omitted.

I think it's inexcusable for a chat program to go down in 2018. * your hdd failed? Use a raid * your power went out? Use a UPS * your DNS went down? Use a fallback (slack2) * your whole datacenter flooded? Good thing you have multiple replicated cloud instances that seamlessly take over See, these are the issues that "the cloud" was supposed to solve. Not give us the same problems as before, just with a recurring bil…

Let me add more reasons: 1) Software human mistake, when some software error/exception throws much larger issues, that require manual restore with service downtime. 2) Geodistributed datacenters is VERY expensive thing, so not implemented fully. 3) Bad system design, full of "one point of failure".

> ) Geodistributed datacenters is VERY expensive thing, so not implemented fully

You buy servers on aws-us-west and aws-us-east, and sync them . How is that very expensive?

Re: Slack outage: Connectivity issues affecting all workspaces

#236

Earlier quoted context omitted.

I think it's inexcusable for a chat program to go down in 2018. * your hdd failed? Use a raid * your power went out? Use a UPS * your DNS went down? Use a fallback (slack2) * your whole datacenter flooded? Good thing you have multiple replicated cloud instances that seamlessly take over See, these are the issues that "the cloud" was supposed to solve. Not give us the same problems as before, just with a recurring bil…

>slack is the most trivial software you can think of This is like saying that food service at 30k feet in a passenger airline is trivial because all the server has to do is walk up and down a narrow aisle handing out food from a cart. Since "you see no reason this service can't be nearly as reliable as life support firmware", one of two things must be true: 1) You know something nobody else knows. In which case great…

Personally I am going to be rich... Plebs

Re: Slack outage: Connectivity issues affecting all workspaces

#237
post #33

You know how much of the community uses one messaging system when 15 minutes after it going down, it has over 40 points on the front page! This says a lot about how it's a single point of failure in modern company comms. It's even worrying to think about how some users probably have production-dependent (dare I postulate it) workflows in Slack that get crippled by its outage... ITT: Chat about decentralisation that w…

I have used Mattermost and been pleased with it. It is an open-source Slack clone you can run on a low-end VM or your own hardware.

Re: Slack outage: Connectivity issues affecting all workspaces

#238

Earlier quoted context omitted.

Let me add more reasons: 1) Software human mistake, when some software error/exception throws much larger issues, that require manual restore with service downtime. 2) Geodistributed datacenters is VERY expensive thing, so not implemented fully. 3) Bad system design, full of "one point of failure".

> ) Geodistributed datacenters is VERY expensive thing, so not implemented fully You buy servers on aws-us-west and aws-us-east, and sync them . How is that very expensive?

You propose just to buy servers in 2 locations to keep Slack services up? Doesn't work, when you need to store gigabytes daily and have dozen thousand reqs/sec synchronized.

Geodistributed datacenter requires multiple direct low-latency multigigabit/sec connectivity, special software to manage, test and check it, skilled devops.

Re: Slack outage: Connectivity issues affecting all workspaces

#239
post #84
post #73

Earlier quoted context omitted.

I say this as someone who almost always prefers the dark theme wherever it is available: I wonder how much this desire for dark interfaces comes from almost every app interface having bright colors on white. Somewhere along the shift to flat design, grays and non-bright colors have been ignored in the visual design of applications.

In civil engineering circles, it's known that a room which is too bright will cause eye strain and fatigue. There is an optimal level of light for the eyes to be most effective. But the computer makers and UI designers don't take this into account. Dark themes transmit less light to the eyes, causing less fatigue over time.

I just installed Dark Mode for Firefox [1], it makes all websites have a dark theme. My eyes are already thanking me.

[1] https://addons.mozilla.org/en-US/firefox/addon/dark-mode-web...

Re: Slack outage: Connectivity issues affecting all workspaces

#240

Earlier quoted context omitted.

Yes, that way we can beat them up for years to come based on whatever mistake they made. It would be even better if they told us which employee made the mistake so we can incessantly mock that employee openly and publicly every time Slack is ever mentioned on HN. When GitHub was purchased by Microsoft, Gitlab came up quite a bit and we got to rehash that whole database outage over again many times over those few days…

I hope you don't work in aviation with that attitude!

As usual people are taking a comment and twisting it any old way they'd like. Which is fine, that's why we have these communications. To start off, no I am not in aviation. I have run quite a few companies and development departments.

I am not suggesting Slack or anyone else should not communicate at all when they have an outage. A public postmortem, which many people are asking for, is one method. Is it the most effective method? I doubt it. Many people are suggesting that as paying customers they would like to know what happened. Does a public postmortem tell the paying customer what happened in an effective way? Maybe, but maybe not.

When I am running a company I care very much what my paying customers think and are feeling about my service. I will communicate issues directly to them. Do I need to explain to the rest of the world in some great technical detail what happened during an incident? Absolutely not. Do I need to have the first post in Google about my company be an outage postmortem? Of course not. I need my PAYING customers to be pleased with the service I offer and to understand how I will mitigate the damage I have done to them. To me, that's a basic principle of business. I don't have to explain to everyone. I owe everything to my paying customers. Gitlab did a postmortem almost immediately after a major outage and some people tried to slaughter them with the information they shared. It was sad and unfortunate. Their openness was met with some horrible results from the community.

Also, I use Slack. My company uses it for everything including ChatOps for my production environment deployment. We have a hundred of so active users. The outage this morning harmed us. But you know what? I don't pay for Slack. I owe a lot to Slack but they don't owe me anything. I can't blame them for my problems this morning. They are a free service to me. I appreciate that their absolutely free service servers my company so well almost all of the time.

Post reply on HN