Live data from Hacker News

Slack outage: Connectivity issues affecting all workspaces

status.slack.com

241–250 of 278 posts

Re: Slack outage: Connectivity issues affecting all workspaces

#241

Earlier quoted context omitted.

Yes it's a single point of failure, but so what? I don't particularly care whether other organizations fail at the same time as I do, I just care whether I fail. Hosting my own chat system does not solve that problem. In fact, it may make it worse because then I have to worry about system administration, and Slack probably has more expertise on that. It's likely that they can fix this problem for all customers faster…

I think it's inexcusable for a chat program to go down in 2018. * your hdd failed? Use a raid * your power went out? Use a UPS * your DNS went down? Use a fallback (slack2) * your whole datacenter flooded? Good thing you have multiple replicated cloud instances that seamlessly take over See, these are the issues that "the cloud" was supposed to solve. Not give us the same problems as before, just with a recurring bil…

>slack is the most trivial software you can think of

Your comment reeks of some serious beginner hubris.

Almost reads like a parody of the classic hubris dismissal we have to endure on HN.

Re: Slack outage: Connectivity issues affecting all workspaces

#242

Earlier quoted context omitted.

Let me add more reasons: 1) Software human mistake, when some software error/exception throws much larger issues, that require manual restore with service downtime. 2) Geodistributed datacenters is VERY expensive thing, so not implemented fully. 3) Bad system design, full of "one point of failure".

> ) Geodistributed datacenters is VERY expensive thing, so not implemented fully You buy servers on aws-us-west and aws-us-east, and sync them . How is that very expensive?

Just hit the magic "sync" button. It's that easy!

Re: Slack outage: Connectivity issues affecting all workspaces

#243
post #209
post #121

Earlier quoted context omitted.

In light of how Slack and other companies haven't been able to get a decent level of uptime, I have to say, the company known to make huge web applications that don't go down in shame every couple of months is probably Google. I can't remember the last time Gmail was down. It just works! If google is down, probably your internet is down. Their expertise and discipline in distributed applications is unrivaled. I'm gue…

Funny you should mention google, as something is down over there right now. lots of reports of chromecasts being dead right, assuming something at google is down which is causing this.

Oh, interesting. Thanks for pointing this out. I was having Chromecast trouble this morning and didn't even think to check if it was a widespread issue.

Re: Slack outage: Connectivity issues affecting all workspaces

#244

Earlier quoted context omitted.

Let me add more reasons: 1) Software human mistake, when some software error/exception throws much larger issues, that require manual restore with service downtime. 2) Geodistributed datacenters is VERY expensive thing, so not implemented fully. 3) Bad system design, full of "one point of failure".

> ) Geodistributed datacenters is VERY expensive thing, so not implemented fully You buy servers on aws-us-west and aws-us-east, and sync them . How is that very expensive?

I imagine you've never actually had to solve any of these hard problems, which is why you think it's so easy to do.

Re: Slack outage: Connectivity issues affecting all workspaces

#245

Earlier quoted context omitted.

> ) Geodistributed datacenters is VERY expensive thing, so not implemented fully You buy servers on aws-us-west and aws-us-east, and sync them . How is that very expensive?

Just hit the magic "sync" button. It's that easy!

I know. there's totally not a command called rsync. And "replication" is just a word you hear on star trek along with teleportation.

Re: Slack outage: Connectivity issues affecting all workspaces

#246

Earlier quoted context omitted.

> ) Geodistributed datacenters is VERY expensive thing, so not implemented fully You buy servers on aws-us-west and aws-us-east, and sync them . How is that very expensive?

Just hit the magic "sync" button. It's that easy!

Oh, yes, someone in Slack clicked the button “Pause” and we all are waiting, when Slack’s hero will click “Resume” :)

Re: Slack outage: Connectivity issues affecting all workspaces

#247

Earlier quoted context omitted.

It's not about assigning blame, it's about sharing lessons learned with the broader community and being transparent and honest with paying customers about issues that may have significant impact on downstream productivity.

I entirely understand what you are saying, believe me I do. But that is not the way some communities take it. We still see messages like "You could move to Gitlab but... you know they dropped their production database a couple of years back? Use them at your own risk!" We learned a lot from the Gitlab outage. It was a simple mistake and not one they will have again, yet people still beat them up for it. I'm not sure…

Wouldn't they have gotten beaten up over the outage even more had they not offered an explanation?

In my experience, customers are often seeking an explanation/post-mortem because their customers are seeking an explanation. If an upstream service goes down for an extended period of time and all you can do is go back to your customers and say, "Your system was down because our provider's system went down for 4 hours. But they won't tell us why.", that not go over well.

Re: Slack outage: Connectivity issues affecting all workspaces

#248

Earlier quoted context omitted.

I hope you don't work in aviation with that attitude!

As usual people are taking a comment and twisting it any old way they'd like. Which is fine, that's why we have these communications. To start off, no I am not in aviation. I have run quite a few companies and development departments. I am not suggesting Slack or anyone else should not communicate at all when they have an outage. A public postmortem, which many people are asking for, is one method. Is it the most eff…

My company does pay for slack, pays a lot, and I expect an RFO

Re: Slack outage: Connectivity issues affecting all workspaces

#249
post #52

Previous outages: https://news.ycombinator.com/item?id=16108912 - 5 months ago (longer discussion) https://news.ycombinator.com/item?id=15597461 - 7 months ago https://news.ycombinator.com/item?id=15597431 - 8 months ago https://news.ycombinator.com/item?id=13811815 - 1 year ago https://news.ycombinator.com/item?id=10616743 - 3 years ago

[deleted]

Re: Slack outage: Connectivity issues affecting all workspaces

#250

Earlier quoted context omitted.

So you are saying people should sow discord?

No. People should use IRC.

In practice, all the attempts I've seen at getting a significant number of non-techies in a company to use IRC have failed. At one company, we almost got everyone on Jabber, but it was never used much outside the tech circles. Slack? Everyone is using it and most seem happy about it, and it does a lot more than IRC.
Post reply on HN