Live data from Hacker News

Slack outage: Connectivity issues affecting all workspaces

status.slack.com

261–270 of 278 posts

Re: Slack outage: Connectivity issues affecting all workspaces

#261
post #201

Maybe unrelated, but my AWS-hosted websockets-using app had an outage starting at the same time. Also a third-party API provider we use for handling inbound phone calls. So this smells like a wider outage than just Slack.

Telegram was down too just half an hour before slack. Dunno if they run on aws?

They do, and GCP.

I recall AWS/GCP public IPs getting banned in Russia when they were trying to block telegram.

Re: Slack outage: Connectivity issues affecting all workspaces

#262
post #225

Earlier quoted context omitted.

IRC netsplits quite a bit, breaking ongoing conversations until it’s recovered, to be fair.

Small ircds that you would run for a single team don't split because it's a single server. Large networks can have the servers go up and down, and it's still not a big deal because of redundancy. DNS round-robin entries mean you don't even have to know the other servers on the network. In 2018 netsplits caused by down links are fairly rare. If you wait six months you might see one.

And if you run a small single-point ircd, at some point, the server or it’s internet connection will fail, and you’re in the same position as when Slack fails.

There’s nothing that gets around technical failure. Either you have a single server that’s going to die at some point due to sheer entropy, or you have a somewhat complex distributed system with the tradeoffs you desire that might fail anyway.

Re: Slack outage: Connectivity issues affecting all workspaces

#263

Earlier quoted context omitted.

Facebook springs to mind as well.

Facebook breaks features very often. Sometimes things go missing and comes back a week later. Dropbox does this a lot too.

It's the same for all sites beyond a certain size. It's never fully up. It's very rarely fully down. It's gradually degraded in ways that you hopefully don't see, but sometimes do. Or maybe you don't see it, but others do. etc etc etc. Availability isn't boolean once you have users.

Re: Slack outage: Connectivity issues affecting all workspaces

#264

Earlier quoted context omitted.

> ) Geodistributed datacenters is VERY expensive thing, so not implemented fully You buy servers on aws-us-west and aws-us-east, and sync them . How is that very expensive?

I imagine you've never actually had to solve any of these hard problems, which is why you think it's so easy to do.

That's bordering on (if not crossing into) ad-hominem.

There was no accusation of "so easy", only so not expensive and supposedly (and previously, demonstraby) solved in the last 30 years.

They may well be "hard" or even "expensive" for some definition of those two words, but if it weren't, it would defeat much of the (stated/advertised) purpose of outsourcing/cloud.

Re: Slack outage: Connectivity issues affecting all workspaces

#265

Earlier quoted context omitted.

Just hit the magic "sync" button. It's that easy!

I know. there's totally not a command called rsync. And "replication" is just a word you hear on star trek along with teleportation.

Although I agree with your premise, I think the delivery takes away from your point a bit.

Specifically, you risk people piling on that rsync isn't good enough in the modern world and referencing the comment criticizing Dropbox as being little more than an rsync replacement [1].

Of course, the specific tool one uses is irrelevant. The data synchronization problem may not be well solved, but it has been very well studied, with a remarkable number of good-enough options.

So, no, there isn't just one "sync" button, as the parent comment snarkily suggested, but there may be two, one where you might lose the last N seconds of chat (perhaps temporarily) and another where you lose the ability to chat entirely for those N seconds.

[1] Although it had other criticisms, such as monetization, which are, naturally, ignored.

Re: Slack outage: Connectivity issues affecting all workspaces

#266
post #42

Thats the problem without self-hosting your essential stuff.

Reasons my self-hosted servers have gone down in the past year: - Scheduled electrical maintenance that facilities manager failed to disclose (even though they knew about it for weeks). - Emergency power-down because two of the four air conditioners failed at the same time. - Someone accidentally powered off the VM. I'd much rather have an hour long outage here and there than incur the cost of defending against these…

>- Someone accidentally powered off the VM.

how is that self-hosting when you don't control the hypervisor in this case?

it usually implies that you at least have some sort of control. either having a real server somewhere (with ups and stuff) or at home, where you know when power is out.

while what you are doing is technically self-hosting, I would have changed the VM provider after the first incident like you described.

Re: Slack outage: Connectivity issues affecting all workspaces

#267

Earlier quoted context omitted.

As usual people are taking a comment and twisting it any old way they'd like. Which is fine, that's why we have these communications. To start off, no I am not in aviation. I have run quite a few companies and development departments. I am not suggesting Slack or anyone else should not communicate at all when they have an outage. A public postmortem, which many people are asking for, is one method. Is it the most eff…

My company does pay for slack, pays a lot, and I expect an RFO

Excellent! If you somehow read my entire message and got out of it that Slack shouldn’t give you detail about the outage this morning, then I somehow did not portray how important it is to emplain issues and resolutions to paying customers. I hope you get a full break down and understand exactly how they will keep you from having this sort of outage again. If they don’t, then it becomes a value issue to decide whether you should move to another system.

My point is only that it does not have to be a large public explanation. You, or the decision maker at your company, who pays a substantial sum of money to slack for their service, should have an explanation until you are satisfied.

Re: Slack outage: Connectivity issues affecting all workspaces

#268
post #82
post #10

Earlier quoted context omitted.

https://en.wikipedia.org/wiki/Netsplit

Anybody knows why the netsplit was written *.net *.split Why those stars and dots ?

It's to avoid showing the server names to the public on some IRC networks for various reasons including security through (some) obscurity.

Re: Slack outage: Connectivity issues affecting all workspaces

#269
post #115
post #52

Previous outages: https://news.ycombinator.com/item?id=16108912 - 5 months ago (longer discussion) https://news.ycombinator.com/item?id=15597461 - 7 months ago https://news.ycombinator.com/item?id=15597431 - 8 months ago https://news.ycombinator.com/item?id=13811815 - 1 year ago https://news.ycombinator.com/item?id=10616743 - 3 years ago

There is also this https://status.slack.com/calendar , but they seem to grossly under report the actual downtime... [edit] note that including this outage, they are reporting to have missed their monthly uptime guarantee 3 months in a row.

Yeah, stripe does the same thing with their status page. I get alerts that they have an outage at least once a week and more often than not it never shows up as anything in their history. Honestly this is my only significant beef with the service and I've been using it for years now with multiple integrations.

Re: Slack outage: Connectivity issues affecting all workspaces

#270
post #257

Earlier quoted context omitted.

If I pay for your service, and you are transparent about mistakes and flaws, I will be more forgiving about mistakes and flaws in the future, and appreciate the work you do to fix them. If I pay for your service, and the only communication is, "We know there is a problem, and we'll let you know when it's fixed", I may assume you are not equipped to thoroughly explain the problem, and therefore not well equipped to so…

It’s more the people who don’t pay for the service, but might, that are quickest to see post-mortems in a negative light. The only reason they have for reading them is looking for justifications for culling the product/service from the list of contenders for when they ever have to evaluate solutions in that category. In other words: post-mortems are good PR, but incredibly bad advertising .

And a world-wide outage followed by "we fixed it and trust us it won't happen again" is going to filter any service off of my list more so than "we had a single point of failure running in our CTO's basement and his cleaning lady pulled the plug. Trust us it won't happen again."
Post reply on HN