Live data from Hacker News

Slack is down

status.slack.com

441–450 of 840 posts

Re: Slack is down

#441

> Customers may have trouble connecting or using Slack I can't stand how marketing speak pervades every sphere of the world. Their entire system is offline (inconvenient certainly, but it happens) and they can't bring themselves to say "Slack is down. We're working on it and will be back ASAP." or something similar. Instead we may have trouble.

I agree with you in principal, but I have had no problem connecting to Slack today (I have a free one I use with friends, not a business account) so to say they are down would also be inaccurate.

Re: Slack is down

#442

> Customers may have trouble connecting or using Slack I can't stand how marketing speak pervades every sphere of the world. Their entire system is offline (inconvenient certainly, but it happens) and they can't bring themselves to say "Slack is down. We're working on it and will be back ASAP." or something similar. Instead we may have trouble.

The funniest part to me is that their status page still says "Uptime for the current quarter: 100%". These uptime messages are so BS. Heroku reports 6 9s of uptime for this month, even though their own status page shows multiple days with incidents >6 hours

Re: Slack is down

#443

Earlier quoted context omitted.

That scenario is what Disaster Recovery plans are for. Every large company I've worked for has had recovery plans in place, including scenarios as disturbing as "All data centers and offices explode simultaneously, and all staff who know how it all works are killed in the blasts." You not only have backups in place, you have documentation in place, including a back-up vendor who has copies of the documentation and ca…

That’s some really good thoughts on DR planning. I have never thought DR to be to such an extent. How many companies really plan for an event where their entire infrastructure goes offline and their entire team gets killed? Does even companies like Google plan for this kind of event?

Yes, Google plans extensively and runs regular drills.

It's hearsay, but I was once told that achieving "black start" capability was a program that took many years and about a billion dollars. But they (probably) have it now.

Re: Slack is down

#444

> Customers may have trouble connecting or using Slack I can't stand how marketing speak pervades every sphere of the world. Their entire system is offline (inconvenient certainly, but it happens) and they can't bring themselves to say "Slack is down. We're working on it and will be back ASAP." or something similar. Instead we may have trouble.

"We're experiencing increased service degradation" is so 201x-ish

I find it hilarious that the status page is still saying the uptime for the current quarter is 100%. I'd think it'd have lost at least one 9 by any obvious definition of "current quarter".

Re: Slack is down

#446
post #343

Earlier quoted context omitted.

What an overwrought headline, the employee in question has already been fired.

It's weird that you describe the headline as "overwrought" and call the person an "employee" when the headline is more accurate than you. This was an executive, not just an employee. That's a huge distinction and I can't help but think you intentionally downgraded his position to cover-up his behavior. "Just an employee" "Not a big deal" But when you read the allegations, they seem like a very big deal that an execut…

Regardless of intent, it's undeniable that at some point there were insufficient controls to prevent this executive, or any executive in the future, from gaining this level of surveillance access.

And it's also undeniable that the consequences for Zoom (really, just needing to fire a few people, and not even the people who designed those controls if there were any) are so minimal that they have no incentive to strengthen those controls.

For some organizations (mine included) the benefits of Zoom outweigh the risks of Zoom having proven itself to not have those controls, namely the possibility of both political and corporate espionage. As with all things, YMMV.

Re: Slack is down

#447

> Customers may have trouble connecting or using Slack I can't stand how marketing speak pervades every sphere of the world. Their entire system is offline (inconvenient certainly, but it happens) and they can't bring themselves to say "Slack is down. We're working on it and will be back ASAP." or something similar. Instead we may have trouble.

Why do you consider that to be "marketing speak?" It appears to be concise, direct, and accurate. The phrase "Slack is down," even if true by some interpretations (it hasn't been "completely down" from what I have seen), is imprecise and informal.

It was pretty clear the 'may' is a euphemism when your whole system is down.

Re: Slack is down

#448

> Customers may have trouble connecting or using Slack I can't stand how marketing speak pervades every sphere of the world. Their entire system is offline (inconvenient certainly, but it happens) and they can't bring themselves to say "Slack is down. We're working on it and will be back ASAP." or something similar. Instead we may have trouble.

How do you know it's down completely? Maybe it's down for you and maybe even down for a majority but still up for some subset. Happens with many products.

Yeah, I don't know, because the Slack status page is so vague.

Re: Slack is down

#449
post #370
post #245

This site seems to be lagging as well. Or is it just me?

It isn’t just you, but due to Hacker News’ right-sized[0] infrastructure, you should sign out unless you need to comment. That way you hit the caches instead of getting the server to make you a new page. 0: https://news.ycombinator.com/item?id=12911461

That’s the wrong way to look at it. If HN struggles in certain situations then it is not right-sized. You don’t beg of users to walk an unintuitive happy path (i.e. logout when not commenting).

Re: Slack is down

#450
post #89

When I was at Uber, we noticed that most incidents are directly caused by human actions that modify the state of the system. Therefore, a large "backlog" of human actions that modify the system state have a much higher chance of causing an incident. My bet is that this incident is caused by a big release after a post-holiday "code freeze".

I have definitely worked in places where the times right before and right after a change freeze were the most unstable, so that could be it. However, as others have mentioned, it's pretty early on the west coast of the US. Unless some engineer was up extra early (perhaps at the behest of an anxious project manager) it seems unlikely to be a release. What it could be is some engineer somewhere coming in after the holi…

Most releases are automated with time lockouts.
Post reply on HN