Live data from Hacker News

Slack is down

status.slack.com

371–380 of 840 posts

Re: Slack is down

#371
post #258

At GitLab our fallback from Slack is Zoom https://about.gitlab.com/handbook/communication/#emergency-c... I'm posting this because I found a lot of people don't know that Zoom includes a complete chat client that includes channels. And #HugOps to the engineers at Slack working on this. I appreciate that they posted a periodic update even when there was no news to report: "There are no changes to report as of yet. We'…

Aren't you worried about so many security vulnerabilities found in Zoom?

Re: Slack is down

#373
post #89

When I was at Uber, we noticed that most incidents are directly caused by human actions that modify the state of the system. Therefore, a large "backlog" of human actions that modify the system state have a much higher chance of causing an incident. My bet is that this incident is caused by a big release after a post-holiday "code freeze".

This is one of the original concepts why to go capital-A Agile. Make smaller releases more often, so at least if something breaks, it's (hopefully) something small, and least it's easier to trace.

(I'm not making a statement if that's good or bad or if it works or whatever. Please don't read an opinion into it.)

Re: Slack is down

#374

For some reason, today, HN seems exceedingly exceedingly slow (to me) after logging in... Without being logged in, things are as fast as they usually are -- but post log-in, SLOWWWER THAN MOLASSESS ... I tried this several times; why this is, I can only wonder... To quote Bill and Ted... "Strange things are afoot at the Circle-K..."

It is easier to cache stuff for users who are not logged in as it is the same for everyone. and everyone is looking up on Hackernews at the moment to see what is wrong with slack, which is probably the cause of the slowness.

Re: Slack is down

#375

Having been in this situation before, with a totally-down-and-not-coming-back-up outage of a payments system, I really feel for their incident response team. I'll take this moment to remind everyone of their human tendency to read meaning into random events. There's no evidence to suggest New Year traffic has caused this, and outages like this can happen in spite of professional and competent preparation. Hugops for…

> I'll take this moment to remind everyone of their human tendency to read meaning into random events. There's no evidence to suggest New Year traffic has caused this, and outages like this can happen in spite of professional and competent preparation.

On the one hand, sure we don't specifically know what's going on. On the other hand, it's the first Monday in the new year and they went down shortly after the start of the business day Eastern time; it could be coincidence, but it would be a remarkable coincidence.

Re: Slack is down

#376

> Customers may have trouble connecting or using Slack I can't stand how marketing speak pervades every sphere of the world. Their entire system is offline (inconvenient certainly, but it happens) and they can't bring themselves to say "Slack is down. We're working on it and will be back ASAP." or something similar. Instead we may have trouble.

I'm still logged in on mobile and can communicate with people from my team, but cannot log in from desktop. With so few people able to connect, it's also unclear whether Slack is eating my messages or there's just no one to respond. So I'd certainly rank that as "trouble using slack" rather than "the system is completely down".

Re: Slack is down

#377
post #126
post #89

When I was at Uber, we noticed that most incidents are directly caused by human actions that modify the state of the system. Therefore, a large "backlog" of human actions that modify the system state have a much higher chance of causing an incident. My bet is that this incident is caused by a big release after a post-holiday "code freeze".

To elaborate a bit more on this point, you have to think about it like any complex system failure - it's almost never one thing, but rather a combination of many different factors. The factors around post NYE releases: - high risk changes that weren't released pre-holidays get released. Depending on the company, this could mean a 1-week to 1-month delay between implementation and release. The greater that interval, t…

I would add here the potential scaling issue - holidays were a dry season - less meeting. So if they have some automation for scaling down to reduce cost, it may have bitten them in their arses now.

People came back to work, and most of them start around the same time (US wise at least).

Hence kids - a vital lesson for all of us - don't start the call at a full hour, give it 3-7 min to make your coworkers confused and give some time for the systems to auto-scale ;)

Re: Slack is down

#379

Earlier quoted context omitted.

How do you know it's down completely? Maybe it's down for you and maybe even down for a majority but still up for some subset. Happens with many products.

https://status.slack.com/ Every service is marked as "Outage" as of now (also when I wrote the comment).

yet also "Uptime for the current quarter: 100%"

Re: Slack is down

#380

> Customers may have trouble connecting or using Slack I can't stand how marketing speak pervades every sphere of the world. Their entire system is offline (inconvenient certainly, but it happens) and they can't bring themselves to say "Slack is down. We're working on it and will be back ASAP." or something similar. Instead we may have trouble.

Why do you consider that to be "marketing speak?" It appears to be concise, direct, and accurate. The phrase "Slack is down," even if true by some interpretations (it hasn't been "completely down" from what I have seen), is imprecise and informal.
Post reply on HN