Live data from Hacker News

Slack is down

status.slack.com

811–820 of 840 posts

Re: Slack is down

#811
post #89

When I was at Uber, we noticed that most incidents are directly caused by human actions that modify the state of the system. Therefore, a large "backlog" of human actions that modify the system state have a much higher chance of causing an incident. My bet is that this incident is caused by a big release after a post-holiday "code freeze".

If you think about it, modifications to state of the system caused by human actions are the sole purpose of computers.

Re: Slack is down

#812
post #126

Earlier quoted context omitted.

To elaborate a bit more on this point, you have to think about it like any complex system failure - it's almost never one thing, but rather a combination of many different factors. The factors around post NYE releases: - high risk changes that weren't released pre-holidays get released. Depending on the company, this could mean a 1-week to 1-month delay between implementation and release. The greater that interval, t…

People returning to work and downloading a huge backlog of messages from the past two weeks.

As an Microsoft Teams Ex-Dev, I can vouch that message retrieval after vacation puts a lot of stress on storage systems before it stabilizes :)

Re: Slack is down

#813

Earlier quoted context omitted.

It depends on the goal you’re trying to accomplish. Are you going for a promotion or bonus? Or instead is your goal to maximize uptime?

I doubt that regularly releasing breaking changes that reduce uptime is a good strategy to get a bonus or promotion.

Assume promotion after releasing 10 changes

Releasing 1 change a year with a 100% chance of working -- no promotion for 10 years

Releasing 10 changes a year each with a 10% chance of breaking something -- 1 in 3 chance of promotion in a year, and a 2 in 3 chance of downtime

Re: Slack is down

#814

Earlier quoted context omitted.

If new hires tends to break production, it's not in the first business day of the calendar year. December gets really quiet for recruiting, typically, as candidates get busy with their social lives, and scheduling interviews gets harder. January is busy for recruiting, but given a week or two of interviewing and negotiating, two weeks notice, it's probably February before new employees are starting, and they're not m…

I think some big company (maybe Facebook) has this rule that you had to deploy something to production on your first day. They seemed pretty confident in their processes and devops teams. A company trying to imitate that policy without doing the work necessary to make it possible would probably have outages on days when lots of new people joined :-P

Could be Facebook as I think production releases are always rolled out in phases e.g. first to 10 users, then 100, then 1000 and so on. That means there's much less chance of even the worst mistake having a serious effect.

Re: Slack is down

#815
post #804

Earlier quoted context omitted.

Just curious - why not use Matternost as a backup? (disclosure: I work at Mattermost, but really just want to know what you think) I’ve advocated for an idea where Mattermost is to be used as a “bunker” where it is hosted on a raspberry Pi (or somewhere else) and acts as a digital bunker if your critical infrastructure (slack, teams, exchange?) is compromised somehow.

Not OP. Good idea. I thought it's integrated into GitLab (on premise omnibus), but I still haven't fiddled with it, but enabled something in the config file, but nothing happened. I know it's a tough spot, but if it were usable from GitLab with zero config that would be great for fallback.

@pas: Mattermost is indeed integrated with GitLab Omnibus.

To enable Mattermost, you can add the Mattermost external URL in the config file, and run `sudo gitlab-ctl reconfigure`. I'm wondering if that's something you've tried? https://docs.gitlab.com/omnibus/gitlab-mattermost/#getting-s...

Re: Slack is down

#816

Earlier quoted context omitted.

If new hires tends to break production, it's not in the first business day of the calendar year. December gets really quiet for recruiting, typically, as candidates get busy with their social lives, and scheduling interviews gets harder. January is busy for recruiting, but given a week or two of interviewing and negotiating, two weeks notice, it's probably February before new employees are starting, and they're not m…

I think some big company (maybe Facebook) has this rule that you had to deploy something to production on your first day. They seemed pretty confident in their processes and devops teams. A company trying to imitate that policy without doing the work necessary to make it possible would probably have outages on days when lots of new people joined :-P

Wow, onboarding new hires here is going good, if they can access slack, O365, LDAP, VPN and clone the repo by the end of the first day. Tho we have the initiation ritual of installing the OS to your laptop.

Re: Slack is down

#817

Earlier quoted context omitted.

But clearly no one is saying that email is too hard to use and we should just use $something_else (or are they?). And you are starting to move goalposts here.. first it was uptime, then it was operations and now it's features... And what about those web based IRC solutions? They are even easier to use than slack, have combined history, file sharing, etc.

They are moving the goalposts because there are several and ultimately very many reasons why IRC won't work, they just didn't bother to think of all the reasons and list them at once.

Ultimately there is only one reason that matters: The person in charge of deciding what communication channel to use likes Slack/Teams/IRC/whatever.

Add to that the SaaS propaganda that hosting literally anything yourself is just too hard (it really isn't). Or this notion people are just too stupid to deal with anything more than the simplest possible web interface - Really? what do those people even do? Stare at Notepad all day? Of course not. They stare at various complicated software packages ranging from CAD, $spreadsheet abominations, SAP to various Adobe software packages. Sprinkle in a bit of hype for the latest new thing and presto..

Re: Slack is down

#818

I'd like to take this moment to mention self-hosted, open source, and federated alternatives like XMPP and Matrix. I'd like to, but unfortunately I don't feel like I can in good faith. Matrix is woefully immature, and suffers from a lot of issues, but I think is closer to being a functional Slack/Discord alternative. XMPP is much more mature, and works very well for chat, but doesn't have a nice package that does all…

Yep, I've stopped recommending Matrix because 1. There is virtually zero user-facing documentation. Need to know how to backup keys, verify another user, or what E2EE means? Ask your server operator. Basically the onus is on operators to document this stuff for their users. Except the stuff we're documenting is hard even for server operators, and especially challenging to document in a way that both nontechnical and…

Have you submitted the requested bug reports?

Also, it seems the FAQ answers several of your points: https://element.io/help

Re: Slack is down

#819
post #661

I'd like to take this moment to mention self-hosted, open source, and federated alternatives like XMPP and Matrix. I'd like to, but unfortunately I don't feel like I can in good faith. Matrix is woefully immature, and suffers from a lot of issues, but I think is closer to being a functional Slack/Discord alternative. XMPP is much more mature, and works very well for chat, but doesn't have a nice package that does all…

> self-hosted How often is Slack/Discord down? I mean it's not perfect, but I really honestly don't think I could match their uptime by self-hosting, as well as more on-call rotations for something that's not core product. I very much prefer that for something that isn't core product, if it goes down I need to do exactly nothing for it to come back up, and that the engineers at Slack will be starting to work on it li…

> it's not perfect, but I really honestly don't think I could match their uptime by self-hosting

This is such a common misconception. The services I self-host was configured by me, if anything goes down (which they very rarely do), I know the exact cause and have it fixed in minutes. When some company's cloud service goes down I'm completely at their mercy. I also spend very little time on maintaining these services, just security updates, which are mostly automated.

Bottom line, maintaining and self-hosting services that has 1 or a few users is much less complex than services with millions of users. Hence, my uptime is better than Google's, Amazon's, and Azure's, etc.

Re: Slack is down

#820
post #768

Earlier quoted context omitted.

Is it really though? If I take a look at a random modern IRC desktop client - how is it more difficult to setup than say your email program? The amount of information needed on setup is about the same: server, username, password (in fact email can get a bit more confusing in big corporate email setups with differing imap and smtp servers, etc.) Also there are plenty of modern web clients for IRC, such as https://thel…

> your email program Reality check: Most people don't use email programs anymore. Also how do you get IRC to sync all conversation data, history, between your several desktops and phones, how do you send files, make calls, and thread conversations?

> Reality check: Most people don't use email programs anymore.

Guess you are not in enterprise.

Post reply on HN