Live data from Hacker News

Slack is down

status.slack.com

781–790 of 840 posts

Re: Slack is down

#781

I'd like to take this moment to mention self-hosted, open source, and federated alternatives like XMPP and Matrix. I'd like to, but unfortunately I don't feel like I can in good faith. Matrix is woefully immature, and suffers from a lot of issues, but I think is closer to being a functional Slack/Discord alternative. XMPP is much more mature, and works very well for chat, but doesn't have a nice package that does all…

I can only offer my own personal experience: Matrix has been working well for me for a couple years now. However, I probably have a more narrow use case than you're thinking of. I run a small homeserver and use it to communicate with a group of about 20 friends. Most of them aren't "technical" people. We use it mostly for chatting and image/video sharing. We never use live calling (audio or video). There have been a…

    and use it to communicate with a group of about 20
    friends. Most of them aren't "technical" people. 
I'm insanely curious about the human side of things here. How did you get them to buy into this idea in the first place? That sounds like quite an achievement.

The non-technical folks in my life generally struggle with paths of least resistance (iMessage, etc) and it's hard to imagine getting them onto some alternative platform/protocol.

Re: Slack is down

#782
post #780

Earlier quoted context omitted.

I agree with both of you, and TMTP supports adding people to a thread after it starts (see PostNotify).

I'm not finding much about TMTP or postnotify with a search through Google. Could you link to some resources?

I've only just begun publicizing it, after getting the client & server implementations to a point where folks can evaluate them.

Protocol: https://github.com/networkimprov/mnm/blob/master/Protocol.md

Why TMTP? https://mnmnotmail.org/rationale.html

Follow: https://twitter.com/mnmnotmail

Re: Slack is down

#783
post #126
post #89

When I was at Uber, we noticed that most incidents are directly caused by human actions that modify the state of the system. Therefore, a large "backlog" of human actions that modify the system state have a much higher chance of causing an incident. My bet is that this incident is caused by a big release after a post-holiday "code freeze".

To elaborate a bit more on this point, you have to think about it like any complex system failure - it's almost never one thing, but rather a combination of many different factors. The factors around post NYE releases: - high risk changes that weren't released pre-holidays get released. Depending on the company, this could mean a 1-week to 1-month delay between implementation and release. The greater that interval, t…

People returning to work and downloading a huge backlog of messages from the past two weeks.

Re: Slack is down

#784

I'd like to take this moment to mention self-hosted, open source, and federated alternatives like XMPP and Matrix. I'd like to, but unfortunately I don't feel like I can in good faith. Matrix is woefully immature, and suffers from a lot of issues, but I think is closer to being a functional Slack/Discord alternative. XMPP is much more mature, and works very well for chat, but doesn't have a nice package that does all…

it's trivial and essentially free (it also doesn't take gigs of memory on client devices.

Re: Slack is down

#785
post #661

I'd like to take this moment to mention self-hosted, open source, and federated alternatives like XMPP and Matrix. I'd like to, but unfortunately I don't feel like I can in good faith. Matrix is woefully immature, and suffers from a lot of issues, but I think is closer to being a functional Slack/Discord alternative. XMPP is much more mature, and works very well for chat, but doesn't have a nice package that does all…

> self-hosted How often is Slack/Discord down? I mean it's not perfect, but I really honestly don't think I could match their uptime by self-hosting, as well as more on-call rotations for something that's not core product. I very much prefer that for something that isn't core product, if it goes down I need to do exactly nothing for it to come back up, and that the engineers at Slack will be starting to work on it li…

There will always be more downtime on Slack/Discord. There are more users, more updates. Slack/Discord is a giant distributed system with nodes all around the world. An IRC/XMPP server on one machine that 100 people use is not going to crash unless intentionally.

Re: Slack is down

#786

Earlier quoted context omitted.

It has, and I’ve been using it since its early days. I still use it. It’s still terrible, just slightly less terrible. And, no, messages don’t consistently send in 100ms on the default home server; there are regularly disruptions that cause significant delays, sometimes as much as 10-20sec. That’s a big problem for a federated chat platform. Edit 1: I want to love it; the design is everything I could ever hope for in…

It's weird that you're calling it Vector when it's now called Element and it was called Riot for years before that.

The constant rebranding and confusion over Matrix/Vector/Riot/Element is another point of pain for me. It’s incredibly difficult to communicate unambiguously about Matrix with people who haven’t been following it for years.

Does Element refer to the ecosystem as a whole, including EMS? The primary client? The core federation? It’s not obvious from a casual visit to element.io. I suppose if I said “Element web app,” that would be fairly clear, but I’m still in the habit of saying “Vector” from the days of Riot.

Re: Slack is down

#787
So many large scale downtimes across multiple large companies in the past month or so. Is this for a bugfix deployment for the SolarWinds hack, or downtime caused by the hack itself ? Or some state-sponsored orgs installing upgraded eavesdropping stuff ?

Re: Slack is down

#788

Earlier quoted context omitted.

That’s some really good thoughts on DR planning. I have never thought DR to be to such an extent. How many companies really plan for an event where their entire infrastructure goes offline and their entire team gets killed? Does even companies like Google plan for this kind of event?

> How many companies really plan for an event where their entire infrastructure goes offline and their entire team gets killed? Since 9/11, more than you might think. For example Empire Blue Cross Blue Shield [1] had its HQ in the WTC. https://www.computerworld.com/article/2585046/empire-blue-[1... cross-it-group-undaunted-by-wtc-attack--anthrax-scare.html

Fixed link: https://www.computerworld.com/article/2585046/empire-blue-cr...

And what a blast from the past:

> Some of the temporary locations, such as the W Hotel, required significant upgrades to their network infrastructure, Klepper said. "We're running a Gigabit Ethernet now here in the W Hotel,'' Klepper said, with a network connected to four T1 (1.54M bit/sec) circuits. That network supports the code development for a Web-based interface to the company's systems, which Klepper called "critical" to Empire's efforts to serve its customers. Despite the lost time and the lost code in the collapse of the World Trade Center towers, Klepper said, "we're going to get this done by the end of the year."

> Shevin Conway, Empire's chief technology officer, said that while the company lost about "10 days' worth" of source code, the entire object-oriented executable code survived, as it had been electronically transferred to the Staten Island data center.

Re: Slack is down

#789
post #126

Earlier quoted context omitted.

To elaborate a bit more on this point, you have to think about it like any complex system failure - it's almost never one thing, but rather a combination of many different factors. The factors around post NYE releases: - high risk changes that weren't released pre-holidays get released. Depending on the company, this could mean a 1-week to 1-month delay between implementation and release. The greater that interval, t…

If new hires tends to break production, it's not in the first business day of the calendar year. December gets really quiet for recruiting, typically, as candidates get busy with their social lives, and scheduling interviews gets harder. January is busy for recruiting, but given a week or two of interviewing and negotiating, two weeks notice, it's probably February before new employees are starting, and they're not m…

I think some big company (maybe Facebook) has this rule that you had to deploy something to production on your first day. They seemed pretty confident in their processes and devops teams. A company trying to imitate that policy without doing the work necessary to make it possible would probably have outages on days when lots of new people joined :-P

Re: Slack is down

#790
post #641

Earlier quoted context omitted.

RocketChat works pretty well for simple team comms. I have no idea if it can do XMPP and/or Matrix.

I suggested RocketChat when the outage was announced and HN community downvoted it quite heavily. I'm not sure why. [0] We ended making the switch and committed to Discord. We're now looking at Rocket.chat as a backup in case Discord goes down. But Slack is now completely out of the picture for our team. [0] https://news.ycombinator.com/item?id=25633047

Just curious - why not use Matternost as a backup? (disclosure: I work at Mattermost, but really just want to know what you think)

I’ve advocated for an idea where Mattermost is to be used as a “bunker” where it is hosted on a raspberry Pi (or somewhere else) and acts as a digital bunker if your critical infrastructure (slack, teams, exchange?) is compromised somehow.

Post reply on HN