Live data from Hacker News

Slack – Degraded service affecting multiple features

status.slack.com

61–70 of 246 posts

Re: Slack – Degraded service affecting multiple features

#61
post #16

Meta - can we not post outages to HN? I understand that post mortem's are interesting learning lessons, but if you're curious whether S3 or whatever is down or not right at this moment, then use the subsequent status pages of those services.

I get most of my outage notifications from HN. If it's a blip, it probably won't make the front page through my standard refresh cycle. If it persists through two refresh cycles, there's a TITSUP somewhere and it's time to take a look.

This community is awesome. Always someone with their finger on the pulse :)

Re: Slack – Degraded service affecting multiple features

#62
reminds me of that one time a genius at IBM created an `ibm-global-announcements` channel and force-invited 200'000 people in it, and then some guy `@channel`d and all ibm's slack workspaces were down for 30 minutes

https://status.slack.com/ibm/2018-03/f01d4c22cd953dd7

https://i.imgur.com/Rk6Kdgp.png

EDIT:

also that channel made using slack impossible for mac book air users, they had around 80% cpu usage for slack. so basically entire marketing and PM part was unable to work that day. developes machines were wasting around 20% on slack.

after people started complaining in that channel, posting in it was limited to admins only, but they didn't lock commenting. so, for approx 6 hours all of IBM was posting memes in THETHREAD as we dubbed it, and @mentioning the genius who created that channel. next day the channel was nuked, not even an archive preserved.

some guy calculated that the entire affair, considering electricity prices, a 20% decrease in developers productivity, and so on resulted in IBM loosing several millions with that stunt

fun times

p.s. shout out to Martinj for that spicy jeff-coffee-mug meme

Re: Slack – Degraded service affecting multiple features

#63
post #20

Earlier quoted context omitted.

I’ve hosted multiple IRC servers over the years and can’t remember ever having to do “maintaining” after the initial setup.

Honestly curious, how long did you last without having to: * update the version of the IRC server * update the version of the os of the machine running the IRC server * repair a broken disk / fan / overflowed disk on the machine * add / remove / reset password / change weird settings of users I'm not saying this is "unbearable", and plenty of organisations have people whose job description would probably correspond t…

>* update the version of the IRC server

Odds are that you’ll have updates worth installing once every couple of years.

>* update the version of the os of the machine running the IRC server

Almost never unless there’s a remotely exploitable code execution vulnerability. Local bugs wont matter unless you run multiple services on the same box.

>* repair a broken disk / fan / overflowed disk on the machine

Depends on your hosting setup. With a cloud setup perhaps never.

I’m certainly not trying to suggest that anyone should use IRC over slack in any situation, just that it’s not a horrible time sink requiring significant maintenance.

Re: Slack – Degraded service affecting multiple features

#65
post #52

Earlier quoted context omitted.

Or not. I'm remote and I cannot ask questions nor coordinate action to solve live production problems due to this outage. Slack is becoming a SPOF for many organizations, especially distributed.

Why can't you use email?

Because Slack is so much better than email!

Re: Slack – Degraded service affecting multiple features

#66
post #47

Earlier quoted context omitted.

Isn't internet a single communication medium anyway?

It is, technically. Good thing there's also phones and SMS.

I have heard a horror story involving a phone line multiplexed over fiber to an IX and someone used it as OOB.

Re: Slack – Degraded service affecting multiple features

#67

Earlier quoted context omitted.

Isn't internet a single communication medium anyway?

The internet is specifically designed to route around downed links such that there is no SPOF, that’s the theory anyway. However, when major backbones go down it takes it a while to recover.

Well, kinda! Just look at this recent CloudFlare blog post[1]. It's true, theres no Single Point of Failure - there are MANY points of failure ;)

[1]: https://blog.cloudflare.com/how-verizon-and-a-bgp-optimizer-...

Re: Slack – Degraded service affecting multiple features

#68
post #16

Meta - can we not post outages to HN? I understand that post mortem's are interesting learning lessons, but if you're curious whether S3 or whatever is down or not right at this moment, then use the subsequent status pages of those services.

Why? These events are often great opportunities to learn about system failures and their impact on other systems in real time. Status pages are often not even genuinely accurate.

Re: Slack – Degraded service affecting multiple features

#69

In other news: an unexplained productivity spike has been recorded today across the tech industry.

Or not. I'm remote and I cannot ask questions nor coordinate action to solve live production problems due to this outage. Slack is becoming a SPOF for many organizations, especially distributed.

Good thing email still exists, right?

Re: Slack – Degraded service affecting multiple features

#70

Earlier quoted context omitted.

You're implying we don't update the status board when there are errors?

The recent outage was very very delayed. I stand by my comment. Also don't get me wrong, you do a great service. It's just a pet peeve that it seems invariably status pages are a lie.

[deleted]
Post reply on HN