Live data from Hacker News

IBM Cloud was down, as well as their status page

cloud.ibm.com

91–100 of 201 posts

Re: IBM Cloud was down, as well as their status page

#91
post #65

So what are HNers using IBM Cloud for and where do you see that it has an edge over AWS offerings (where an overlap exists, obviously)? (I figure either you’re in devops and you are putting out fires too busy to read this thread or you’re not and your work is halted because of the incident so you might have time to read and reply ;)

I suspect nobody really uses it outside of weird outsourced financial modelling/planning tools like TM1 and other apps people stopped wanting to manage themselves.

I work as a consultant with big enterprise companies and I can assure you big enterprise companies are using IBM very heavily. As well as Oracle and HP and other uncool tech companies.

Re: IBM Cloud was down, as well as their status page

#92

Honest slightly cynical question: most probably someone inside the responsible team said some day that it would be very stupid to host the status page inside the same infrastructure being monitored, but they were probably ignored... what should that person do now? Say "toldya!" out loud in the postmortem meeting or simply shut up and move on because reality is that we are hired to do some stupid task and not to think…

Ideally, someone should write a postmortem with a timeline of what happened and recommended fixes. These would then be fixed, and nobody would be blamed. (This is called a "blameless postmortem".)

But whether you can get away with that depends on culture.

Re: IBM Cloud was down, as well as their status page

#93

Earlier quoted context omitted.

Never humiliate a coworker in public. Instead say "both options were considered but ultimately it was decided to select option B for reason Y."

My professional experience tells me that the next question will be who decided for B given Y, then you answer it and then you have a target on your SRE back, I'm afraid. Remember that the trickle-down economics works only when the shit hits the fans and what trickles down is not money.

[deleted]

Re: IBM Cloud was down, as well as their status page

#94

Seems pretty dumb to host a status page in a way that it could go down, when it should be a static page that is trivially hosted on CDN's worldwide.

Their status page also seems integrated into their internal support ticketing system. It's not a traditional status page. They wanted to maintain a consistent garbage interface to keep it inline with the rest of their administrative service.

Re: IBM Cloud was down, as well as their status page

#95

Honest slightly cynical question: most probably someone inside the responsible team said some day that it would be very stupid to host the status page inside the same infrastructure being monitored, but they were probably ignored... what should that person do now? Say "toldya!" out loud in the postmortem meeting or simply shut up and move on because reality is that we are hired to do some stupid task and not to think…

Never humiliate a coworker in public. Instead say "both options were considered but ultimately it was decided to select option B for reason Y."

So, the appeal to anonymous authority, I see

Re: IBM Cloud was down, as well as their status page

#97
post #90

Earlier quoted context omitted.

Do you actually have data on that or are you conjecturing? Because I would really love to see data about that if it exists somewhere.

I'm talking from experience. Most things do get post mortems, but there's a lot of crap they also don't give us post mortems for "because customer data." It's my number 1 complaint, and I fight with managers about this all the time. We have a ton of hypervisor problems, and a lot of networking issues (generally over private network) and they tend to get very very secretive about it.

IBM cloud specifically or just in general?

Re: IBM Cloud was down, as well as their status page

#98
post #97
post #90

Earlier quoted context omitted.

I'm talking from experience. Most things do get post mortems, but there's a lot of crap they also don't give us post mortems for "because customer data." It's my number 1 complaint, and I fight with managers about this all the time. We have a ton of hypervisor problems, and a lot of networking issues (generally over private network) and they tend to get very very secretive about it.

IBM cloud specifically or just in general?

I'm talking about IBM Cloud specifically, yes.

Re: IBM Cloud was down, as well as their status page

#99
post #79

Earlier quoted context omitted.

Lol yes. It's intentional. It stops spam bots from trying to sign up, since we aren't open for signups yet. But don't worry, you're not the first to mention it. I suppose I should just fix it and deal with the spam like normal. I liked the unintended effect of cutting down on spam. I guess a lot of spam bots are written on top of standard libraries that reject bad certs. :) Also, this was ironically a great way to pu…

This has to be the most unconventional anti-spam technique I've ever heard about.

> I stumbled on it by accident. I was lazy and let the cert lapse, but then noticed that spam signups basically stopped. One day maybe I'll make a post about it with graphs, although I'm not sure I actually have the data.

This is intriguing. I'm going to remember this but I'm too anal about perfect A+ TLS and renewal is already fully automated these days anyway :-\

I wonder if one could setup their TLS stack to get this effect without the tradeoff...

Re: IBM Cloud was down, as well as their status page

#100
post #48

Honest slightly cynical question: most probably someone inside the responsible team said some day that it would be very stupid to host the status page inside the same infrastructure being monitored, but they were probably ignored... what should that person do now? Say "toldya!" out loud in the postmortem meeting or simply shut up and move on because reality is that we are hired to do some stupid task and not to think…

I built a status page for a top cloud provider, and this was question number one from SREs.

I built a bunch of CloudWatch monitoring for an AWS stack, and duplicated critical monitoring using a 3rd party monitoring service as well. So, because the universe hates me, the 3rd party service migrated their hosting into AWS ~18 months later... :sigh:
Post reply on HN