Live data from Hacker News

IBM Cloud was down, as well as their status page

cloud.ibm.com

121–130 of 201 posts

Re: IBM Cloud was down, as well as their status page

#121
post #80

Earlier quoted context omitted.

[deleted]

I thought BGP shouldn't take longer than half an hour to propagate?

That's true however an external party might still be sending the bad advertisements that cause the issue. IBM can't really fix that.

Re: IBM Cloud was down, as well as their status page

#122
post #87

So what are HNers using IBM Cloud for and where do you see that it has an edge over AWS offerings (where an overlap exists, obviously)? (I figure either you’re in devops and you are putting out fires too busy to read this thread or you’re not and your work is halted because of the incident so you might have time to read and reply ;)

We used Softlayer (rebranded to IBM Cloud, and affected by this) at my last job. For the most part, their service pretty much just works; clearly not today. :) We had a couple thousand bare metal servers, and barely used any of their API stuff. As with any facility, there were occasional issues with electrical transfer switches, core router failures, fiber cuts, etc. Stuff happens, but we got pretty good communicatio…

Last job for me was also a few thousand bare metal servers at SoftLayer. Acquired and moved to that infrastructure instead. Wonder if its the same acquisition? :-)

Re: IBM Cloud was down, as well as their status page

#123

Earlier quoted context omitted.

Does it smell of fuckup, or of intentional BGP hijack?

If it was a BGP peer who normally sends you 3 prefixes with under a /20 in aggregate and they suddenly started sending you a whole table, or if you allowed a peer to send you a default route, then both of those are highly avoidable through session configuration and filtering. If the route which caused the madness came in via a large settlement-free peer (like a big eyeball/access network) or a transit (which is proba…

[deleted]

Re: IBM Cloud was down, as well as their status page

#124
post #114
post #87

Earlier quoted context omitted.

We used Softlayer (rebranded to IBM Cloud, and affected by this) at my last job. For the most part, their service pretty much just works; clearly not today. :) We had a couple thousand bare metal servers, and barely used any of their API stuff. As with any facility, there were occasional issues with electrical transfer switches, core router failures, fiber cuts, etc. Stuff happens, but we got pretty good communicatio…

You guys were one of the best use cases for the SL model, which really hasn't changed in 10+ years. You had very few dependencies on the less-reliable (read: all of them) services inside the SL stack and mostly managed everything on box and in software. In a few POPs you guys were running about 50% of the total SL backbone bandwidth. There were a lot of sad panda hats when you guys started to transition away.

> There were a lot of sad panda hats when you guys started to transition away.

For us as well. It was so nice to have things work one day and the next and the next, although I guess they wouldn't have worked today.

Favorite firefighting moment was when wdc lost half the fiber in ~ 2014, and we had to move all of our traffic out, so that there was capacity. Our guy asked why we had to move? and your guy said something like 'Because if you guys move, we only need one customer to move.' :D

Re: IBM Cloud was down, as well as their status page

#125

So what are HNers using IBM Cloud for and where do you see that it has an edge over AWS offerings (where an overlap exists, obviously)? (I figure either you’re in devops and you are putting out fires too busy to read this thread or you’re not and your work is halted because of the incident so you might have time to read and reply ;)

My previous job used Softlayer heavily.

Two of the biggest advantages were:

Price for hardware. As a base price, their bare-metal gear was significantly cheaper than equivalent-specced AWS gear (if it was even possible to get something like that). We managed to snag quite a few 'interesting' configurations of things at various times that you just couldn't get at all in AWS. Things like PCI SSDs, very large RAM configs, or High-Frequency low-core count CPUs.

Free international/regional transfer. We took significant advantage of this to move data around. We'd replicate TBs of data around.

At various times management and dev teams would complain and say that we should move everything to AWS (or whatever cloud provider they'd just met with at a conference).

We consistently showed higher performance and lower cost by significant margins. On cost alone, we were paying a small fraction of what it'd cost on AWS, even after taking into consideration ways to reduce cost on AWS such as scaling, spot instances and reserved-instances.

Re: IBM Cloud was down, as well as their status page

#126
post #123

Earlier quoted context omitted.

If it was a BGP peer who normally sends you 3 prefixes with under a /20 in aggregate and they suddenly started sending you a whole table, or if you allowed a peer to send you a default route, then both of those are highly avoidable through session configuration and filtering. If the route which caused the madness came in via a large settlement-free peer (like a big eyeball/access network) or a transit (which is proba…

[deleted]

That wasn't my point though. If you normally take 3 prefixes and suddenly you are receiving 800k+ prefixes then that likely classifies as "unexpected" (and avoidable since you can define max prefixes accepted per session).

I wasn't suggesting you could, I was questioning whether you should. :-)

Re: IBM Cloud was down, as well as their status page

#127
post #124
post #114

Earlier quoted context omitted.

You guys were one of the best use cases for the SL model, which really hasn't changed in 10+ years. You had very few dependencies on the less-reliable (read: all of them) services inside the SL stack and mostly managed everything on box and in software. In a few POPs you guys were running about 50% of the total SL backbone bandwidth. There were a lot of sad panda hats when you guys started to transition away.

> There were a lot of sad panda hats when you guys started to transition away. For us as well. It was so nice to have things work one day and the next and the next, although I guess they wouldn't have worked today. Favorite firefighting moment was when wdc lost half the fiber in ~ 2014, and we had to move all of our traffic out, so that there was capacity. Our guy asked why we had to move? and your guy said something…

Yeah the move from FreeBSD to Linux wouldn't have been fun for you guys either. And yeah, the WDC POPs were some of the most overbuilt from a bandwidth perspective and that was almost entirely because of you guys. Pretty sure there's a Cisco sales rep enjoying a nice holiday home in Connecticut as a result of the growth you guys did.

Re: IBM Cloud was down, as well as their status page

#129

Honest slightly cynical question: most probably someone inside the responsible team said some day that it would be very stupid to host the status page inside the same infrastructure being monitored, but they were probably ignored... what should that person do now? Say "toldya!" out loud in the postmortem meeting or simply shut up and move on because reality is that we are hired to do some stupid task and not to think…

Never humiliate a coworker in public. Instead say "both options were considered but ultimately it was decided to select option B for reason Y."

What if Y = ignorance?
Post reply on HN