Does this happen often with OVH?
OVH Incident in Strasbourg
171–180 of 207 posts
Re: OVH Incident in Strasbourg
#172Earlier quoted context omitted.
> - naaah it's fine, we do not have time nor resources" Yup, been there multiple times in smaller hosting companies. It's basically how it goes. They don't get serious about outages until revenue is severely affected and the brand damaged, they don't get serious about security until there's been a big breach or sales are lost because of lack of certification.
Calling OVH a "smaller hosting company" (which you did do indirectly) is rather funny.
Re: OVH Incident in Strasbourg
#173Re: OVH Incident in Strasbourg
#174I imagine Mr Good Guy at OVH telling some others: "guys we have a single point of failure in our architecture with SBG, maybe we should... - naaah it's fine, we do not have time nor resources" Then shit happens. edit: I have no idea what is happening exactly, but OVH being what it is, it seems extremely weird that all datacenters "can" get down at the same time, and it looks like a serious architecture problem to me…
It is rather counter-intuitive. In neteng and dcops, "I don't know and I cannot find out. I can only attempt to mitigate what i think might have caused it for next time" is a very reasonable answer to 99% of the "why this happened?" questions because in order to replicate the situation to test the theory one needs to recreate the same problem again on the same scale.
This also means that certain things cannot be tested. Most of generator tests are garbage - turning on generator and running it without production load delivered over the transfer switch does not test anything other than that one can turn on a generator and run it. The problem typically happens not because the generator ( also is there the generator or the first and the second generator? Why is there no generator bank for a non monkey-sized company? ) does not start - the problem is because over time transfer switch develops a problem and unlike generators it is not possible to test a transfer switch where in the event of a test failure the customers won't lose power unless the data center is designed from the beginning to deliver A and B powers over separate circuits to every single customer and every single customer has per system ( not per rack ) transfer switches.
Of course it costs a lot more money, something that companies are reluctant to spend.
Re: OVH Incident in Strasbourg
#175wow, yesterday I was playing with their public cloud because considering choosing them. I had some connection problem with my private networking there (deleted it more than once) and opened a ticket. If it was me... sorry, haha. Not good advertisement but it can happen to everyone.
Hey I would suggest not completely disregarding OVH. Their prices are good and their network is usually very reliable and fast. You simply won't find a North American hosting provider with those prices and reliability. You won't find one at double OVH prices, either (and I have tried).
Re: OVH Incident in Strasbourg
#176Earlier quoted context omitted.
just in case you followed the wrong devops borat, it's @DEVOPS_BORAT (the other one is a spam bot).
Azamat suggest to follow for your Kazakh tech need @DNS_BORAT https://twitter.com/DNS_BORAT @InfoSecBorat https://twitter.com/InfoSecBorat @KanbanBorat https://twitter.com/KanbanBorat @mysqlborat https://twitter.com/mysqlborat @NetEng_Borat https://twitter.com/NetEng_Borat @secure_borat https://twitter.com/secure_borat @SecurityBorat https://twitter.com/SecurityBorat @Sysadm_Borat https://twitter.com/Sysadm_Borat
Re: OVH Incident in Strasbourg
#177Earlier quoted context omitted.
While at $bigco we halted testing of generation equipment because it was sending DCs offline more often than it kept them up. Lawyers were involved, things got ugly
I'm completely unfamiliar with electrical generators/power generation, so take this question in the spirit of ignorance: Is there not a way to test generators without actually having them power the live datacenter infrastructure? I mean, simulate the exact generation and load requirements that the generators will face? I don't know if it's feasible to dump all that power to ground or whatever, but that way you could…
That would exercise transfer switches in addition to generators. Transfer switches are always energized, except for the time an equivalent of a big red mechanical switch is flipped into "OFF" position. When it is in the off position, neither main or generator are going to provide power to the customer. The biggest power consumer is actually a cooling system. If a HVAC system stops functioning in a typical data center, the temperature would quickly rise to the level of Arizona desert, destroying metric tons of equipment. Transfer switches have certain properties where after a flip they may not go back to the correct state on power restore ( say 1% ). Normally it is not a big deal because you really need to have lose power under full load often to get bit by it and if you are losing power that much you should probably address it with the utility company. However, if you are doing your full test once a week, in one year you would introduce 52 power failures.
Say your transfer switch is stuck in a wrong position. Now you need to shutdown all heat generating equipment to when you drop power to fix/replace the failed component you don't melt your data center... Congratulations, your data center now has to go offline.
Re: OVH Incident in Strasbourg
#178Earlier quoted context omitted.
Azamat suggest to follow for your Kazakh tech need @DNS_BORAT https://twitter.com/DNS_BORAT @InfoSecBorat https://twitter.com/InfoSecBorat @KanbanBorat https://twitter.com/KanbanBorat @mysqlborat https://twitter.com/mysqlborat @NetEng_Borat https://twitter.com/NetEng_Borat @secure_borat https://twitter.com/secure_borat @SecurityBorat https://twitter.com/SecurityBorat @Sysadm_Borat https://twitter.com/Sysadm_Borat
Anyone here remember https://en.m.wikipedia.org/wiki/Bastard_Operator_From_Hell
Re: OVH Incident in Strasbourg
#179Apparently, the root cause of that issue is a critical software bug in Cisco NCS 2000 transponders.
Re: OVH Incident in Strasbourg
#180Not "all datacenters". Only 2 of them. They have 22, not counting all the POPs.
9 by their count (7xRBX+2xSBG). When this was posted, the CEO wrote that everything except BHS (Canadian DC) was offline (tweet now deleted). Presumably he was in RBX and noticed that they couldn't reach any of the other DCs. So it looked like all DCs were down for a while.