Microsoft had three staff at Australian data centre campus when Azure went out
11–20 of 52 posts
Re: Microsoft had three staff at Australian data centre campus when Azure went out
#12Guessing affected customers had to spend time and effort on top of ongoing high cloud bills I've slept so much better since I began hosting, producing energy, and cooling on-prem
What about data center colocation? When you simply rent the energy, cooling, etc, but the hardware is yours? Do you think it's a nice middle ground?
It is.
The cloud fanbois will tell you until their blue in the face that its not.
I fully accept that the cloud is great for bursty workloads where you're doing nothing and then suddenly half the planet needs your service for a couple of days. That is clear.
But if you've got a reasonably stable baseload running 24x7x365 and a few modest bursts here and there then honestly people need to do the math, because if you look at beyond the short-term figures, the cloud tends to work out much more expensive than colo if you look at for example a three-year period.
Most people don't need the scale the cloud gives. They think they do, but really most people will never grow to FANG scale as much as they may dream it !
Re: Microsoft had three staff at Australian data centre campus when Azure went out
#13Earlier quoted context omitted.
The secret is hosting across failure boundaries so that a single outage like this does not impact you. Self-hosting is fine if you can afford the capex for two physically separate data centers (like really separate - like 100+ miles etc (or more!) to cope with natural disasters) and the staff to operate & maintain them 24/7. For many, this is not realistic. For those that do need to use cloud, just make sure you are…
>like 100+ miles etc (or more!) to cope with natural disasters) People talk about this often but this failure mode seems to never happen? When was the last time us-east-1 went down because of a natural calamity compared to some technical issue?
.... Of course then you have latency issues to think about, but that is often quite application-specific and potentially a good problem to have if a slightly slow website or database or whatever is the biggest problem you have when the alternative would have been a total shutdown.
There are also occasional fires and stuff that take out a whole building (I think OVH had this in France recently?). Ensure that your failure zones are physically separate places, and not just logically-separate zones in the same physical building, or in a building that is next to the one on fire :)
Re: Microsoft had three staff at Australian data centre campus when Azure went out
#14Guessing affected customers had to spend time and effort on top of ongoing high cloud bills I've slept so much better since I began hosting, producing energy, and cooling on-prem
The secret is hosting across failure boundaries so that a single outage like this does not impact you. Self-hosting is fine if you can afford the capex for two physically separate data centers (like really separate - like 100+ miles etc (or more!) to cope with natural disasters) and the staff to operate & maintain them 24/7. For many, this is not realistic. For those that do need to use cloud, just make sure you are…
I would be very, very surprised if the companies mentioned, in particular banks, weren't running on multiple AZs, but I wouldn't be surprised if the scenario of severing down an AZ was not tested.
Re: Microsoft had three staff at Australian data centre campus when Azure went out
#15I know very little on the datacenter operations side of things - I guess 3 people is not a lot, but what is normal? How many operations people are at say AWS US-East-1? I presume it doesn't scale with number of servers, that would not scale well. What is a 'normal' level? 10? 100? It can't be more than 100, can it?
Bear in mind that outside of the US and maybe one or two other locations ex-US, almost all of the magic cloud operates out of third-party datacentres, not their own.
They will have a small office on-site where 3–5 people sit, and those people are exclusively dedicated to the cloud equipment itself. The datacentre ops side is, by definition, handled by the third-party datacentre operator.
The guys onsite are clearly only there for "intelligent hands" purposes, as everything else will be done remotely from silicon valley or wherever.
Re: Microsoft had three staff at Australian data centre campus when Azure went out
#16I know very little on the datacenter operations side of things - I guess 3 people is not a lot, but what is normal? How many operations people are at say AWS US-East-1? I presume it doesn't scale with number of servers, that would not scale well. What is a 'normal' level? 10? 100? It can't be more than 100, can it?
Re: Microsoft had three staff at Australian data centre campus when Azure went out
#17Earlier quoted context omitted.
What about data center colocation? When you simply rent the energy, cooling, etc, but the hardware is yours? Do you think it's a nice middle ground?
> Do you think it's a nice middle ground? It is. The cloud fanbois will tell you until their blue in the face that its not. I fully accept that the cloud is great for bursty workloads where you're doing nothing and then suddenly half the planet needs your service for a couple of days. That is clear. But if you've got a reasonably stable baseload running 24x7x365 and a few modest bursts here and there then honestly pe…
Also on the price side, I'm not comparing the price of cloud vs colo, but the price of cloud vs what the company IT department charges my department for being allowed to use one of 'their' colo servers, and that is many times what a cloud server costs. (as a real world example, the place I used to work internally invoiced $150/server/month for a virtual server that would cost me $20/server/month on AWS before any discounts).
Cloud lives not by competing against smart people running their own servers, but against inefficient internal IT services, and there they have them beat both on price and quality.
Re: Microsoft had three staff at Australian data centre campus when Azure went out
#18I know very little on the datacenter operations side of things - I guess 3 people is not a lot, but what is normal? How many operations people are at say AWS US-East-1? I presume it doesn't scale with number of servers, that would not scale well. What is a 'normal' level? 10? 100? It can't be more than 100, can it?
0 - 2 staff in a typical DC is not unusual at all, with people who are on-call usually within a 30 minute drive. Larger DCs can and do have more staff on-site 24/7 and typically the amount of staff on-site at any given time is driven by SLAs. I expect the DC in TFA to return to lower staff levels once they've worked on reducing their total "time to restart chiller" or reduced the amount of manual work involved in doi…
Re: Microsoft had three staff at Australian data centre campus when Azure went out
#19The only people who should be shocked in this thread are the people who have been hoodwinked into thinking operations is so hard you need thousands of staff. I know AWS/GCP/Azure like to charge us as if we were hiring an army of sysadmins, but the truth is that day-to-day DC ops does not require so many people. Hardware failures are more rare than you think and you can work around them without panicking anyway.
Re: Microsoft had three staff at Australian data centre campus when Azure went out
#20Earlier quoted context omitted.
0 - 2 staff in a typical DC is not unusual at all, with people who are on-call usually within a 30 minute drive. Larger DCs can and do have more staff on-site 24/7 and typically the amount of staff on-site at any given time is driven by SLAs. I expect the DC in TFA to return to lower staff levels once they've worked on reducing their total "time to restart chiller" or reduced the amount of manual work involved in doi…
Still we read how DCs are a job generator. E.g. "One hopes for hundreds of jobs for locals." https://cryptoquorum.com/oman-opens-cryptocurrency-mining-ce...