Live data from Hacker News

EC2 Maintenance Update II

aws.amazon.com

41–50 of 60 posts

Re: EC2 Maintenance Update II

#41
post #40

Earlier quoted context omitted.

> Their philosophy, as is pervasive in the hosting industry, is that they have paying customers so whatever they are doing must be right. Obviously Amazon often also seems to act this way, but this particular maintenance was handled well afaict, and availability zones showed their value. This could explain why Rackspace was shopped around by Morgan Stanley. They may be profitable now, but Amazon and Google are going…

Shopped around, with no takers. At last count it was reported they've given up on that.

Indeed they have. Now they're doubling down on their existing platforms (and also, OnMetal).

I think their best move would be to pivot to be a firm that manages solutions for corporations that refuse to move off on-premises equipment for whatever reason. Their CapEx costs fall away, and they already have a deep ocean of talent to draw on.

There are already large orgs that already do this, but Rackspace has the potential to suck A LOT less than they do at the same task.

Re: EC2 Maintenance Update II

#42

Earlier quoted context omitted.

Don't lie. I had ~120 instances I had to juggle between 3 availability zones, and never were two AZs down/rebooted at once. Our environment suffered no downtime, as we had at least 2-3 days notice per AZ to move instances around. Rackspace's handling of the situation was a joke. They sent notification emails out at 9:30pm on a Friday night, and then proceeded to do the reboots Saturday at peak traffic times.

Yeah Rackspace is hardly to be held up as a standard bearer here, they don't even _have_ availability zones. We have around 200 instances, had about 59 reboot, and specifically were able to plan around these happening on different days. We weren't super excited when the window seemed to go to 4h right before it started, but we were prepared. I'm an ex Racker and I've told people high up at Rackspace for years that un…

something like availability zones

How is that different from a datacenter?

Re: EC2 Maintenance Update II

#44
post #42

Earlier quoted context omitted.

Yeah Rackspace is hardly to be held up as a standard bearer here, they don't even _have_ availability zones. We have around 200 instances, had about 59 reboot, and specifically were able to plan around these happening on different days. We weren't super excited when the window seemed to go to 4h right before it started, but we were prepared. I'm an ex Racker and I've told people high up at Rackspace for years that un…

something like availability zones How is that different from a datacenter?

With az's you can at least have some fault tolerance and not have to pay the outbound costs or latency associated with going across the wan.

Re: EC2 Maintenance Update II

#45
post #21

Earlier quoted context omitted.

Amazon had to perform the maintenance due to XSA 108, and the timetable they had to meet was set by the Xen project. They set up the maintenance to have as little impact as possible by splitting availability zones onto separate days so that people utilizing multiple zones for high-availability would not lose multiple zones at once. Afterwards, they gave a detailed breakdown, linking to the vulnerability and explainin…

Doesn't matter why they did the maintenance, so your first four sentences are simply rationalizations. The parent was complaining all the zones were restarted AT THE SAME TIME. Jeff explains why they had to do the maintenance, and nobody is disputing why they had to do it. What is being disputed is that they did it all at once, without rolling the restarts. This makes ZERO sense given their advice to 're-examine your…

I was pretty unhappy with how Rackspace handled the maintenance window.

1. Rackspace's maintenance announcement was sent at 9:00 PM on Friday night (Pacific time). Seriously?! I had already left for a weekend vacation without my laptop, so I couldn't do anything to get my company prepared. Even if the patch wasn't ready until Friday night, Rackspace could have scheduled the maintenance windows and announced them to customers much earlier.

2. The maintenance window for all three USA regions were scheduled at the same time. We couldn't just move to a different region without going to another continent.

3. Each maintenance window was 24 hours -- that's just too long. Even though our servers were only down for 10 minutes, we had to be on call and ready for 24 hours.

4. Although we have redundant servers in every region, we still couldn't guarantee that those redundant servers wouldn't be rebooted at the same time. As it turns out, we did lose both of our servers in ORD at the same time.

Re: EC2 Maintenance Update II

#46

@jeffbarr I just wanted to say thanks for not only posting this, but also sticking around in the comments section. Makes you + Amazon seem way more human :-) Also, in case you have any "cloud servers" you want to decommission: https://news.ycombinator.com/item?id=8373394

I am always happy to help, time and circumstances allowing. Before joining Amazon I earned my living by consulting for startups. I could always tell when they were about to run out of money when they would offer to pay me in servers. This was always the cue to find my next gig.

I would like to request, that if you're going to suggest: "Run instances in two or more Availability Zones." as a fault-tolerant architecture method (which everyone SHOULD do) - can we please have a flag or tag to say that "Traffic between Zone-1-HOST-A and Zone-2-HOST-A" is specifically resiliancy/HA/fault-tolerant connection traffic could benefit from a discount in zone transfer fees?

We are moving many hundreds servers between zones every day and sometimes by the hour, specifically to handle all sorts of constraints (Spot, capacity, failure, load, etc)...

I have paid quite a lot to zone transfer fees, especially when dealing with AWS network issues or spot price insanity.

I would love to classify some traffic as resiliency traffic and be charged differently for that traffic as opposed to general traffic used to service my user-base.

Re: EC2 Maintenance Update II

#47
post #8

"Pay attention to your Inbox and to the alerts on the AWS Management Console." Especially with incidents like these (and other cases where instances are scheduled to be taken down), it really annoys me that AWS doesn't offer any push alerts besides emailing the account owner.

What kinds of alerts would work for you? Let me know and I will pass them along to the team.

I got more than 2,000 pager duty alerts :) Would you like to be on-call?

Re: EC2 Maintenance Update II

#48

Earlier quoted context omitted.

Yeah Rackspace is hardly to be held up as a standard bearer here, they don't even _have_ availability zones. We have around 200 instances, had about 59 reboot, and specifically were able to plan around these happening on different days. We weren't super excited when the window seemed to go to 4h right before it started, but we were prepared. I'm an ex Racker and I've told people high up at Rackspace for years that un…

> Their philosophy, as is pervasive in the hosting industry, is that they have paying customers so whatever they are doing must be right. Obviously Amazon often also seems to act this way, but this particular maintenance was handled well afaict, and availability zones showed their value. This could explain why Rackspace was shopped around by Morgan Stanley. They may be profitable now, but Amazon and Google are going…

Amazon and Google are worthless to any company without a robust sysadmin/devops team, which is most companies in the world.

It's easy to get stuck in the tech-savvy bubble here, where most people can write code, pick up Chef in a week, and are trying to build cheaply at "web scale." Those people don't need, or want to pay for, support with their servers.

But most companies need some help to run a few servers for web and email. Rackspace is the only large hosting provider who provides that across the board.

That said, Rackspace needs to beef up their devops support, or they risk limiting their own abilities to grow with their customers.

Re: EC2 Maintenance Update II

#49
post #42

Earlier quoted context omitted.

Yeah Rackspace is hardly to be held up as a standard bearer here, they don't even _have_ availability zones. We have around 200 instances, had about 59 reboot, and specifically were able to plan around these happening on different days. We weren't super excited when the window seemed to go to 4h right before it started, but we were prepared. I'm an ex Racker and I've told people high up at Rackspace for years that un…

something like availability zones How is that different from a datacenter?

AZs in AWS are essentially distinct "datacenters" from a logical perspective. There is no shared infrastructure between them; if AZ B drops off, as long as you have your data and instances replicated and serving in another AZ, you will see no downtime.

Re: EC2 Maintenance Update II

#50

Earlier quoted context omitted.

> Their philosophy, as is pervasive in the hosting industry, is that they have paying customers so whatever they are doing must be right. Obviously Amazon often also seems to act this way, but this particular maintenance was handled well afaict, and availability zones showed their value. This could explain why Rackspace was shopped around by Morgan Stanley. They may be profitable now, but Amazon and Google are going…

Amazon and Google are worthless to any company without a robust sysadmin/devops team, which is most companies in the world. It's easy to get stuck in the tech-savvy bubble here, where most people can write code, pick up Chef in a week, and are trying to build cheaply at "web scale." Those people don't need, or want to pay for, support with their servers. But most companies need some help to run a few servers for web…

Huh?

Elastic Beanstalk, OpsWorks, Google sites, Google apps, AWS Marketplace, etc..

Going by what you said about most companies, most registered businesses in the world are likely just looking for a single dinky site with a mailbox pointing @theirbusiness.com, definitely no need for more than a shared server. Google, Wordpress, Github pages, Shopify, and dozens of others make this very simple to setup and use. You said Rackspace is the only large hosting provider that provides this across the board, that is not true and they aren't even in my top 10 if I was looking for a provider.

For a single website that gets less than 10 visits / day with 5 html pages I was just quoted $75/mo minimum by Rackspace with some server management on my part.

Post reply on HN