Live data from Hacker News

EC2 Maintenance Update II

aws.amazon.com

31–40 of 60 posts

Re: EC2 Maintenance Update II

#31
post #8

Earlier quoted context omitted.

What kinds of alerts would work for you? Let me know and I will pass them along to the team.

I would like to be able to have alerts go right into SNS.

Oh the possibilities! This would be super handy.

I'd add SNS support for non-US SMS notifications :p Then I guess I would not need Pingdom anymore.

Re: EC2 Maintenance Update II

#32
post #21

Earlier quoted context omitted.

> you may also want to take this opportunity to re-examine your AWS architecture to look for possible ways to make it even more fault-tolerant Basically they won't allow blame to be placed on them for anything they do. According to AWS, it's your fault this happened.

Amazon had to perform the maintenance due to XSA 108, and the timetable they had to meet was set by the Xen project. They set up the maintenance to have as little impact as possible by splitting availability zones onto separate days so that people utilizing multiple zones for high-availability would not lose multiple zones at once. Afterwards, they gave a detailed breakdown, linking to the vulnerability and explainin…

Doesn't matter why they did the maintenance, so your first four sentences are simply rationalizations. The parent was complaining all the zones were restarted AT THE SAME TIME. Jeff explains why they had to do the maintenance, and nobody is disputing why they had to do it. What is being disputed is that they did it all at once, without rolling the restarts. This makes ZERO sense given their advice to 're-examine your (AWS based) architecture' for fault tolerance. Don't patronize me while you're doing something that affects me in a way I can't control.

I've had direct experience with AWS in this regard, and was equally disappointed in the outcome. If you want to see a class act in taking ownership of issues arising from such matters, have a look at Rackspace's response:

  Dear Rackspace Customer,

  I’m writing to apologize for the downtime and inconvenience that you and others of our customers have suffered in recent days.  Like other major cloud providers, we were forced to reboot some of our customers’ servers to patch a security vulnerability affecting certain versions of XenServer, a popular open-source hypervisor.  This maintenance was especially difficult for many of you because it had to be performed on short notice, and over the weekend.

  Now that this issue has been fully remediated, without any reports of compromised data among our customers, I’d like to explain what happened, and why.

  Whenever we at Rackspace become aware of a security vulnerability, whether in our systems or (as in this case) in third-party software, we face a balancing act.  We want to be as transparent as possible with you, our customers, so you can join us in taking actions to secure your data.  But we don’t want to advertise the vulnerability before it’s fixed — lest we, in effect, ring a dinner bell for the world’s cyber criminals.

  That’s the dilemma that we faced over the Xen bug. Such vulnerabilities are regularly found in software, whether proprietary or open source. The key, once a bug is identified, is to fix it swiftly and quietly.  This particular vulnerability could have allowed bad actors who followed a certain series of memory commands to read snippets of data belonging to other customers, or to crash the host server. We wanted to flag the issue as quickly as possible to those of you using our Standard, Performance 1, and Performance 2 Cloud Servers, and our Hadoop Cloud Big Data service.  But we didn’t want to do so until we had a software patch in place to address the vulnerability.

  When we learned of the security issue and realized its significance early last week, our engineers worked with our Xen partners to develop and test a patch, and organize a reboot plan.  The patch wasn’t ready until the evening of Friday, Sept. 26.  And the technical details of the vulnerability were scheduled to be publicly released on Wednesday, Oct. 1.  We were faced with the difficult decision of whether to start our reboots over the weekend, with short notice to our customers, or postpone it until Monday. The latter course would not allow us to sufficiently stagger the reboots.  It would jeopardize our ability to fully patch all the affected servers before the vulnerability became public, thus exposing our customers to heightened risk.

  We decided the lesser evil was to proceed immediately, at which time we notified you, and our partners in the Xen community, of the need for an urgent server reboot.  Even then, to avoid alerting cyber criminals, we didn’t mention Xen as the reason for the reboot. Another major cloud provider did attribute its reboot to security problems with Xen, which put all users of the affected versions of that hypervisor at heightened risk.  But we’re relieved to report that, as of now, we’ve learned of no data compromise among Rackspace customers.  Now that the vulnerability has been fully remediated, the Xen community has lifted its embargo on talking about it.

  Those of you who are longtime Rackspace customers know that we have a strong record of open, timely communication with you. We reach out to you whenever there’s an issue.  We answer the phone whenever you call. We do everything we can to find a solution.  This past weekend, our engineers worked tirelessly with customers and partners to remediate the Xen vulnerability.  

  This maintenance affected nearly a quarter of our 200,000-plus customers, and in the course of it, we dropped a few balls.  Some of our reboots, for example, took much longer than they should.  And some of our notifications were not as clear as they should have been. We are making changes to address those mistakes.  And we welcome your feedback on how we can better serve you.

  As a veteran Racker who is proud of our commitment to our customers and their businesses, I am personally sorry for any inconvenience or downtime that we caused you during this incident.

  Sincerely, 

  Taylor Rhodes
  CEO and President
  Rackspace

  taylor.rhodes@rackspace.com

Re: EC2 Maintenance Update II

#33
One of my instances was scheduled to be rebooted between 22pm-2am.

But the reboot has never actually occurred.

It's not that I feel left out but did anyone else experience the same?

Re: EC2 Maintenance Update II

#35
post #14

Earlier quoted context omitted.

More generally, it would be cool if every account automatically got two free read-only SQS queues: one for events and one for upcoming events. Every known event (startup, terminate, permission change, network error, etc) could be published to the queue. Not high-priority for us but the parent comment sparked the idea (I'm not in devops so maybe this exists via a different mechanism.)

These are both great ideas. I will share them with the team today. Keep them coming!

seems like realtime CloudTrail->{SQS,SNS} could satisfy his request and several that I have.

(instead of waiting for events to batch from CloudTrail to S3)

Re: EC2 Maintenance Update II

#36
post #21

Earlier quoted context omitted.

Amazon had to perform the maintenance due to XSA 108, and the timetable they had to meet was set by the Xen project. They set up the maintenance to have as little impact as possible by splitting availability zones onto separate days so that people utilizing multiple zones for high-availability would not lose multiple zones at once. Afterwards, they gave a detailed breakdown, linking to the vulnerability and explainin…

Doesn't matter why they did the maintenance, so your first four sentences are simply rationalizations. The parent was complaining all the zones were restarted AT THE SAME TIME. Jeff explains why they had to do the maintenance, and nobody is disputing why they had to do it. What is being disputed is that they did it all at once, without rolling the restarts. This makes ZERO sense given their advice to 're-examine your…

Don't lie. I had ~120 instances I had to juggle between 3 availability zones, and never were two AZs down/rebooted at once. Our environment suffered no downtime, as we had at least 2-3 days notice per AZ to move instances around.

Rackspace's handling of the situation was a joke. They sent notification emails out at 9:30pm on a Friday night, and then proceeded to do the reboots Saturday at peak traffic times.

Re: EC2 Maintenance Update II

#37

Earlier quoted context omitted.

Doesn't matter why they did the maintenance, so your first four sentences are simply rationalizations. The parent was complaining all the zones were restarted AT THE SAME TIME. Jeff explains why they had to do the maintenance, and nobody is disputing why they had to do it. What is being disputed is that they did it all at once, without rolling the restarts. This makes ZERO sense given their advice to 're-examine your…

Don't lie. I had ~120 instances I had to juggle between 3 availability zones, and never were two AZs down/rebooted at once. Our environment suffered no downtime, as we had at least 2-3 days notice per AZ to move instances around. Rackspace's handling of the situation was a joke. They sent notification emails out at 9:30pm on a Friday night, and then proceeded to do the reboots Saturday at peak traffic times.

Yeah Rackspace is hardly to be held up as a standard bearer here, they don't even _have_ availability zones.

We have around 200 instances, had about 59 reboot, and specifically were able to plan around these happening on different days.

We weren't super excited when the window seemed to go to 4h right before it started, but we were prepared.

I'm an ex Racker and I've told people high up at Rackspace for years that until they implement something like availability zones, they're a joke for any kind of production. Their philosophy, as is pervasive in the hosting industry, is that they have paying customers so whatever they are doing must be right. Obviously Amazon often also seems to act this way, but this particular maintenance was handled well afaict, and availability zones showed their value.

Re: EC2 Maintenance Update II

#38

Earlier quoted context omitted.

Don't lie. I had ~120 instances I had to juggle between 3 availability zones, and never were two AZs down/rebooted at once. Our environment suffered no downtime, as we had at least 2-3 days notice per AZ to move instances around. Rackspace's handling of the situation was a joke. They sent notification emails out at 9:30pm on a Friday night, and then proceeded to do the reboots Saturday at peak traffic times.

Yeah Rackspace is hardly to be held up as a standard bearer here, they don't even _have_ availability zones. We have around 200 instances, had about 59 reboot, and specifically were able to plan around these happening on different days. We weren't super excited when the window seemed to go to 4h right before it started, but we were prepared. I'm an ex Racker and I've told people high up at Rackspace for years that un…

> Their philosophy, as is pervasive in the hosting industry, is that they have paying customers so whatever they are doing must be right. Obviously Amazon often also seems to act this way, but this particular maintenance was handled well afaict, and availability zones showed their value.

This could explain why Rackspace was shopped around by Morgan Stanley. They may be profitable now, but Amazon and Google are going to eat their lunch.

Re: EC2 Maintenance Update II

#39
We had 50 servers reboot and 15... never came back. It was painful. 10 were in one AZ. We're putting servers in more AZs now, but preparing for a reboot and preparing to rebuild a redis slave, 2 zookeepers, 10 cassandras, and an API server are different things.

Re: EC2 Maintenance Update II

#40

Earlier quoted context omitted.

Yeah Rackspace is hardly to be held up as a standard bearer here, they don't even _have_ availability zones. We have around 200 instances, had about 59 reboot, and specifically were able to plan around these happening on different days. We weren't super excited when the window seemed to go to 4h right before it started, but we were prepared. I'm an ex Racker and I've told people high up at Rackspace for years that un…

> Their philosophy, as is pervasive in the hosting industry, is that they have paying customers so whatever they are doing must be right. Obviously Amazon often also seems to act this way, but this particular maintenance was handled well afaict, and availability zones showed their value. This could explain why Rackspace was shopped around by Morgan Stanley. They may be profitable now, but Amazon and Google are going…

Shopped around, with no takers. At last count it was reported they've given up on that.
Post reply on HN