So they rolled out a performance update (not a critical security fix) to all their datacenters at once? This sounds incredibly amateur for a provider the size of Azure.
Hey nnx, this is Corey from the Azure engineering team. We have a standard protocol in the team of applying production changes in incremental batches. Due to an operational error, this update was made across most regions in a short period of time. I really apologize for the disruption.
Update on Azure Storage Service Interruption
11–18 of 18 posts
Re: Update on Azure Storage Service Interruption
#12Earlier quoted context omitted.
Hey nnx, this is Corey from the Azure engineering team. We have a standard protocol in the team of applying production changes in incremental batches. Due to an operational error, this update was made across most regions in a short period of time. I really apologize for the disruption.
So, you had a bug in your code. That happens to everyone and I think we all understand. However, there are a number of other issues here which seem systemic and much more troubling. First, that your "flighting" did not catch the problem. Why was that? If the bug caused an infinite loop on all the live storage systems, that seems like it should have been fairly obvious on the customer systems you tested on. Second, th…
Re: Update on Azure Storage Service Interruption
#13I hope to read more in the post-mortem RCA but I am curious what their flighting missed, is flighting so limited that is does not see the cross region scale or something? I also had the feeling from watching Mark Russinovich discuss previous failures that their patch rollouts were much more controlled.
Re: Update on Azure Storage Service Interruption
#14> The configuration change for the Blob Front-Ends exposed a
> bug in the Blob Front-Ends, which had been previously
> performing as expected for the Table Front-Ends.
> This bug resulted in the Blob Front-Ends to go into an
> infinite loop not allowing it to take traffic.
An infinite loop that has not been discovered during the partial role out. That is clearly weird. Also, doesn't put too much trust in their current monitoring scheme.
Re: Update on Azure Storage Service Interruption
#15Earlier quoted context omitted.
Hey nnx, this is Corey from the Azure engineering team. We have a standard protocol in the team of applying production changes in incremental batches. Due to an operational error, this update was made across most regions in a short period of time. I really apologize for the disruption.
So, you had a bug in your code. That happens to everyone and I think we all understand. However, there are a number of other issues here which seem systemic and much more troubling. First, that your "flighting" did not catch the problem. Why was that? If the bug caused an infinite loop on all the live storage systems, that seems like it should have been fairly obvious on the customer systems you tested on. Second, th…
Re: Update on Azure Storage Service Interruption
#16Earlier quoted context omitted.
So, you had a bug in your code. That happens to everyone and I think we all understand. However, there are a number of other issues here which seem systemic and much more troubling. First, that your "flighting" did not catch the problem. Why was that? If the bug caused an infinite loop on all the live storage systems, that seems like it should have been fairly obvious on the customer systems you tested on. Second, th…
Thanks. We are continuing to investigate this and driving needed improvements in our process and technology to avoid similar issues in the future.
The OP mentions that Microsoft representatives gave info via public forums. When the issue appeared I looked in different places trying to find info, but only I found was a statement saying that We are aware of issues. I looked at Azure twitter/blog, ScottGu twitter/blog, Hanselmans, MSDN forums. I also tried this forum and reddit. Do you know where I should have gone to receive details?
Re: Update on Azure Storage Service Interruption
#17Earlier quoted context omitted.
Thanks. We are continuing to investigate this and driving needed improvements in our process and technology to avoid similar issues in the future.
The last two times there was a big issue the same thing happened with the status dashboard (it became inaccessible). I remember the same issue when the certs expired 1,5 years ago. I really like Microsoft and was convinced "you" would somehow isolate the dashboard and host it separately, but it turns out I was wrong. Do you happen to know the reasons for hosting the status dashboard inside of Azure? It seems so count…
For general communications, we did most of our early communication on the event using twitter, announcing the incident and giving updates. We need to build a more formal multi-pronged approach to communicating, including faster responses in the MSDN forums and here in HN to make sure we are reaching as many of our customers and partners as possible. Thanks again for the feedback!!
Re: Update on Azure Storage Service Interruption
#18Please send with high importance so it pops in our inbox and we will dig in.