Live data from Hacker News

Summary of June 8 outage

fastly.com

71–80 of 110 posts

Re: Summary of June 8 outage

#71
post #64

Earlier quoted context omitted.

Smoke tests or gradual rollout won’t help with the change by fastly on May 12 when they deployed the buggy change. It’s likely that this was gradually rolled out and looked ok with all customer configs existing on May 12. Obviously there should be other safeguards in place but gradual rollout in itself wouldn’t have helped since it would look green at a 100% rollout weeks ago.

I meant smoke testing and gradual rollout of the config, not the code, although obviously that's important too.

You’d think that individual customer configuration changes should only ever affect that customer, and that gradual vs instant rollout would be an option the customer handles when changing their configuration!

Re: Summary of June 8 outage

#72

Earlier quoted context omitted.

Test “it”? The change in question wasn’t by fastly but a customer of theirs making a config change. It’s possible that this customer did validate their change somehow. Fastly obviously didn’t test their code (with the bug) enough, but testing of course can never prove the absence of bugs. Testing for a global deployment like a massive CDN happens to a large extent in prod because you don’t have another globe. You can…

Fastly even say it was a valid change. > We experienced a global outage due to an undiscovered software bug that surfaced on June 8 when it was triggered by a valid customer configuration change. in the first sentence

Their change was bad, that was May 12. Since that seemed OK on May 13,14,… there wasn’t much indicating that change would blow up weeks later. For example if they roll it out gradually, they would reach 100% rollout with all lights being green

The customer change was a valid configuration. That was yesterday.

Re: Summary of June 8 outage

#73
This is a great write-up. When describing my own bugs and issues to stakeholders I often find the following difficult to communicate:

* Bug was introduced on date X but only caused problems on date Y ("if the bug was introduced on date X then we would have seen it on date X, so you’re wrong")

* Doing X led to the outage but X wasn’t the fault, X was a valid thing to do, the code should have been able to handle X, the fact the code couldn't handle it was the actual problem which needs to be fixed ("look, you said X caused the problem, so the solution is just not to do X right?")

This article conveys both these points clearly and effortlessly. I might borrow some terminology from this in the future.

Re: Summary of June 8 outage

#74
post #32

I love that somewhere out there is a developer who doesn't even work at Fastly but just innocently pushed a change to their Fastly config and basically broke the entire internet. I'm actually jealous. If it was me, I'd put that on my resume.

Isn't it a shame and a cause for alarm that the supposed decentralised, damage resistant Internet has been reduced to this?

Re: Summary of June 8 outage

#75

Earlier quoted context omitted.

What are Europe's whales?

Lots of tech companies have engineering offices with SRE responsabilities in Europe timezones. Could be anyone really.

My guess is this level of change would come from senior devs in the main office.

Re: Summary of June 8 outage

#76

Wasn't there a failover or some redundency from Fastly in place during this outage?

No doubt, but there's always a way to smoke the whole thing, even with all the fail safe and redundancy in the world.

Except when there is no connection (except for chaos theory) between certain systems (strict isolation).

Re: Summary of June 8 outage

#77

Earlier quoted context omitted.

A web filled with DDOS attacks and scraping is a web that needs cloudflare and fastly. I’m not sure how to avoid this sorry state of things.

Invisible Internet Project (I2P) is decentralized and defends from such attacks quite well.

How is that project doing? It's been around for years and does not come up often.

Re: Summary of June 8 outage

#78
I don't expect fastly to name and shame a customer who made a valid change, nor do I expect fastly to give us a detailed explanation of what the bug is.

I'm still a little annoyed at their status page [0]. It says:

> We're currently investigating potential impact to performance with our CDN services.

yet in the blog post we're talking about here it says:

> Early June 8, a customer pushed a valid configuration change that included the specific circumstances that triggered the bug, which caused 85% of our network to return errors.

85% of your network returning errors is _not_ a potential performance impact.

[0] https://status.fastly.com/incidents/vpk0ssybt3bj

Re: Summary of June 8 outage

#79

Earlier quoted context omitted.

Lots of tech companies have engineering offices with SRE responsabilities in Europe timezones. Could be anyone really.

My guess is this level of change would come from senior devs in the main office.

Are you implying that no senior devs or main offices are located in europe?

Also, just because the rollout to fastly happened at a morning EU time, doesn't mean that the change was made. If there's a deployment pipeline, it could have been made 2-3 hours earlier, or even the day before

Re: Summary of June 8 outage

#80
post #79

Earlier quoted context omitted.

My guess is this level of change would come from senior devs in the main office.

Are you implying that no senior devs or main offices are located in europe? Also, just because the rollout to fastly happened at a morning EU time, doesn't mean that the change was made. If there's a deployment pipeline, it could have been made 2-3 hours earlier, or even the day before

I'm saying a big network change is probably going to be executed while HQ is awake and in the loop. In my experience the most skillful devs aren't the most senior.
Post reply on HN