Live data from Hacker News

Summary of June 8 outage

fastly.com

61–70 of 110 posts

Re: Summary of June 8 outage

#61

This is annoyingly vague. What was the software bug, and what was the valid customer configuration change? It's perhaps a bit premature to demand it at this point, but I'm hoping a full post-mortem will outline precisely how this change was not picked up in pre-prod. Surely all valid customer configurations must be tested prior to rollout.

If they tell everyone exactly how to trigger the bug before they finish rolling out a fix, people will trigger it on purpose to bring down websites.

Re: Summary of June 8 outage

#62
post #7

Absolutely no details about the bug or why a single customer configuration effected global state on the server, or why this wasn't caught by configuration change safety mechanisms/smoke tests/gradual rollout. Also, what is up with their partitioning? Do they seriously have one customer that gets served from 85% of their servers? Is it a whale? Good on them for getting a statement out right away (although they basical…

Smoke tests or gradual rollout won’t help with the change by fastly on May 12 when they deployed the buggy change. It’s likely that this was gradually rolled out and looked ok with all customer configs existing on May 12. Obviously there should be other safeguards in place but gradual rollout in itself wouldn’t have helped since it would look green at a 100% rollout weeks ago.

Re: Summary of June 8 outage

#63
post #56

Earlier quoted context omitted.

This attitude is why we have only 2½ search engines on the entire Internet. Only Google, Bing, and Yandex run crawlers. Everybody else is just a reseller for them. Web crawlers are a feature not a bug. If your site shouldn't be crawled, it doesn't belong on the Internet.

Search engines scrapping your content is not the problem. Competitors scraping your content is.

If you don't want your content crawled, don't put it on the public Internet.

Your profitability is not the Internet's problem.

Re: Summary of June 8 outage

#64
post #7

Absolutely no details about the bug or why a single customer configuration effected global state on the server, or why this wasn't caught by configuration change safety mechanisms/smoke tests/gradual rollout. Also, what is up with their partitioning? Do they seriously have one customer that gets served from 85% of their servers? Is it a whale? Good on them for getting a statement out right away (although they basical…

Smoke tests or gradual rollout won’t help with the change by fastly on May 12 when they deployed the buggy change. It’s likely that this was gradually rolled out and looked ok with all customer configs existing on May 12. Obviously there should be other safeguards in place but gradual rollout in itself wouldn’t have helped since it would look green at a 100% rollout weeks ago.

I meant smoke testing and gradual rollout of the config, not the code, although obviously that's important too.

Re: Summary of June 8 outage

#65
post #38

This is annoyingly vague. What was the software bug, and what was the valid customer configuration change? It's perhaps a bit premature to demand it at this point, but I'm hoping a full post-mortem will outline precisely how this change was not picked up in pre-prod. Surely all valid customer configurations must be tested prior to rollout.

It's not just annoying value. It's insultingly vague. If my data centre provider suffered a complete outage, then I demand to get a detailed post-mortem of what happened (in due time). If they just tell me bullshit PR speak about "We value our customers", I'll be looking at switching providers. As a Fastly customer whose site went down, I'm entitled to know exactly what happened. If they don't tell me, I'm switching…

It does say they haven't finished rolling out the permanent fix, (i.e. such a customer configuration(/exploit) could still bring down some servers) and will/are conducting a full post-mortem. So hopefully a juicier post to come.

Re: Summary of June 8 outage

#66
post #36

Earlier quoted context omitted.

They forgot to test it.

Test “it”? The change in question wasn’t by fastly but a customer of theirs making a config change. It’s possible that this customer did validate their change somehow. Fastly obviously didn’t test their code (with the bug) enough, but testing of course can never prove the absence of bugs. Testing for a global deployment like a massive CDN happens to a large extent in prod because you don’t have another globe. You can…

Fastly even say it was a valid change.

> We experienced a global outage due to an undiscovered software bug that surfaced on June 8 when it was triggered by a valid customer configuration change.

in the first sentence

Re: Summary of June 8 outage

#67
post #7

Absolutely no details about the bug or why a single customer configuration effected global state on the server, or why this wasn't caught by configuration change safety mechanisms/smoke tests/gradual rollout. Also, what is up with their partitioning? Do they seriously have one customer that gets served from 85% of their servers? Is it a whale? Good on them for getting a statement out right away (although they basical…

> one customer that gets served from 85% of their servers

There could be a feedback loop that is the opposite of a smoke test.

1. Validate customer configuration, if passes, assume it can roll out 2. Roll out customer configuration to node 3. Node goes down 4. Migrate all customers on node to new nodes 5. Node that problematic customer was migrated to goes down 6. Rinse and repeat as problematic customer migrates to every node and takes out every last one.

Re: Summary of June 8 outage

#68

Thought it surprising that amazon.com uses fastly when they have their own cdn

Someone answered this yesterday. CloudFront is good for video and large download assets (plus very low margins) but not for images and smaller stuff which Fastly is much faster at:

https://www.streamingmediablog.com/2020/05/fastly-amazon-hom...

Re: Summary of June 8 outage

#69
post #37

You took down my site and a good swath of the whole internet. I am entitled to know, in detail, what happened, so I can be more informed and assess any actions I might need to take. I don't want any more of your PR speak or "we value our customers". That's crap and insults my intelligence. STOP getting PR to write your comms; just speak to engineers like engineers. I'd rather get no response than this post. I hope th…

But you're really not entitled to the details, no matter how you feel about it. You're entitled to whatever compensation is defined in your contract if an SLA was breached. Beyond that, Fastly is going to provide the level of detail they feel is necessary to reassure their major customers that they've addressed the issue and will do their best to keep it from happening again. Unfortunately one of the downsides of using the services of other companies is knowing that something is going to happen at some point, and there's not going to be anything you can do about it except hope that your contingency plans are adequate or wait it out.

Obviously on this site we tend to be rather technical people, so we want to know as much detail as possible, but that's something we desire, not something we are entitled to.

Re: Summary of June 8 outage

#70
post #49

Yeah - not really a post mortem, is it? "We had a bug and we fixed the bug." They could probably cut and paste that same page for 90% of future outages. Maybe they need to read this: https://artsy.github.io/blog/2014/11/19/how-to-write-great-o...

The post says "We are conducting a complete post mortem of the processes and practices we followed during this incident. "

So I don't think they are claiming this is a post mortem.

Post reply on HN