This is annoyingly vague. What was the software bug, and what was the valid customer configuration change? It's perhaps a bit premature to demand it at this point, but I'm hoping a full post-mortem will outline precisely how this change was not picked up in pre-prod. Surely all valid customer configurations must be tested prior to rollout.
Summary of June 8 outage
61–70 of 110 posts
Re: Summary of June 8 outage
#62Absolutely no details about the bug or why a single customer configuration effected global state on the server, or why this wasn't caught by configuration change safety mechanisms/smoke tests/gradual rollout. Also, what is up with their partitioning? Do they seriously have one customer that gets served from 85% of their servers? Is it a whale? Good on them for getting a statement out right away (although they basical…
Re: Summary of June 8 outage
#63Earlier quoted context omitted.
This attitude is why we have only 2½ search engines on the entire Internet. Only Google, Bing, and Yandex run crawlers. Everybody else is just a reseller for them. Web crawlers are a feature not a bug. If your site shouldn't be crawled, it doesn't belong on the Internet.
Search engines scrapping your content is not the problem. Competitors scraping your content is.
Your profitability is not the Internet's problem.
Re: Summary of June 8 outage
#64Absolutely no details about the bug or why a single customer configuration effected global state on the server, or why this wasn't caught by configuration change safety mechanisms/smoke tests/gradual rollout. Also, what is up with their partitioning? Do they seriously have one customer that gets served from 85% of their servers? Is it a whale? Good on them for getting a statement out right away (although they basical…
Smoke tests or gradual rollout won’t help with the change by fastly on May 12 when they deployed the buggy change. It’s likely that this was gradually rolled out and looked ok with all customer configs existing on May 12. Obviously there should be other safeguards in place but gradual rollout in itself wouldn’t have helped since it would look green at a 100% rollout weeks ago.
Re: Summary of June 8 outage
#65This is annoyingly vague. What was the software bug, and what was the valid customer configuration change? It's perhaps a bit premature to demand it at this point, but I'm hoping a full post-mortem will outline precisely how this change was not picked up in pre-prod. Surely all valid customer configurations must be tested prior to rollout.
It's not just annoying value. It's insultingly vague. If my data centre provider suffered a complete outage, then I demand to get a detailed post-mortem of what happened (in due time). If they just tell me bullshit PR speak about "We value our customers", I'll be looking at switching providers. As a Fastly customer whose site went down, I'm entitled to know exactly what happened. If they don't tell me, I'm switching…
Re: Summary of June 8 outage
#66Earlier quoted context omitted.
They forgot to test it.
Test “it”? The change in question wasn’t by fastly but a customer of theirs making a config change. It’s possible that this customer did validate their change somehow. Fastly obviously didn’t test their code (with the bug) enough, but testing of course can never prove the absence of bugs. Testing for a global deployment like a massive CDN happens to a large extent in prod because you don’t have another globe. You can…
> We experienced a global outage due to an undiscovered software bug that surfaced on June 8 when it was triggered by a valid customer configuration change.
in the first sentence
Re: Summary of June 8 outage
#67Absolutely no details about the bug or why a single customer configuration effected global state on the server, or why this wasn't caught by configuration change safety mechanisms/smoke tests/gradual rollout. Also, what is up with their partitioning? Do they seriously have one customer that gets served from 85% of their servers? Is it a whale? Good on them for getting a statement out right away (although they basical…
There could be a feedback loop that is the opposite of a smoke test.
1. Validate customer configuration, if passes, assume it can roll out 2. Roll out customer configuration to node 3. Node goes down 4. Migrate all customers on node to new nodes 5. Node that problematic customer was migrated to goes down 6. Rinse and repeat as problematic customer migrates to every node and takes out every last one.
Re: Summary of June 8 outage
#68Thought it surprising that amazon.com uses fastly when they have their own cdn
https://www.streamingmediablog.com/2020/05/fastly-amazon-hom...
Re: Summary of June 8 outage
#69You took down my site and a good swath of the whole internet. I am entitled to know, in detail, what happened, so I can be more informed and assess any actions I might need to take. I don't want any more of your PR speak or "we value our customers". That's crap and insults my intelligence. STOP getting PR to write your comms; just speak to engineers like engineers. I'd rather get no response than this post. I hope th…
Obviously on this site we tend to be rather technical people, so we want to know as much detail as possible, but that's something we desire, not something we are entitled to.
Re: Summary of June 8 outage
#70Yeah - not really a post mortem, is it? "We had a bug and we fixed the bug." They could probably cut and paste that same page for 90% of future outages. Maybe they need to read this: https://artsy.github.io/blog/2014/11/19/how-to-write-great-o...
So I don't think they are claiming this is a post mortem.