Live data from Hacker News

Summary of June 8 outage

fastly.com

31–40 of 110 posts

Re: Summary of June 8 outage

#31
The "valid customer configuration change" seems to cover the angle that the user input was validated and valid but the backend implementation of said configuration was the buggy part. Look forward to actual details from them.

Re: Summary of June 8 outage

#32
I love that somewhere out there is a developer who doesn't even work at Fastly but just innocently pushed a change to their Fastly config and basically broke the entire internet. I'm actually jealous. If it was me, I'd put that on my resume.

Re: Summary of June 8 outage

#33

So, a valid customer configuration change triggered a bug. One thing I don't see in this writeup is a commitment to ensure that customer configurations cannot break the whole system. Cloudflare does seem to make this promise with their zero trust architecture, https://www.cloudflare.com/learning/security/glossary/what-i...

Why do so many big companies use Fastly when Cloudflare (from the outside, as someone who doesn't know much about the space) looks to be so much cleaner and more technically sophisticated? Am I being brainwashed by their blog posts?

I don't know about you but I find the prospect of a Cloudflare monoculture pretty worrying, especially since they've already demonstrate a willingness to kick off users they don't like. (I also think the https veneer that they offer is misleading to end users and bad for everyone on the internet, though not everyone will agree with that).

Re: Summary of June 8 outage

#34
post #32

I love that somewhere out there is a developer who doesn't even work at Fastly but just innocently pushed a change to their Fastly config and basically broke the entire internet. I'm actually jealous. If it was me, I'd put that on my resume.

I can't actually imagine something like that can happen. Single person with a simple change in a config can cause this.

Re: Summary of June 8 outage

#35
I might sound naive, but as a true hyperscale internet company to plan for disaster scenario like this as a consumer of fastly?

How could you plan for an outage like this by fastly and how could you mitigate this?

Re: Summary of June 8 outage

#36
post #32

I love that somewhere out there is a developer who doesn't even work at Fastly but just innocently pushed a change to their Fastly config and basically broke the entire internet. I'm actually jealous. If it was me, I'd put that on my resume.

I can't actually imagine something like that can happen. Single person with a simple change in a config can cause this.

They forgot to test it.

Re: Summary of June 8 outage

#37
You took down my site and a good swath of the whole internet. I am entitled to know, in detail, what happened, so I can be more informed and assess any actions I might need to take.

I don't want any more of your PR speak or "we value our customers". That's crap and insults my intelligence. STOP getting PR to write your comms; just speak to engineers like engineers. I'd rather get no response than this post.

I hope there are actual details as they complete their investigation. If there isn't a public post-mortem, I am switching away from Fastly.

Re: Summary of June 8 outage

#38

This is annoyingly vague. What was the software bug, and what was the valid customer configuration change? It's perhaps a bit premature to demand it at this point, but I'm hoping a full post-mortem will outline precisely how this change was not picked up in pre-prod. Surely all valid customer configurations must be tested prior to rollout.

It's not just annoying value. It's insultingly vague.

If my data centre provider suffered a complete outage, then I demand to get a detailed post-mortem of what happened (in due time). If they just tell me bullshit PR speak about "We value our customers", I'll be looking at switching providers.

As a Fastly customer whose site went down, I'm entitled to know exactly what happened. If they don't tell me, I'm switching CDNs as a matter of priority.

Re: Summary of June 8 outage

#39

I might sound naive, but as a true hyperscale internet company to plan for disaster scenario like this as a consumer of fastly? How could you plan for an outage like this by fastly and how could you mitigate this?

Use a short TTL on the CDN subdomain you use. Then setup an alternative CDN provider in advance, so that you can switch from one to the other in a matter of minutes.

Re: Summary of June 8 outage

#40
post #7

Absolutely no details about the bug or why a single customer configuration effected global state on the server, or why this wasn't caught by configuration change safety mechanisms/smoke tests/gradual rollout. Also, what is up with their partitioning? Do they seriously have one customer that gets served from 85% of their servers? Is it a whale? Good on them for getting a statement out right away (although they basical…

> Is it a whale? TikTok is my guess. ByteDance is valued at 250 billion. Plus, the change was pushed in the middle of the night, which would be daytime in Asia. Certainly there are other development teams in Asia, but considering the scale of the change it likely comes from HQ, and Fastly's whale in Asia would be them. edit They may have lost TikTok at the end of last year, either partially or completely [1]. Anyone…

My guess is reddit. It was down yesterday due to this.
Post reply on HN