Live data from Hacker News

Fastly Outage

fastly.com

541–550 of 740 posts

Re: Fastly Outage

#541

This seems to be impacting a number of huge sites, including the UK government website[0]. [0] https://www.gov.uk/ https://m.media-amazon.com/ https://pages.github.com/ https://www.paypal.com/ https://stackoverflow.com/ https://nytimes.com/ Edit: Fastly's incident report status page: https://status.fastly.com/incidents/vpk0ssybt3bj

Is their anything these big sites could do in this situation, or must they choose between running and maintaining all of their own infra or relying on a single CDN?

Re: Fastly Outage

#542
post #381

Earlier quoted context omitted.

The point isn't to dance around the incident, but to not blame people. You can blame systems, design, engineering culture, processes, but don't blame people. Even if someone accidentally pressed the 'destroy prod' button, that's not the fault of that person, it's the fault of that button existing and being accessible in the first place. I have no empathy for Fastly-the-company. I hate the fact that the Internet is ce…

I disagree. People implemented those systems, so if you are correct that it is the systems fault, then it is also a persons fault. People must be held accountable to have good incentives to reduce such outtages in the future. I do agree though that we should always be compassionate and realistic with other humans.

> People must be held accountable to have good incentives to reduce such outtages in the future.

Holding specific people "accountable" for outages doesn't incentivize reducing outages; it incentivizes not getting caught for having caused the outage.

As a result, post-mortems turn into finger-pointing games instead of finding and resolving the root cause of the issue, which costs the company more money in the long run when a political scapegoat is found but the actual bug in the code is not.

Re: Fastly Outage

#543

Earlier quoted context omitted.

> [0] https://www.gov.uk/ Just checked, thank god the NHS vaccine site is still available - vaccines just got rolled out for under 30s today.

Edit: I didn’t mean anything negative here! Just slightly shocked that as the UK is opening up under 30 vaccinations, the US is struggling to find any more willing takers. It’s really probably a sign that there’s fewer anti-vaxxers in the UK more than anything. And that universal healthcare is more efficient at distribution than an inherently for profit system. I don’t know, but I just didn’t realize it was so differ…

US and UK have very similar vaccination rates despite the US being open to more age ranges. This indicates that a higher percentage of eligible people have gotten the vaccine in the UK, and the US has somewhat hit a wall in terms of vaccinations (though there is the concern that the rates will slow down in the UK also).

I must admit, it has been strange seeing my US peers getting the vaccine months before I can in the UK, but I guess I take comfort knowing that both countries are still doing pretty well!

Re: Fastly Outage

#544

This seems to be impacting a number of huge sites, including the UK government website[0]. [0] https://www.gov.uk/ https://m.media-amazon.com/ https://pages.github.com/ https://www.paypal.com/ https://stackoverflow.com/ https://nytimes.com/ Edit: Fastly's incident report status page: https://status.fastly.com/incidents/vpk0ssybt3bj

But https://news.ycombinator.com/ is UP! :) Prepare those HN servers for massive influx in 3...2..1..

Might not be the case anymore, but a few years back, Hackernews was just running on a single server.

Re: Fastly Outage

#545

This seems to be impacting a number of huge sites, including the UK government website[0]. [0] https://www.gov.uk/ https://m.media-amazon.com/ https://pages.github.com/ https://www.paypal.com/ https://stackoverflow.com/ https://nytimes.com/ Edit: Fastly's incident report status page: https://status.fastly.com/incidents/vpk0ssybt3bj

Fastly Engineer 1: Seems like a common error message. Can you check stackoverflow to see if there's an easy fix? Fastly Engineer 2: I have some very bad news...

Oh man, how do we keep a pocket copy of SO? All of our jobs depend on it.

Re: Fastly Outage

#546
post #544

Earlier quoted context omitted.

But https://news.ycombinator.com/ is UP! :) Prepare those HN servers for massive influx in 3...2..1..

Might not be the case anymore, but a few years back, Hackernews was just running on a single server.

fairly sure it still is.

Re: Fastly Outage

#547

Yep, seems like: Reddit BBC News Twitch.tv Twitter emoji cdn? are all down 503 service error

Ah didn't cop that Twitter emoji issue was related! Thought an ad-blocker was stepping up its filters aggressively :) Stack Overflow, The Guardian, Gov.uk too as some other biggish names getting hit.

Various bits of GitHub on the Web (committing edits, editing releases) were broken for the same reason. Failure modes of JS-heavy GUIs are interesting.

Re: Fastly Outage

#548

Stupid question: why didn't sites "just" fail over to their actual servers to handle the traffic, albeit slowly? I guess they won't be sized to handle the load in a lot of cases, and Fastly was responding, so DNS fail over didn't work?

Probably a different answer for each site. I'm not a DNS expert but I think you're right on both counts. Having failover also requires a duplicate CDN architecture at the fallback location, which is an increase of costs in time, money & maintenance for relatively little benefit. Often there's a fair amount of background integration with a CDN, and each function slightly differently, so it's not simply plug & play.

Re: Fastly Outage

#549
post #494

Earlier quoted context omitted.

Meh. Losing sleep sounds like an over-reaction. No system is foolproof. Of course Fastly should do what they can to prevent downtime, but it's still expected that they will go down. I would blame anyone who claimed otherwise or couldn't deal with it while not having a fallback.

I hear that you're suggesting that those involved shouldnt feel bad because its a systemic / just a job / etc. But the reality is that incidents like this can be very traumatic for those involved and thats not something they can control. If it was that simple to manage, depression and anxiety would not be a thing. Think its best to show a large amount of support and empathy for the individuals having a really bad day…

I wanted to show support to the engineers in the sense that I don't think you should encourage a working culture where you have "massive post-mortems" and expect people to feel bad for extended periods of time over simple mistakes. By not making a big deal out of it, you can also support your staff.

But I think our disagreement mainly stems from how we interpreted the parent comment. I thought it was very double, at one hand claiming to show support, at the other hand emphasizing how big of a catastrophy this was.

I just wanted to say that I think it most likely was a completely natural mistake, only exerbarated by the scale of the company, and that while you should take some action to prevent it in the future, you should not spend so much time dwelling on it. Shit happens, it's fine.

Re: Fastly Outage

#550

Yeah so it's been mentioned in the comments already, but to everyone in Fastly right now: I feel for you. Something like this must be insanely stressful, and not just during the outage. There will be (should be) a massive post-mortem. People will be losing sleep over this for days, weeks, months. :( Edit: There seems to be a major empathy outage in this thread. Disgusted but not surprised, unfortunately.

They shouldn't lose sleep over it, though.

Everyone's got responsibilities and aspirations. To be fair, I was thinking more of the jobbing engineer who's going to face anxiety about losing their job over this, but it extends to all levels. Having a fat bank balance helps get through periods without employment, but it's not just about money. There's anxiety, shame, embarrassment, the whole gamut. Going through a major incident at work is a shitty experience.
Post reply on HN