Live data from Hacker News

Fastly Outage

fastly.com

561–570 of 740 posts

Re: Fastly Outage

#561

Earlier quoted context omitted.

StackOverflow and all the StackExchange family of sites are down. I suspect the lost productivity from that will be more costly over the whole economy than potential lost sales via Shopify. People can go back to shopify so those transactions not definitely blocked for ever, any time "lost" due to reference resources being unavailable can't so easily be claimed back.

I don't think you understand how ecommerce works. A very significant amount of people won't go back. It's why the most effective marketing campaign by far is retargeting those people to convince them to come back. Unfortunately that's not possible in this case since you can't track the users as the site is unusable.

> A very significant amount of people won't go back

So they didn't need what they were about to purchase and saved their money. Doesn't sound like a net loss to me.

Re: Fastly Outage

#562
post #399

Earlier quoted context omitted.

Call me old fashioned but the latest trend of showing "empathy" for a serious incident, then proceeding to dance around the aftermath of it, whilst people give themselves a pat on back in a retro/post-mortem, isn't the way to do it. People need to be blamed, and responsibility for actions taken (without covering asses)

Blame culture isn't the way forward here. Do a post-mortem, work out root causes, work as a unit to ensure this doesn't happen again. Obviously if there are levels of gross negligence or misconduct discovered during post-mortem, that will need to be dealt with accordingly, but coming into this with an attitude of "we must find someone to blame and incur repercussions" isn't healthy at all. We are humans - don't forge…

[deleted]

Re: Fastly Outage

#563
post #494

Earlier quoted context omitted.

I hear that you're suggesting that those involved shouldnt feel bad because its a systemic / just a job / etc. But the reality is that incidents like this can be very traumatic for those involved and thats not something they can control. If it was that simple to manage, depression and anxiety would not be a thing. Think its best to show a large amount of support and empathy for the individuals having a really bad day…

I wanted to show support to the engineers in the sense that I don't think you should encourage a working culture where you have "massive post-mortems" and expect people to feel bad for extended periods of time over simple mistakes. By not making a big deal out of it, you can also support your staff. But I think our disagreement mainly stems from how we interpreted the parent comment. I thought it was very double, at…

I agree, and I think I picked on your comment a bit because it was the top one.

Re: Fastly Outage

#564

This seems to be impacting a number of huge sites, including the UK government website[0]. [0] https://www.gov.uk/ https://m.media-amazon.com/ https://pages.github.com/ https://www.paypal.com/ https://stackoverflow.com/ https://nytimes.com/ Edit: Fastly's incident report status page: https://status.fastly.com/incidents/vpk0ssybt3bj

Fastly Engineer 1: Seems like a common error message. Can you check stackoverflow to see if there's an easy fix? Fastly Engineer 2: I have some very bad news...

Well, with SO, at least you can search on Google and view the version cached by Google just fine.

With Reddit however, these days almost all comments are locked behind “view entire discussion” or “continue this thread”. In fact, just now I searched for something for which the most relevant discussion was on Reddit; Reddit was down so I opened the cached version, and was literally greeted by five “continue this thread”s and nothing else. What a joke.

Re: Fastly Outage

#565

This seems to be impacting a number of huge sites, including the UK government website[0]. [0] https://www.gov.uk/ https://m.media-amazon.com/ https://pages.github.com/ https://www.paypal.com/ https://stackoverflow.com/ https://nytimes.com/ Edit: Fastly's incident report status page: https://status.fastly.com/incidents/vpk0ssybt3bj

Is their anything these big sites could do in this situation, or must they choose between running and maintaining all of their own infra or relying on a single CDN?

Use two CDNs and DNS providers for redundancy. Gets expensive, but at scale, probably doesn't make a huge difference. More complexity for the site operators to manage, however.

Re: Fastly Outage

#566
post #480
post #181

Earlier quoted context omitted.

That is true, but the typo is because they presumably meant to reference this: https://en.wikipedia.org/wiki/Guru_Meditation

...and naturally someone has already updated the page today to include (and highlight) its use and mispelling on Varnish.

Someone (I can't unfortunately due to IP block) needs to change that. The part about the spelling is false, apparently [1] it's an intentional change by Fastly so that they can tell if it's their own Varnish or a customer's Varnish that is throwing an error.

[1] https://news.ycombinator.com/item?id=27433139

Re: Fastly Outage

#567

This seems to be impacting a number of huge sites, including the UK government website[0]. [0] https://www.gov.uk/ https://m.media-amazon.com/ https://pages.github.com/ https://www.paypal.com/ https://stackoverflow.com/ https://nytimes.com/ Edit: Fastly's incident report status page: https://status.fastly.com/incidents/vpk0ssybt3bj

Fastly Engineer 1: Seems like a common error message. Can you check stackoverflow to see if there's an easy fix? Fastly Engineer 2: I have some very bad news...

Haha! That explains why the internet was down for a while!

Re: Fastly Outage

#568

Earlier quoted context omitted.

Call me old fashioned but the latest trend of showing "empathy" for a serious incident, then proceeding to dance around the aftermath of it, whilst people give themselves a pat on back in a retro/post-mortem, isn't the way to do it. People need to be blamed, and responsibility for actions taken (without covering asses)

I'm sure you've never made a mistake. The best way (in a team), to tackle mistakes, is to ensure the process in place corrects these mistakes. The only way to do that, is a post-mortem/learning from the mistake. If you blame it on some engineer who did it, that guy will eventually be replaced by some other guy, who may make the same mistake.

You also need to be proactive about other possible failure modes. Avoiding a culture of blame may or may not help. There needs to be a strong incentive for the organization to expend the resources to do so, and a mere "oops my bad" doesn't provide that without SLAs with teeth.

Re: Fastly Outage

#569
post #452
post #324

Earlier quoted context omitted.

Believe it or not but "the internet" and "the world wide web" are not synonyms.

True. But the vast majority of use goes via "WWW". For example email - the other big "internet-user" is technically not part of the WWW, but most (? I don't have any stats, just a guess) of our mailclients run on the WWW, nonetheless.

BitTorrent was half of all Internet traffic for a while, though it has decreased with the rise of legal and convenient streaming services.

Re: Fastly Outage

#570

Earlier quoted context omitted.

Fastly Engineer 1: Seems like a common error message. Can you check stackoverflow to see if there's an easy fix? Fastly Engineer 2: I have some very bad news...

Oh man, how do we keep a pocket copy of SO? All of our jobs depend on it.

no they don't, but if yours does you can download a complete datadump of SO from them.
Post reply on HN