Live data from Hacker News

Twitter Completely Down

isup.me

161–170 of 194 posts

Re: Twitter Completely Down

#161
post #146

Earlier quoted context omitted.

> I'm fairly certain that the higher-ups in Twitter weren't told "We have pretty good failover protection, but there is a small risk of catastrophic failure where everything will go completely down." Whoever was in charge of disaster recovery obviously didn't really understand the risk. Is this really a valid conclusion to come to at this point? I expect downtime in any service I operate. It's just how the world work…

Do you expect Google search to have downtime?

Google search has downtime.

Re: Twitter Completely Down

#162
post #78

I think this is another good example of how we as an industry are still unable to adequately assess risk properly. I'm fairly certain that the higher-ups in Twitter weren't told "We have pretty good failover protection, but there is a small risk of catastrophic failure where everything will go completely down." Whoever was in charge of disaster recovery obviously didn't really understand the risk. Just like the recen…

Whoever was in charge of disaster recovery obviously didn't really understand the risk. That's not necessarily true. People don't die when twitter is down, and whatever twitter's business model actually is, I am not even sure there is a monetary penalty to them being down (unlike, say, Amazon being down which results in lost orders). They may have made the calculation that it was not cost effective engineering-wise t…

> [Edit: Pedantry shield: Ok, ok, should have said people don't die because twitter is down. Obviously people are dying all the time, and some will indeed expire while twitter is down].

And here I was hoping we could just take down Twitter and live forever...

Re: Twitter Completely Down

#163

Earlier quoted context omitted.

Accidents happen. (expanding my reply) No risk assessment in the world will stop a cage monkey from tripping over a pile of 1Us and falling onto the big red button. Figure out what your pain threshold is and live with it.

This is called "acceptable risk". If it would cost $3m to eliminate the risk and the damages would only total $1.5m, that would be an acceptable risk.

What I was trying to point out is that catastrophe is inevitable and risk assessment isn't a panacea. I'd rather invest money in proactive countermeasures to unknown risks than try to think of all the things that could go wrong (which you'll never have the money in your life to fix anyway).

But also to the thread parent's comment about Twitter execs not realizing there was a minor risk of catastrophic failure: Executives only care that their money-making baby keeps running. In the past i've seen execs demand that an engineer call them at 3AM if the production site goes down for more than 5 minutes... even though that call is pointless. I think they just assume there's no point in getting involved with the plan because the plan will never be perfect, but at least they can be aware of a problem so they can cover their asses and tell a higher-up that it's being worked on. At the end of the day, even the guys at the top don't really give a shit about the product, they just care about their paycheck.

Re: Twitter Completely Down

#164

Earlier quoted context omitted.

xx Next time your favorite waste of time is down, just shut up about it. xx Twitter is infrastructure for us in media. It's well worth discussion.

that's what you get for not using RSS. also, "us in the media", what kind of whore talk is that even? twitter is correctly referrerd to in w3c docs as medium preventing intelligent discussion. so it was down? GOOD. people are inconvenienced? even better! it cannot possibly have hit anyone or anything that was worth fuck all.

Cool it on the ad hominem please

I'm not sure why you suggest RSS is somehow synonymous with Twitter, but I will say that in addition to RSS buttons almost every major media company on the planet has a Twitter button on its article page (I work on one, which is why I say "us in media"). Because many sites don't do proper async JS when it comes to social buttons, an outage on socnets can be crippling

Here's an example: http://techcrunch.com/2012/06/01/facebook-outage-affects-oth...

Re: Twitter Completely Down

#165
post #78

I think this is another good example of how we as an industry are still unable to adequately assess risk properly. I'm fairly certain that the higher-ups in Twitter weren't told "We have pretty good failover protection, but there is a small risk of catastrophic failure where everything will go completely down." Whoever was in charge of disaster recovery obviously didn't really understand the risk. Just like the recen…

Whoever was in charge of disaster recovery obviously didn't really understand the risk. That's not necessarily true. People don't die when twitter is down, and whatever twitter's business model actually is, I am not even sure there is a monetary penalty to them being down (unlike, say, Amazon being down which results in lost orders). They may have made the calculation that it was not cost effective engineering-wise t…

An e-commerce site being down does not lead to exactly the amount of orders lost as is average for that time. Most people will try again later, with the possible exception of first time buyers and likely exceptin of impatient commodity buyers with alternative accounts.

Re: Twitter Completely Down

#167
post #96

Earlier quoted context omitted.

> But this sort of news story about something that's happening right now is symptomatic of the useless '24 hour news' cycle of noise. The single most important part of the internet is the immediate availability of news (to me, anyway). I've never heard anyone complain about that before; why do you think it's not worth knowing and talking about events as they happen? 'Twitter is down' isn't noise. Years and years ago…

So, here in the UK Twitter is back for me. It looks like I was without it for about 45 minutes. Call it an hour for a nice round figure. Think about these two scenarios: 1. During that one hour you spend your time focussed on talking about this event as it's happening, speculating, having an emotional response (because you can't access something you want to and find a group of people experiencing the same thing and a…

As a person whose love of news started at age five scrapbooking newspaper clippings about Nixon, Ford, and Carter, your point makes me very happy.

The 24 hr news cycle ruined TV news, and SEO has made Internet news worse (first links win).

Re: Twitter Completely Down

#169

I think this is another good example of how we as an industry are still unable to adequately assess risk properly. I'm fairly certain that the higher-ups in Twitter weren't told "We have pretty good failover protection, but there is a small risk of catastrophic failure where everything will go completely down." Whoever was in charge of disaster recovery obviously didn't really understand the risk. Just like the recen…

You've never done DR have you? It's a business process with a cost and there are RTO's (recovery time objectives) and RPO's (recovery point objectives). Systems can and will go down. So long as the recovery meets the defined objectives, then DR has been performed correctly. There is a limited amount of money and resources that businesses can spend on DR, COOPs, etc. You should understand that.

One hour RTO and RPO will cost way more than 24 hour recovery. Edit... and the business managers decide how much they wish to spend on DR. It's a trade-off and anyone who has ever done it, understands that.

Re: Twitter Completely Down

#170
post #74

I think this is another good example of how we as an industry are still unable to adequately assess risk properly. I'm fairly certain that the higher-ups in Twitter weren't told "We have pretty good failover protection, but there is a small risk of catastrophic failure where everything will go completely down." Whoever was in charge of disaster recovery obviously didn't really understand the risk. Just like the recen…

Your examples describe two entirely different systems. The failover of a software product is drastically different from the failover of a power system. Trying to map everything back to a common best practice under the category of "risk" seems like it would miss out on important intricacies.

Risk management is about determining how to identify risks, as such, it is applicable everywhere. However, much like security is applicable everywhere, securing Fort Knox is a very different endeavor than securing a web site.
Post reply on HN