Live data from Hacker News

Coding Horror and blogs.stackoverflow.com experience "100% Data Loss"

codinghorror.com

141–150 of 175 posts

Re: Coding Horror and blogs.stackoverflow.com experience "100% Data Loss"

#141
post #115
post #7

To be fair, Atwood thought his hosting provider (CrystalTech) was backing up his system. As it turns out, their entire VM backup solution failed silently, so everyone thought the backups were being made. If anything, I'd say this is a sign not to work with CrystalTech.

It doesn't seem to be particularly fair to the hosting provider, really, to put their name on the failpage and to tweet about how it's half their fault. Especially after you've advocated redundant backups, implied that you have them, written about the advantages of hosting your images on S3. When a mishap reveals that you've actually done none of these things, it is more than a little disingenuous to try to emphasize…

Relatively minor?! They failed to do what they said they would. They destroyed data and borked up the backups.

If it wasn't Jeff Atwood but me, would it be more their fault? Or would I just be less disingenuous?

Nobody's calling them names, or threatening to take business elsewhere or anything. But it's good to get sunlight in there, show them consequences, when they fail you. When a company providing you with service fails, you're allowed to scream it from the rooftops if you feel like it.

Re: Coding Horror and blogs.stackoverflow.com experience "100% Data Loss"

#142
I sent jeff tarballs from blekko's webcrawl for www.codinghorror.com, blog.stackoverflow.com, www.fakeplasticrock.com and haacked.com - about 6300 pages overall. He's got Coding Horror back up from the basic html.

Unfortunately we don't have the images, but it looks like most of the site is back up at least. It will probably be more work for him to re-integrate it into the cms though.

Re: Coding Horror and blogs.stackoverflow.com experience "100% Data Loss"

#143

Earlier quoted context omitted.

The bank analogy is interesting and a good point, but the reality of ISPs is that they are less reliable, and less regulated, than banks. If my income is dependent on data, I make sure it gets backed up. If it's an active project, I use my backups to build my dev environments so it's fairly obvious when a backup has failed.

You have good habits, and that's sincerely commendable. However, it really isn't fair to blame this kind of problem on the end user. Are website developers also expected to keep the servers secure? If Apache isn't patched and up to date, is that the website admin's fault? Service providers are paid to do a job. This one failed tremendously, and should lose a large chunk of business for it. I think Atwood is being far…

To put it another way: the only reason to maintain your own backups of your site data -- aside from healthy paranoia -- is because you expect your service provider to fail at doing their job. And if that's the case, shouldn't you be finding a service provider that does it better?

If you don't expect your service provider to fail, you don't know anything about service providers.

Everybody fails.

Our colocation facility has redundant generator systems. They're tested regularly, and have handled failures previously. Yet, when the power went out, three of the backup generators failed, and our site (as well as Craigslist, Yelp, and others) was out for 45 minutes.

The cause? A bug in the backup generator's software: http://365main.com/status_update.html

Shit happens. Sometimes it's not your fault. You still need to prepare for it.

Re: Coding Horror and blogs.stackoverflow.com experience "100% Data Loss"

#144
post #125

Earlier quoted context omitted.

You can't outsource liability. So you've tested whether your toaster has proper grounding and other safety precautions, in case it shortcircuits? I bet you haven't and that why you should stop repeating that stupid soundbite. We all 'outsource' liability all the time: we pay others to perform services for us and hold them responsible for the proper execution of those services. This includes hosting content and backin…

So you've tested whether your toaster has proper grounding and other safety precautions, in case it shortcircuits? Nope. I have a fully paid and verified home owners insurance policy though that covers any probable loss from a toaster fire. It's sort of like having an offsite backup of my important personal stuff. Your example is really not a parallel to a data backup. It's difficult to fully test all of the toasters…

An insurance policy is equivalent to an offsite backup? So you don't have any personal items of any intrinsic value at all; you could lose it all, get a cheque in return, and be happy? Wow. Well, awesome, but I doubt that is common.

And look, I agree with you in many ways - people need to take responsibility for their own backups, sure. But there is a division of responsibility. I mean, even if all my backups are 100% perfect, I am still trusting the hosting provider to, well, keep the power on. Pay the peering bill. Keep the server temp down. Not go bankrupt tomorrow.

You can just follow this chain as far as you want. Whether you like it or not, you're utterly dependent on the DNS root server admins. There is nothing at all you can do to prepare yourself for their failure. I bet you can't generate your own electricity or grow your own food, either.

All of civilisation is built on co-dependency and delegation of responsibility. It's the only way to do anything complex. At some point, you must delegate. Atwood should have checked - but he was paying his host to do it. That's like having an employee whose job it is to do backups. At some point, you just have to let go and trust them. Otherwise you can never really do anything; you're caught up in checking minutia; I can give examples of this kind of leadership failure until my keyboard breaks.

Re: Coding Horror and blogs.stackoverflow.com experience "100% Data Loss"

#145
post #115

Earlier quoted context omitted.

It doesn't seem to be particularly fair to the hosting provider, really, to put their name on the failpage and to tweet about how it's half their fault. Especially after you've advocated redundant backups, implied that you have them, written about the advantages of hosting your images on S3. When a mishap reveals that you've actually done none of these things, it is more than a little disingenuous to try to emphasize…

Relatively minor?! They failed to do what they said they would. They destroyed data and borked up the backups. If it wasn't Jeff Atwood but me, would it be more their fault? Or would I just be less disingenuous? Nobody's calling them names, or threatening to take business elsewhere or anything. But it's good to get sunlight in there, show them consequences, when they fail you. When a company providing you with servic…

Their role in making him look foolish (what I said, incidentally) is indeed quite minor. If he'd actually done the things he advocated and, frankly, implied that he'd done, he'd still have his images, his downtime would have been close to zero and he'd probably get to write a triumphant article about the value of following his own sage advice.

Someone else pointed out in another comment, there's a good analogy to be drawn to Atwood's own words here: http://www.codinghorror.com/blog/archives/001079.html

Re: Coding Horror and blogs.stackoverflow.com experience "100% Data Loss"

#146

I want to make a snarky comment, but I just feel for the guy.

I feel for the guy too, but it really isn't the first time in the last couple of months that some service provider fails at their #1 stated goal: to keep your data safe. And it isn't just the small ones either. For the love of - insert your favorite deity here - please try to restore your backups, and try to do so on a regular basis. If not all you might have is the illusion of a backup. It is a very easy trap to fal…

There's an added bonus to restoring backups on a regular (daily) basis: a one day old instance of the production environment is always available as a playground for training, dev or qa.

Re: Coding Horror and blogs.stackoverflow.com experience "100% Data Loss"

#147

Earlier quoted context omitted.

Joining the chorus of other people saying "WRONG". Reason 1: I pay for my bank account one way or another. It's bank's responsibility to keep the system running and secure, not mine. Sometimes we just can't do everything ourselves and have to rely on others. Reason 2: Many of us host something somewhere. How many do backups? How many of us check the backups? How many do check the backups checking process? (enter infi…

Reason 1: I pay for my bank account one way or another. It's bank's responsibility to keep the system running and secure, not mine. Sometimes we just can't do everything ourselves and have to rely on others. You are responsible for reading your statement and ensuring that all activity is valid, much in the same way that you are responsible for ensuring the viability of your recovery strategy. Reason 2: Many of us hos…

You cut a lot from reason 2. The was an important part. I've seen a system where the backups were made. They were verified too. Only after an actual data loss it was discovered that the backup verification was faulty and half of the "verified" and "properly backed up" data was missing.

My point is that you can spend lots of hours trying to backup your data and verify its correctness. But unless you put it back into the actual working environment and check every single bit of it, you cannot be sure it was a proper backup - not with 1GB of information and certainly not with 1TB. And then after you verified that you can verify that you have backups, something will fail and a bunch of people will say "if you didn't check the backups properly, it's your fault". I've seen backup systems fail in amazing ways and will probably never again believe that you can be "sure".

Re: Coding Horror and blogs.stackoverflow.com experience "100% Data Loss"

#148
post #62

1. Register an Amazon AWS S3 account - https://aws-portal.amazon.com/gp/aws/developer/registration/... 2. Download my S3 backup script (or anyone's S3 Backup script) - http://github.com/leftnode/S3-Backup 3. Set up a cron to push hourly/daily/whatever tar's of your vhost's directory to S3. Spend, like, $10 a month. Thats 30gb of storage, 30gb up and 30gb down. Now, I know that may not be a lot, but I doubt codinghorr…

If you push your backups, make sure the account used does not have permissions to modify or delete the files it is creating offsite.

Yes, this! Push backups are inherently risky -- if at all possible, backups should either be pulled by the target, or should be mediated by a third system. Otherwise, you risk an attacker deleting all your backups along with your data.

Re: Coding Horror and blogs.stackoverflow.com experience "100% Data Loss"

#149

Earlier quoted context omitted.

The bank analogy is interesting and a good point, but the reality of ISPs is that they are less reliable, and less regulated, than banks. If my income is dependent on data, I make sure it gets backed up. If it's an active project, I use my backups to build my dev environments so it's fairly obvious when a backup has failed.

You have good habits, and that's sincerely commendable. However, it really isn't fair to blame this kind of problem on the end user. Are website developers also expected to keep the servers secure? If Apache isn't patched and up to date, is that the website admin's fault? Service providers are paid to do a job. This one failed tremendously, and should lose a large chunk of business for it. I think Atwood is being far…

>Are website developers also expected to keep the servers secure? If Apache isn't patched and up to date, is that the website admin's fault?

Getting hacked isn't as potentially catastrophic as not having backups. With backups, being hacked can be recovered from and the service provider changed.

>To put it another way: the only reason to maintain your own backups of your site data -- aside from healthy paranoia -- is because you expect your service provider to fail at doing their job.

A business doesn't expect its premises to burn down, but most have fire insurance in the event that this happens. Even if I don't expect my service provider to fail, there's no way to know that they won't and it makes sense to deal with this risk if the cost of dealing with it is reasonable and the cost of not dealing with it is catastrophic.

Re: Coding Horror and blogs.stackoverflow.com experience "100% Data Loss"

#150
post #145

Earlier quoted context omitted.

Relatively minor?! They failed to do what they said they would. They destroyed data and borked up the backups. If it wasn't Jeff Atwood but me, would it be more their fault? Or would I just be less disingenuous? Nobody's calling them names, or threatening to take business elsewhere or anything. But it's good to get sunlight in there, show them consequences, when they fail you. When a company providing you with servic…

Their role in making him look foolish (what I said, incidentally) is indeed quite minor. If he'd actually done the things he advocated and, frankly, implied that he'd done, he'd still have his images, his downtime would have been close to zero and he'd probably get to write a triumphant article about the value of following his own sage advice. Someone else pointed out in another comment, there's a good analogy to be…

"Their role in making him look foolish (what I said, incidentally) is indeed quite minor."

Fair enough. But he's not complaining about looking foolish (nor is grandparent as far as I can see). He's complaining about losing data.

In that link he isn't saying "it's your fault no matter what the world throws at you, suck it up." He's saying, "look harder at yourself before you decide somebody else is to blame." Not an issue here; can we agree it is their fault?

Post reply on HN