Live data from Hacker News

“Let's talk about a hypothetical public-facing service”

reddit.com

41–50 of 72 posts

Re: “Let's talk about a hypothetical public-facing service”

#41
post #29

Earlier quoted context omitted.

It costs about a hundred and fifty grand annually (fully loaded, base salary 80-90k) to have the lowest end person who could conceivably keep one of these things running well for three years. n+1 means you need two of them in case one gets the flu. Tell me again about how Amazon's a ripoff, please.

You're still going to need n+1 sysadmins to maintain the rest of your infrastructure. Maintaining a storage system like this really isn't much different for the day-to-day operations. So, yes, it's one more system, but you're already going to need those admins anyways.

I don't believe maintaining a high availability system with proper checking is less than 3 hours a day of work, or ~10% of someone's time on duty if we're talking around the clock.

So we're talking about $75,000-90,000 a year in salary to maintaining the cluster if you want 24x365 coverage (which Amazon provides), as you'll need 2 people per shift, 3 shifts per day at a minimum to have people in house, even if they're only spending 10% of their time actually working on that particular issue. In reality, these are unrealistically small numbers. Each employee will have a cost to the corporation of $125,000-150,000 a year. Employees of that caliber spending 10% of their time on supervising a cluster for a year is $75,000-90,000. I'm amortizing the amount of work over your 3 datacenters by imagining you already have the staff and only counting out the number of hours needed for just this work.

So the reality is that you left off something like $225,000-270,000 in actual cost of running your own cluster from your analysis, because while it's not "much different", it is a few hours a week from at least 6 employees if you're really talking about running managed, highly reliable storage.

Re: “Let's talk about a hypothetical public-facing service”

#42
post #2

I was hoping they would stay offline. It's really disappointing to see a good service turn into a money wringing desperation.

The way I discovered SF was down was when I was trying to install some Slackbuilds and kept getting MD5 mismatches on the source downloads. I noticed I was getting the same MD5 for several different packages, so I tried to manually grab a package and that's when I hit SF's error page (which sbopkg was happily downloading and trying to pass off as somesourcepkg.tar.gz). I found myself in the odd position of hoping SF…

Sadly I too found it was screwed a few days ago when trying to pull some source code hosted there (ironically so I would have my own copy in case things went wrong there...). I just got their 'Disaster Recovery Mode' notice and a subset of binary packages available.

Lots of google searching for what's wrong with Sourceforge revealed nothing (other than complaints about their packaged installers/crapware). Now I have the answer to the question I was really asking.

I just need to wait and see if the developers of the packages I need will somehow migrate to github so I can get the source...

Re: “Let's talk about a hypothetical public-facing service”

#43
post #25

Got a good chuckle. Such is life in IT. Gotta admit I missed the "Extremely Massive Corporation" hint.

It reminds me of Douglas Crockford's talk where he mentioned that a company asked him for an exemption to the "do no evil" clause in the JSON License, and he said he didn't want to name the company as that would embarrass them so he would instead give their initials: IBM.

It's not embarrassing; their lawyers are just being cautious, as befits a company with huge customers like theirs. The "do no evil" clause is stupid and utterly ambiguous. If Crockford becomes a fundie, then is any gay-related org in violation of the license?

It's his right to make up ridiculous licenses. Like the sisterware license (you can use the software if you send me a pic of your sister if you have one), people shouldn't take it seriously and avoid code licensed like that. That Crockford doesn't get this is either him trolling or being clueless.

Re: “Let's talk about a hypothetical public-facing service”

#44

Earlier quoted context omitted.

You're still going to need n+1 sysadmins to maintain the rest of your infrastructure. Maintaining a storage system like this really isn't much different for the day-to-day operations. So, yes, it's one more system, but you're already going to need those admins anyways.

I don't believe maintaining a high availability system with proper checking is less than 3 hours a day of work, or ~10% of someone's time on duty if we're talking around the clock. So we're talking about $75,000-90,000 a year in salary to maintaining the cluster if you want 24x365 coverage (which Amazon provides), as you'll need 2 people per shift, 3 shifts per day at a minimum to have people in house, even if they'r…

I feel like these comparisons oversell the level of support that actually comes with Amazon. Yes AWS as a whole very rarely goes down, but instances have problems all the time, and who do you call?

When you have your own servers and staff, even if they are just on-call with a pager, you know they are going to work for you on your problem until it is fixed.

In comparison, the sentiment about AWS is this: better build redundancy into the application because at the server layer, you get whatever you get.

Re: “Let's talk about a hypothetical public-facing service”

#45

Earlier quoted context omitted.

You're still going to need n+1 sysadmins to maintain the rest of your infrastructure. Maintaining a storage system like this really isn't much different for the day-to-day operations. So, yes, it's one more system, but you're already going to need those admins anyways.

I don't believe maintaining a high availability system with proper checking is less than 3 hours a day of work, or ~10% of someone's time on duty if we're talking around the clock. So we're talking about $75,000-90,000 a year in salary to maintaining the cluster if you want 24x365 coverage (which Amazon provides), as you'll need 2 people per shift, 3 shifts per day at a minimum to have people in house, even if they'r…

The first post was a back of the envelope estimation, excluding externalities on both sides, to show why someone would want to keep all that "old, obsolete hardware". For Amazon, I didn't put in the costs of pushing data to amazon in terms of API calls used, bandwidth, etc. Additionally, I posted the cost estimate using their slowest storage system with the least amount of flexibility. The cloud doesn't always save you time and money.

Re: “Let's talk about a hypothetical public-facing service”

#46
post #26

Earlier quoted context omitted.

> you can stop buying disks from Amazon every few months This is something that has been puzzling me. Many years ago I purchased 4x 2TB 5900 RPM drives for a 4 bay ReadyNAS (cost about ~$300 for drives plus ReadyNAS). They have been spinning nonstop for ~4 years [1] and haven't had to replace a single one. Not even an increase in errors to signal that the drive is going. Yet - I've worked on a SAN that cost hundreds…

AFR buddy, AFR. AFR is the annual failure rate, its typically 2 - 5% of the population per year. 4 drives you don't have a large enough set to see this in action, just every day you're in danger of losing a drive by a small statistical amount. In the Blekko cluster we have just under 10,000 drives. We have a two 20 drive 'boxes' (40 drives) from Western Digital, as drives fail we pull replacements from the 'new/refur…

> just every day you're in danger of losing a drive by a small statistical amount.

Which is why for important data I always use some sort of RAID (or cloud syncing). If I lose a drive I won't lose all my data (presumably though if I bought both drives at the same time there is a chance that both could fail at the same or close to the same time).

> That said, if you're running your ReadyNAS with raid 10 (mirrored drives in a RAID 0 config)

ReadyNAS and other products use a special type of RAID that is actually kind of clever (that will use all space regardless of disk size). I know many would criticize the special RAID but I've had more issues with Linux software RAID than I have had with ReadyNAS (nothing related to data loss). But suffice it to say if 1 drive fails then in theory the data should still be ok (at least that is what their claim to fame is).

My personal opinion - I feel like the ReadyNAS will die before the drives. I'm not saying they couldn't or won't die - but I feel like that would be highly unlikely and even more unlikely for more than 1 to fail at once.

Maybe I'm just paranoid - but when I worked on my last SAN it felt like the company used the cheapest possible drives with a short MTBF for the simple reason that my employer would have to keep using them and keep paying for their warranty service (this wasn't a big name like Dell).

Re: “Let's talk about a hypothetical public-facing service”

#47
post #46

Earlier quoted context omitted.

AFR buddy, AFR. AFR is the annual failure rate, its typically 2 - 5% of the population per year. 4 drives you don't have a large enough set to see this in action, just every day you're in danger of losing a drive by a small statistical amount. In the Blekko cluster we have just under 10,000 drives. We have a two 20 drive 'boxes' (40 drives) from Western Digital, as drives fail we pull replacements from the 'new/refur…

> just every day you're in danger of losing a drive by a small statistical amount. Which is why for important data I always use some sort of RAID (or cloud syncing). If I lose a drive I won't lose all my data (presumably though if I bought both drives at the same time there is a chance that both could fail at the same or close to the same time). > That said, if you're running your ReadyNAS with raid 10 (mirrored driv…

[deleted]

Re: “Let's talk about a hypothetical public-facing service”

#48

Earlier quoted context omitted.

I don't believe maintaining a high availability system with proper checking is less than 3 hours a day of work, or ~10% of someone's time on duty if we're talking around the clock. So we're talking about $75,000-90,000 a year in salary to maintaining the cluster if you want 24x365 coverage (which Amazon provides), as you'll need 2 people per shift, 3 shifts per day at a minimum to have people in house, even if they'r…

The first post was a back of the envelope estimation, excluding externalities on both sides, to show why someone would want to keep all that "old, obsolete hardware". For Amazon, I didn't put in the costs of pushing data to amazon in terms of API calls used, bandwidth, etc. Additionally, I posted the cost estimate using their slowest storage system with the least amount of flexibility. The cloud doesn't always save y…

I'm just saying that you omitted a major component, as no one would argue that the trade-off in engineering time isn't one of the main cost-benefit components of considering AWS (as we can see here, where it weighed in at a substantial fraction of your estimate).

It's like forgetting to count the price of hardware, and only talking about the relevant cost of electricity.

Ed: Accidentally a negative.

Re: “Let's talk about a hypothetical public-facing service”

#49

Earlier quoted context omitted.

I don't believe maintaining a high availability system with proper checking is less than 3 hours a day of work, or ~10% of someone's time on duty if we're talking around the clock. So we're talking about $75,000-90,000 a year in salary to maintaining the cluster if you want 24x365 coverage (which Amazon provides), as you'll need 2 people per shift, 3 shifts per day at a minimum to have people in house, even if they'r…

I feel like these comparisons oversell the level of support that actually comes with Amazon. Yes AWS as a whole very rarely goes down, but instances have problems all the time, and who do you call? When you have your own servers and staff, even if they are just on-call with a pager, you know they are going to work for you on your problem until it is fixed. In comparison, the sentiment about AWS is this: better build…

We are talking about storage. S3 and Glacier have never gone down, and never will.
Post reply on HN