Live data from Hacker News

Introducing Google Cloud Storage Nearline

googlecloudplatform.blogspot.com

171–180 of 188 posts

Re: Introducing Google Cloud Storage Nearline

#171
post #32

I added GCS nearline to my object storage comparison: http://gaul.org/object-store-comparison/

you are missing runabove edit: also missing constant.com

I have not previously encountered these providers; could you submit a pull request here:

https://github.com/andrewgaul/object-store-comparison

Re: Introducing Google Cloud Storage Nearline

#172
post #73

Quote: "This is a Beta release of Nearline Storage. This feature is not covered by any SLA or deprecation policy and may be subject to backward-incompatible changes." So should I believe in Google's good-will? I would be fine trying out some services, which are in Google Beta. But my valuable data? They should have a SLA right from the start to gain the user's trust.

If your data really is that valuable, any compensation promised by an SLA is likely to be meaningless.

Far better to use multiple redundant solutions, and although Nearline is only beta* it offers us an easy and cheap way to increase storage diversity.

*How long did we all rely on Gmail while it was beta?

Re: Introducing Google Cloud Storage Nearline

#173

Earlier quoted context omitted.

No, one simply should wait for companies to offer services with guarantees for one's own data, if the data is important. Just that.

If they lose your data, yes, that's one thing. But if the 3 seconds becomes 6 seconds... Or the price goes up... Or they announce they're end-of-life'ing the product... Just move your data somewhere else. Sure, it'd be inconvenient (and maybe expensive) to move. So, you balance all of that out in your mind, and maybe this is the right service for you, and maybe it's not.

I'm not even sure it'd be "expensive to move". Are people _really_ considering using this (or Glacier/rsync.net/whatever) as their _only_ copy of their data? I can't imagine looking my boss/customers in the eye and saying "We're going to have multiple terabytes of mission critical business data, and it's all going to live _only_ on AWS/Google/Cloud-service-de-jour!"

If I lose my AWS Glacier stored data (or Amazon bump the prices intolerably), I'll upload it to a competitor _from my local copies_...

Admittedly, I've only had to deal with storage topping out in the tens of terabytes range, so I've never needed to go beyond a dozen or two consumer-grade drives to keep a pair of rsynced copies locally - but I think that same kind of techniques scale all the way out to building your own Backblaze style storage pod if needed.

Re: Introducing Google Cloud Storage Nearline

#174
post #138
post #133

Earlier quoted context omitted.

arq (which we love) + rsync.net = success. We have a "HN readers" discount which is fairly substantial.

= success rsync.net is 20x more expensive than Google Nearline ($0.20/GB). Why would anyone choose to use it for Arq when Arq supports both?

I encourage you to email us and ask about the HN readers discount.

Further, Amazon S3 is a better comparison, as far as pricing goes - we're fully live, online, random access storage - not nearline or weirdline like google/glacier.

Re: Introducing Google Cloud Storage Nearline

#175
post #130

Earlier quoted context omitted.

Unless Glacier just changed their pricing, how is it cheaper and simpler? Glacier is also $.01/GB storage, but only $.09/GB retrieval, which is cheaper than google's $.12/GB. Glacier also comes with 1st GB retrieval free/mo.

"Unless Glacier just changed their pricing, how is it cheaper and simpler? Glacier is also $.01/GB storage, but only $.09/GB retrieval, which is cheaper than google's $.12/GB. Glacier also comes with 1st GB retrieval free/mo." As long as we are comparing effective pricing, which is the only number that matters with Glacier and "Google Nearline" (and, to some degree, S3) it should be noted that rsync.net PB-scale is 3…

Not to nitpick, but your CEO page is a bit misleading.

    > "a two year contract is required." [1]
    > "There are no contracts, overages, fees, or license charges at rsync.net." [2]
[1] http://www.rsync.net/products/petabyte.html

[2] http://rsync.net/products/ceopage.html

Re: Introducing Google Cloud Storage Nearline

#176
post #81
post #73

Quote: "This is a Beta release of Nearline Storage. This feature is not covered by any SLA or deprecation policy and may be subject to backward-incompatible changes." So should I believe in Google's good-will? I would be fine trying out some services, which are in Google Beta. But my valuable data? They should have a SLA right from the start to gain the user's trust.

If we place too many restrictions on how companies should offer preview releases, they'll just stop entirely. Are you suggesting that even those comfortable taking on risk to get a sneak peak should be forced to wait for general availability?

think of the poor poor companies!

Re: Introducing Google Cloud Storage Nearline

#177
post #159

Earlier quoted context omitted.

Why does Amazon charge like this? It seems like on the rare occasion that you need to send some person to go grab the tape/disk from storage and bring it online, Amazon would want you to get all the data you need and put it back in storage. Incentivizing users to bring the data online once a day to trickle it out seems bad for all involved.

I think it makes most sense to think of it as paying for access to the tape robot. If Amazon store your data in a tape archive (I don't know if they do, but they at least seem to have similar constraints), they can only access a small portion of the stored data at a time, so they need to control how often people request data. They could just rate limit everyone, but this way allows people to pay for priority in an em…

Note: The following is speculation best I know. I'm recalling from memory what I've read on the internet written by someone who did not have a direct source.

Glacier uses low-speed (5400 RPM) consumer drives, which they then clock down lower to save energy. Any given 'rack' only has enough power to power a few drives on that rack, the rest is powered down.

To prevent multiple customers from all trying to pull their data out they needed to introduce a rate limiting system, which they did with this exorbitant pricing.

Re: Introducing Google Cloud Storage Nearline

#179

Earlier quoted context omitted.

If they lose your data, yes, that's one thing. But if the 3 seconds becomes 6 seconds... Or the price goes up... Or they announce they're end-of-life'ing the product... Just move your data somewhere else. Sure, it'd be inconvenient (and maybe expensive) to move. So, you balance all of that out in your mind, and maybe this is the right service for you, and maybe it's not.

I'm not even sure it'd be "expensive to move". Are people _really_ considering using this (or Glacier/rsync.net/whatever) as their _only_ copy of their data? I can't imagine looking my boss/customers in the eye and saying "We're going to have multiple terabytes of mission critical business data, and it's all going to live _only_ on AWS/Google/Cloud-service-de-jour!" If I lose my AWS Glacier stored data (or Amazon bum…

In a word: yes.

You're thinking about data that can't be lost, or your customers are screwed. Not all data is like that.

Log files come to mind. They're _nice_ to archive for a long time, but in many businesses, they're certainly not _critical_ to archive for a long time.

Intermediate files, too. You retain the original files in secured storage. But because the intermediate files are large and expensive to re-create, you keep them here in AWS.

Re: Introducing Google Cloud Storage Nearline

#180
post #4

from the documentation: "You should expect 4 MB/s of throughput per TB of data stored as Nearline Storage. This throughput scales linearly with increased storage consumption."

That would mean that if you had about 1TB stored it would take more than 3 days to retrieve it (with an initial 3 second delay before it starts).

That's a good observation. I wonder how it works for multiple objects that are significantly smaller than 1TB. If I'm streaming one at 4MBps and then after a bit, decide to start another object download, does the original slow to 2MBps?
Post reply on HN