I must not interact with big enough data to understand, but what is the point of this? I have watched/read several of the Snow product descriptions from AWS and I am still not quite sure I understand the point. It seems like a way to sync local big data to AWS cloud storage for data that is too large to realistically transfer via the internet. So is this simply a sneakernet external harddrive because physically shipp…
For our use case it was genome sequencing. Each individual genome was a terabyte or so depending on what types of data we sent. We can crank out a few hundred genomes per week at full production so when our collaborators need to receive their data they usually neither have the capacity in bandwidth or storage for cohort of thousands. So it gets shipped to AWS.
We've made improvements to get more direct connections to cloud services so we don't use them anymore but for awhile we were filling a snowball and shipping it every week or two. We also used the Google Transfer Appliance once or twice (480tb) to transfer projects all at once instead of piecemeal.