Live data from Hacker News

Fuck the Cloud (2009)

ascii.textfiles.com

201–210 of 235 posts

Re: Fuck the Cloud (2009)

#201
Salutations, ass-end of the Tech Elite.

As someone who has generated a pretty hefty sandbag of verbiage over my decades online, it's always amusing to see what the Grand Eye of internet arbitration decides is an incredibly important and pertinent subject to discuss in my back catalog. Whether it's my work in guiding volunteers for in-browser emulation (http://archive.org/details/softwarelibrary), my delightful coterie of 1980s BBS textfiles (http://www.textfiles.com) or perhaps my documentaries on BBS culture (http://www.bbsdocumentary.com) Text Adventures (http://www.getlamp.com) or the DEFCON Hacker Conference (https://www.youtube.com/watch?v=rVwaIe6CiHw) ... or, as it is today, one of my many long-form written-down thoughts on all manner of this silly medium many of us have chosen to live our lives.

Oh yes, also that my cat is on twitter and has a million followers. (http://www.twitter.com/sockington) - Lots of people are loaded with knapsacks of opinion about that one as well.

I have found that Hacker News (which is, be clear, an unexpectedly lively extension of Y Combinator) is composed of several diverse groups, all with variant approaches to a linked subject. A linked subject which, as some have pointed out, I wrote 6 years ago, deep in the mists of time.

One group is literally in it for the Money, the gain, the ROI, the endless quest for the "Unicorn", and all their commentary is pungent with the bias and filter of either finding the precious gold coin at the bottom of the shitpile, or are rife with attempts to promote or play up subjects and links of great interest to their financial agenda. Be assured that I could not care less about the current status of the beating of your heart.

Another group seems to be happy to drill down as deep as they can into the mathematics, algorithms, and code of a situation, thinking that if they napkin-blart out enough "facts", they will win some sort of day. I find these people tend to be unhappy about flowery language or effusive phrasing, simply because they've left-brain-dominated themselves into deep pits of nut-sorting and bolt-counting. They use "TL;DR" a lot, as well as, I assume, Adderall. Their heartbeat status is of greater interest to me, if only because I think they are coming from a good place, even if that place smells of Cheetos and sweat.

And, of course, there are Opinion Tourists, my favorite, who might as well be equated with a loud and cantankerous pit of waving hands, waiting for the newly linked (if not newly written) event/opinion/image for their to raise in a mighty roar with a hastily cooked "hot take" on the item. Some of them even optimize the process to not even click on the provided link before the horn honking ensues.

So, "Fuck the Cloud" was written in the deep miasma of when everyone used the term "Cloud" interchangeably with "Magic"; that it was an approach and glory that would lead the experience of computing to a new shangrila. Like any old-timers rife with memories of how we got into that world (and of the echoes of cloud-dom going back 50 years), I decided to write out some of my own thoughts, especially on this attempt to dumb down the populace and separate them from not just responsibility, but control and agency with their data. I have been entirely correct in the general theme - there is a divide within the technical community, of people with admin access and the ability to control any aspect of their work, and then a very large, almost overwhelming set of users who are, essentially, meat stock. And in the same way that meat stock has no particular seat at the table when negotiations of an agricultural nature are conducted, so in the same way are the "users" left out in the cold as a whole range of abilities and ersatz "rights" are stripped away, under the guise of "ease of use" and "leave it to us".

All of this was written without the revelations of the deep, intense surveillance apparatus that is now in place, ensuring that any of this data you control or thought was within your own private space is actually destined to meet you again in an investigation, a courtroom, a warrantless intrusion or a physical SWAT attack. That wasn't even the point.

The point was that user data, treated as something to abuse, monetize, and ultimately discard as a whim, was a complete betrayal of the early promises and experimentation of the Internet. To counteract this trend, I co-founded Archive Team (http://www.archiveteam.org) and our delightful success in many areas would warrant a completely different essay itself - and it has, along with myriad speeches and presentations in the years hence.

I'm sure it might be delightful entertainment for Hacker News to find this or that out on the net and go off, endlessly, in the loop of "This Needs Me" and "Fuck You For Thinking That", but ultimately, these are ridiculous showboat-dances of "what if" and "why not", and I've discovered in the years hence that truly, actions and achievement s speak louder, ever so louder, than words.

Enjoy your day.

And fuck "The Cloud".

Re: Fuck the Cloud (2009)

#202
post #182

Earlier quoted context omitted.

> PS: A home loan might seem like the bank owns your house, but they can't say no when you sell it. They can, unless you satisfy the loan by paying it off as part of the process, at which point they no longer have an ownership-like interest. In the unusual cases where you try to sell a house without doing that, the bank absolutely can -- and often will -- say no.

If you have the cash to pay them off, then they don't get to say no even if it's a short sale. Alternatively, you can generally walk away and sell it to them for the value of the loan.

> If you have the cash to pay them off, then they don't get to say no even if it's a short sale.

If you actually pay them off, then they don't get to say no, because paying them off is, essentially, buying out their interest in the property, under the terms of an existing contract. That doesn't negate the fact that they have legally-enforceable rights in the property until and unless you do that.

> Alternatively, you can generally walk away and sell it to them for the value of the loan.

Only if the mortgage is governed by the law of a jurisdiction where pursuit of mortgage deficiency isn't allowed (either in general or for mortgages in the specific conditions yours has.)

Re: Fuck the Cloud (2009)

#203
post #8

While Jason Scott raises interesting points six years ago, the principles of data management remain the same as the days of dedicated servers with on premises systems. If you have only one copy of data and that machine goes down, then the data goes with it. In my experience in a business context 'the cloud' discussion is about the business' desire to expand capacity without the need for the capital expenditure needed…

(Tedious disclaimer: not speaking for anybody else, my opinion only, etc. I'm an SRE at Google.)

The key piece that's missing here is the idea that risk is something you have to compare, and can combine in interesting ways, then trade off against costs.

There are a bunch of ways in which you can do compute, storage, and networking. You get to pick zero or more of these ways. One of them is "buy a bunch of iron and make a pile of it in your bedroom". Major risk factors here are your house burning down, you getting evicted, or there being a power cut. Another is "rent those services from an infrastructure provider". Risk factors here are much harder for you to visualise, but include things like "governments ban that company from operating in your country".

You can look at the risk of any of these options, and quantify it with an SLO, like "we intend for this compute resource to be available 99.99% of the time in a given quarter". You can then have an SLA that defines what will happen if that objective is not met, and measure how often this is complied with over time. There are lots of ways to analyse this information, but let's suppose that you can reduce it to a single number measuring how safe the resource is for your use case.

If you only look at a single option, and say "this has a safety of X", then the only thing you can get out of this effort is anxiety. This only becomes interesting when you start looking at differences between alternatives, like "the safety of servers in my bedroom is X, but the safety of buying resources on GCE is Y, so I can get this much of an improvement by spending that amount of money", or "by doing both of these things I improve my safety to Z, and I am willing to pay the additional cost of doing so". Or perhaps your position would be "this option is less safe but much cheaper and I'm willing to accept the extra risk".

The problem I have with the "fuck the cloud" article is that it doesn't do any of this. All it says is "the safety of this option is only X, you should experience anxiety". Is X higher or lower than that pile of iron in your bedroom? You still don't know.

(Realistically, unless you have the ability to build a system in your bedroom that has continental diversity for storage, N+2 of everything for hardware failure, etc, your bedroom is likely to be far less safe than the major cloud services - unless you live in a country which regularly bans American companies from doing business with you, which a sixth of the world's population does.)

Re: Fuck the Cloud (2009)

#204
post #197

Earlier quoted context omitted.

How will a contract save your ass if your data is lost? You can't outsource your responsibility.

Unless you're personally memorizing all the data then you're outsourcing the responsibility somewhere - either you're paying for a service that you hope will be compliant with its specification, or you're paying for hardware that you hope will be compliant with its specification and employees that you hope will be as skilled as they claim to be. There's no way to eliminate the risk entirely, all you can do is take th…

> These days that's often the cloud.

For plenty of use cases I actually agree with that conclusion. With the caveat that I would add: accompanied by a suitable non-cloud back-up of critical data, code and configuration data held on a medium that is not accessible by the same people that administer the cloud stuff.

Re: Fuck the Cloud (2009)

#205
post #8

While Jason Scott raises interesting points six years ago, the principles of data management remain the same as the days of dedicated servers with on premises systems. If you have only one copy of data and that machine goes down, then the data goes with it. In my experience in a business context 'the cloud' discussion is about the business' desire to expand capacity without the need for the capital expenditure needed…

That's what the author is getting at, "The Cloud" is sold as this awesome thing that will never die or break, and worse yet it's sold to people who often don't know any better. No one realizes that the Cloud is running on the same crap we've always had and is vulnerable to the same issues as everything else. MAYBE the company is better at data management, MAYBE the employees take pride in their job and do it properly…

Eh... I think that the maybes you are using are a little bit misleading. When a company sells space in their "cloud", like maybe Microsoft or Amazon, there is a business guarantee that goes along with it. If Amazon were to randomly lose a big chuck of Netflix data, AWS's business would tank immediately. AWS is a giant system that uses its scale and number of customers to efficiently provide a more stable system at a lower cost than all of those individual customers could achieve by building and maintaining their own IT.

I think it is sort of like a delivery system. If USPS or FedEx or USP started losing a massive number of packages (I know that they do lose some) then they would get abandoned, just like "Cloud" companies have an incentive to maintain a baseline level of quality. The alternative is that every business would have to create their own shipping services. I think it makes sense to assume that in most cases, unless the business is already massive enough to warrant it, that it is cheaper and more reliable to use the aggregate, dedicated ones for hire. The "Cloud" will be cheaper than individual implementations, and it won't be nearly as suspect to individual implementation errors because the identical system will have been proven by many other customers (otherwise it would be abandoned).

Re: Fuck the Cloud (2009)

#206

Salutations, ass-end of the Tech Elite. As someone who has generated a pretty hefty sandbag of verbiage over my decades online, it's always amusing to see what the Grand Eye of internet arbitration decides is an incredibly important and pertinent subject to discuss in my back catalog. Whether it's my work in guiding volunteers for in-browser emulation ( http://archive.org/details/softwarelibrary ), my delightful cote…

I could have done without your insulting tone, no matter how proud of your opinions you are.

Re: Fuck the Cloud (2009)

#207
post #197

Earlier quoted context omitted.

Unless you're personally memorizing all the data then you're outsourcing the responsibility somewhere - either you're paying for a service that you hope will be compliant with its specification, or you're paying for hardware that you hope will be compliant with its specification and employees that you hope will be as skilled as they claim to be. There's no way to eliminate the risk entirely, all you can do is take th…

> These days that's often the cloud. For plenty of use cases I actually agree with that conclusion. With the caveat that I would add: accompanied by a suitable non-cloud back-up of critical data, code and configuration data held on a medium that is not accessible by the same people that administer the cloud stuff.

I think for most levels of safety you would target, a pure cloud solution is going to be cheaper. E.g. if you decide you need a backup that's independent of your primary cloud provider and any staff with access to that, you can probably accomplish that more cheaply with a second cloud provider than by self-hosting.

Re: Fuck the Cloud (2009)

#208

Earlier quoted context omitted.

That's what the author is getting at, "The Cloud" is sold as this awesome thing that will never die or break, and worse yet it's sold to people who often don't know any better. No one realizes that the Cloud is running on the same crap we've always had and is vulnerable to the same issues as everything else. MAYBE the company is better at data management, MAYBE the employees take pride in their job and do it properly…

Eh... I think that the maybes you are using are a little bit misleading. When a company sells space in their "cloud", like maybe Microsoft or Amazon, there is a business guarantee that goes along with it. If Amazon were to randomly lose a big chuck of Netflix data, AWS's business would tank immediately. AWS is a giant system that uses its scale and number of customers to efficiently provide a more stable system at a…

> If Amazon were to randomly lose a big chuck of Netflix data, AWS's business would tank immediately.

http://www.theregister.co.uk/2015/09/20/aws_database_outage/

AWS is the new Microsoft / IBM, nobody ever got fired for picking AWS.

Re: Fuck the Cloud (2009)

#209

Earlier quoted context omitted.

That's what the author is getting at, "The Cloud" is sold as this awesome thing that will never die or break, and worse yet it's sold to people who often don't know any better. No one realizes that the Cloud is running on the same crap we've always had and is vulnerable to the same issues as everything else. MAYBE the company is better at data management, MAYBE the employees take pride in their job and do it properly…

Eh... I think that the maybes you are using are a little bit misleading. When a company sells space in their "cloud", like maybe Microsoft or Amazon, there is a business guarantee that goes along with it. If Amazon were to randomly lose a big chuck of Netflix data, AWS's business would tank immediately. AWS is a giant system that uses its scale and number of customers to efficiently provide a more stable system at a…

That must be why Amazon's SLA is defined as follows[1] :

If amazon loses more than 3 datacenters (only total loss of external connectivity for all of your instances in an entire availability zone, or total loss of hard disk access, again only counts if all your instances completely lose hard disk/EBS access) for more than 45 minutes in a month you get 10% of what you pay as a voucher for future ec2 usage. If they lose it for more than 7 hours you get 30%.

So no, Amazon, or at least their legal department, does not trust their own competency. Or at least, they're not willing to risk any revenue on that, but they're willing to give you a small future discount to encourage you to restart using the service. Oh and you only get that if you explicitly ask for it.

If they lose your data on EBS/S3/Dynamo/..., you get nothing. So having any data exclusively on any Amazon service should be cause for getting fired, and this of course also means that using Dynamo for storing anything non-trivial is a big no-no from a disaster recovery standpoint.

So I have to say, I would suggest you do not trust Amazon with either your data, nor with keeping your site online. Yes, historically their performance has been better than this, but ...

This reads worse than the SLAs on internet connectivity from places like level3 and cogent (pay 10% less if they fuck up completely for more than 2 days).

[1] https://aws.amazon.com/ec2/sla/

Re: Fuck the Cloud (2009)

#210
post #181

Earlier quoted context omitted.

And so, ideally, I want to be able to tell my accountant to store my archival tax records in my safe-deposit box, not in his office. Compute infrastructure is a commodity—doesn't matter who's paying for it—but all the services you depend on should rely on storage infrastructure you have an SLA agreement with. I don't care about the "distributed computation" promises of Diaspora or Sandstorm.io; I think they're wrongh…

If I'm a Facebook engineer, I would never agree to this because there is no possible way to optimize performance in this scenario. What happens to Facebook when your data provider goes down, or just gets slow? What if they mess up permissions or change their API? Maybe you're thinking "that's fine, if my provider isn't reliable, my Facebook account becomes unavailable and it's up to me to choose a better provider." B…

Think of the Datomic architecture[1]: some nodes are "storage", while other nodes are "transactors." The transactor nodes pull "chunks" of rows/objects/documents from storage nodes, compute relational indexes locally, answer queries from those computed indexes, and finally persist computed index "chunks" back to the storage nodes.

Now, note that the "canonical" storage that gets read from doesn't have to be the same storage that the indexes get written to. The first can be owned by the user, while the second can be owned by Facebook.

Presuming an architecture like this, the latency and availability of the user's "Facebook account" database is relatively immaterial. While writes would have to be synchronous (so, like you said, Facebook would have to give the user a "sorry, your account is unavailable" error), reads could be asynchronous. Think of an online RSS feed reader service: the "primary sources" are the third-party sites with their RSS feeds. Sometimes those sites go down. When they do, the reader-service can't retrieve the feed–so that feed just goes stale.

Things like Facebook's graph, meanwhile, are fundamentally indexes. The base-level "documents" in the graph are relationship assertions—a copy of "B accepted A's friend request" stored in B's database. The graph is a computed value built on a pile of those. When Facebook can't reach someone's database, things like these relationship assertions just go stale.

The crucial idea, here, is that for Facebook to do its job, it probably has to cache a majority of the stuff in the user's database in one form or another—just like an RSS reader caches RSS feeds. But this is purely a cache, in a fundamental sense. Users who don't "check in" with Facebook could be cache-evicted from its database. Other users would still have relationships with them and be able to post on their wall and such (they'd be putting those documents in their own outbox); Facebook would just no longer bother computing anything that's personal to the user, like their news feed. There would be every incentive to set up the architecture such that user data that wasn't needed would be "garbage collected" off of Facebook's servers, because it could always get put back on, the moment that user's account-instance woke up again and said hello.

This would also mean that Facebook wouldn't need to store any of their user-data in anything resembling a relational normal form. Every table would be a "view" table. The canonical database, owned by the user, could be relational and full of nice constraints and triggers (and the user could even add these themselves); but since Facebook can just query out of that to get any data it's missing, it wouldn't need anything like a "users" table. (Fascinatingly, if Facebook was built as a microservice architecture, this means that each microservice would probably separately query the canonical data from the database in order to generate its own indexes; the Search service would know one "face" of you while the Photos service would know quite another. These could even—in theory—be separately ACLed within your own DB instance, giving the user true, actual control over what Facebook can do with their data, component by component.)

[1] http://docs.datomic.com/architecture.html

Post reply on HN