Live data from Hacker News

Getting my personal data from Amazon was weeks of confusion and tedium

theintercept.com

71–80 of 193 posts

Re: Getting my personal data from Amazon was weeks of confusion and tedium

#71
post #68
post #60

Earlier quoted context omitted.

What data does Amazon have that you haven't given them?

I think you misunderstand the comment the other commenter made - there is a lot of info Amazon has about one that is collected via dark patterns. Also, Don't they also buy data from 3rd parties to augment what you give them? Like stats of credit card purchases and stuff? Always assuming that all these big players do that.

>Also, Don't they also buy data from 3rd parties to augment what you give them? Like stats of credit card purchases and stuff? Always assuming that all these big players do that.

They do! That's even mentioned in this article.

Re: Getting my personal data from Amazon was weeks of confusion and tedium

#72
post #30

Earlier quoted context omitted.

> Distributed systems are hard. It takes time to determine where all possible information could live. And you have to make sure you're providing the correct information. And do this flawlessly, every single time lest you open yourself up to bad press and potential fines. This all takes time in systems as large and distributed as Amazon. This implies that Amazon is serving GDPR data requests manually, rather than the…

To create a script/bot/application/whatever that can access all potential data, you have to give something read privileges to possibly hundreds of backend systems and products. This is horrendously bad idea security wise. If that service account gets compromised (either from an external or internal threat), you have a single account that has access to everything Amazon stores. This is bad for the company and bad for…

Sorry, I don't buy any of this.

Automating the process doesn't need to imply that there's a single service with direct access to all of the data. Just from a basic software engineering perspective, it makes a ton of sense each product's data export to be a separate service owned by the product team, so no disagreements there. But by talking about how hard it is to figure out what data you have stored and export it correctly, you were implying that you had no such per-product service either, and each export is an artisanal custom job.

The question of safeguards is interesting. I don't really see how having a human in the loop is adding any real security: a computer is going to be far better at deciding whether the request is valid or not. As an operator, being assigned a ticket to do an export of account 123456, what are you going to do other than do that export? A computer, on the other hand, can actually verify whether the request is actually authorized. That can be done in a way where a compromise of your central data export service account can't be used to fake the authorization.

(A quick design sketch for one option: each account has a public key encryption keypair, managed by the identity system. When the central data export service requests an email verification, that is done via asking the identity system to sign a ticket. The identity system triggers a flow that asks the user to validate the request, and as part of the flow informs them of just what operation they are validating. User approval of the request signs the ticket with their private key. This ticket is sent to each data export service, which checks that the user id they're exporting has signed the ticket, and that the ticket contents match the request: i.e. same userid, operating is a data export, the data export covers this service. You will need to trust your identity system to not be compromised, but if it is, you're completely screwed anyway.)

> And assuming you could securely create this automated workflow, you'd still need a person manually verifying the end result to ensure that all the data scraped is in fact owned by the person who made the request. Within the past couple of years, there was a news story where someone got a different person's Alexa data after asking Amazon for their own data. That can't happen again.

The odds of a human doing a good job of this kind of validation are basically zero. Either they are following a checklist that a computer could execute more reliably, or they are just randomly poking at some 1 GB data dump trying to find the needle in the haystack.

Re: Getting my personal data from Amazon was weeks of confusion and tedium

#73
post #43

Earlier quoted context omitted.

> they do not have the resources to do that. Good - the aim is for them to not store personal data in the first place, much less build business models that rely upon it. Rather than allowing the population to take on the negative externality of surveillance capitalism, it is absolutely right that the burden must fall on those creating the problem. I don't see this as any different to the complain that small restauran…

You're giving the argument too much credit. It's more akin to a large restaurant arguing that small restaurants could be put out of business by health inspections, so maybe we should hold off on the idea. Rather, keeping a clean kitchen is something they all should be doing anyway from the get go. Any pain for Amazon in Amazon's process is entirely Amazon's fault. If systems are built with the requirement of letting…

> If systems are built with the requirement of letting users export their data, then the additional effort to do so is trivial.

It’s unreasonable, IMO, to think that companies should have had the foresight to see legislation that would happen two decades after the company had already existed and as a result build a system for retrieving user data that has no profit generating potential.

GDPR is good because prior to it there really wasn’t any economic incentive to provide this information.

Re: Getting my personal data from Amazon was weeks of confusion and tedium

#74

I run a part of the data request process at our company. This article is an example where people expect anything technology related to be magic. We have to go through EVERY tech stack we own and look for that person's data. It's amazingly manual and tedious and takes about 6 people about an hour per request. We're working to automate it, but needless to say we try not to broadcast it too broadly. I hate that everyone…

It's revealing how hard this stuff is when Google's Data Liberation Front needed 4 years to release Google Takeout – which I consider to be best-in-class for personal data access.

It is a hard problem, but the GDPR went into effect 3 years and 10 months ago. That date didn't come as a surprise, but was known 6 years ago. Anything newer than that should have taken data requests into account from the design stage. Anything older than that has had ample time to adjust. More than that 4 years you quote for Takeout!

Re: Getting my personal data from Amazon was weeks of confusion and tedium

#75
post #73

Earlier quoted context omitted.

You're giving the argument too much credit. It's more akin to a large restaurant arguing that small restaurants could be put out of business by health inspections, so maybe we should hold off on the idea. Rather, keeping a clean kitchen is something they all should be doing anyway from the get go. Any pain for Amazon in Amazon's process is entirely Amazon's fault. If systems are built with the requirement of letting…

> If systems are built with the requirement of letting users export their data, then the additional effort to do so is trivial. It’s unreasonable, IMO, to think that companies should have had the foresight to see legislation that would happen two decades after the company had already existed and as a result build a system for retrieving user data that has no profit generating potential. GDPR is good because prior to…

You're implying that arbitrary "legislation" just arose out of the blue. Rather, it's based on a long held idea that companies are merely trustees for customers' data. So their position is more akin to having built a shed straddling a property line a decade ago, and now complaining that they couldn't have known that their neighbor might eventually want it moved.

Re: Getting my personal data from Amazon was weeks of confusion and tedium

#76
I've been trying for over two years to get my data from Amazon.

I eventually got to a point where Amazon provided a web-page, which has no less than sixty-two download links on, each of which would have to be manually operated.

It's properly tantamount to obstruction.

After finally reaching this point, Support were arrogant and high-handed - "We will not do any more than we have. We look forward to seeing you on Amazon in the future."

I still do not have my data.

I tried to start the process off a second time, but it went nowhere. I chased it, and then had some very disconnected and confusing responsese from Support (email from some random guy in Support who by the looks of it had been told to email me, but neither he had been told what for, nor I that it would happen).

I've not spent more time on it since then.

I stopped using Amazon about two years ago, because I've come to the view that the stories about how Amazon treats warehouse staff are accurate.

I want to get my personal data, so I can close the account.

Amazon of course refuse point blank (in the usual, slimey, support-talking-past-you way) to delete any personal data, so all you can do is delete the account and hope in the end Amazon expire the data.

Re: Getting my personal data from Amazon was weeks of confusion and tedium

#77
post #21

It's not very often that each and every point in an article just feels "fabricated" or over the top. It starts with finding the page: Amazon -> Customer Service -> Search for "personal data" -> Search result #1 is "Request Your Personal Information" which nicely explains what to do and links directly to that page. The need to verify or activate a data request via clicking a link? Of course required so some third part…

> each and every point in an article just feels "fabricated" or over the top. What I thought were valid points from the article: - Unclear data: "cryptic strings of numbers like '26,444,740,832,600,000” for various search queries." This is easily the worst offender IMO. - A wait time of 19 days - Separating the download into 74 buttons

Yup. I agree. The wait time doesn't make sense. They should be able to spin up extra servers from the spot market in seconds. Even if they're using Glacier, that should only be a few hours.

I wonder if they execute the 74 data queries in serial to drag it out.

And the multiple downloads is just bogus.

That being said, I agree with the general point that the article is a bit overly dramatic. Amazon does a pretty good job with the request. It just takes too long.

Re: Getting my personal data from Amazon was weeks of confusion and tedium

#78
post #64

Earlier quoted context omitted.

Perhaps all of this would be a lot easier if you actually built some simple automation to process requests. What could possibly take 3 days to process? The only plausible reason is that you’re wasting developers time on what really belongs in one of the myriad tools AWS itself provides for such tasks.

There’s nothing that’s simple when you’re dealing with 10s of thousands of different datasets across many different internal team and service boundaries with their own security setup depending on the data that’s being stored. The cost of automating and properly securing it (since “gather all customer data into one place” is generally not great as it’s a single point of failure from a security perspective). All of tha…

My reply was to one person presumably on a pizza team at AWS. Surely they would realize some savings from automating their own retrieval requests.

As others have pointed out aggregating all of the different reports into one download is a trivial task itself suited well for automation.

Re: Getting my personal data from Amazon was weeks of confusion and tedium

#79
People in glass houses shouldn't throw stones. The author may want to read the privacy policy[0] for the site they are publishing their story on. They are collecting all sorts of data that they don't need to. And IANAL but apparently your rights to access the data they hold on you are restricted only to locations where they legally have to allow it.

[0] https://theintercept.com/privacy-policy/

Re: Getting my personal data from Amazon was weeks of confusion and tedium

#80
post #28

I had to click through more than 100 links to download all the data, how can this be acceptable? Specially coming from Amazon. How hard is it for them to create an archive with all the data? This is ridiculous, I can't imagine how was the meeting when they decided to produce purposefully such garbage UX.

Exactly the problem I had.

It would take Amazon almost no effort to make a single archive with all those files in.

I cannot help but view this as deliberate obstruction.

Post reply on HN