Live data from Hacker News

Data exfiltration in Keepa Price Tracker

palant.info

21–30 of 38 posts

Re: Data exfiltration in Keepa Price Tracker

#21
post #9

From the Keepa addon settings: > Allow the add-on to gather Amazon prices to improve our price data I thought it was common knowledge that Keepa uses the addon to gather prices. Though with GDPR it probably needs to be more explicitly said.

There is a difference between gathering prices and loading extra URLs to gather those prices. From that text, I would not assume they are using my computer as a part of a botnet.

I knew it was doing that as well (the distributed scraping the article talks about). But I cannot figure out where I read it. Maybe they used to have it somewhere on their site, and now it's gone?

What is strange that people asked for e.g. Amazon.nl support. This isn't implemented as Keepa relies on Amazon (this is their answer in the forums). But if they scrape, why do they still need Amazon?

Re: Data exfiltration in Keepa Price Tracker

#22
If the additional Amazon pages are loaded on days when the user hasn’t browsed Amazon, or done once a day, that could be cookie stuffing, explicitly prohibited by Amazon Affiliate terms. The Amazon affiliate cookies last 24 hours, so triggering a session when a user doesn’t do it, might extent their affiliate window and is not right at all.

Re: Data exfiltration in Keepa Price Tracker

#23
post #4

And not sure if Amazon would agree to this as it essentially threatens the privacy and integrity of their users. Interestingly, Keepa is also an Amazon Affiliate, so they are in a direct business relationship with Amazon.

As far as I know, Keepa is not an Amazon affiliate. They used to be and got kicked out like many similar tools around 5 years ago.

They moved to the current model of providing an API for Amazon data (which seems to use the extensions users to scrape data).

Re: Data exfiltration in Keepa Price Tracker

#24
post #20

> Unless of course you don’t consider the information collected here personal. I don’t. The author even goes out of their way to point out that these requests aren’t generated by the user and so there’s no latent interest information there. I agree that they should cover this behavior in the privacy policy explicitly, but there’s a tone of moral outrage in this piece that seems unearned.

Note: I am the author of this article.

I’m really unsure how you would come to this conclusion. Even if you only read the summary at the beginning or only the conclusions section at the end, you should notice that Keepa is doing both. It will extract data from your Amazon visits (personal information) and do its own scraping (merely wasting your bandwidth if implemented correctly which I am unconvinced of).

Re: Data exfiltration in Keepa Price Tracker

#25
post #22

If the additional Amazon pages are loaded on days when the user hasn’t browsed Amazon, or done once a day, that could be cookie stuffing, explicitly prohibited by Amazon Affiliate terms. The Amazon affiliate cookies last 24 hours, so triggering a session when a user doesn’t do it, might extent their affiliate window and is not right at all.

Keepa is a data company though, not an Amazon Affiliate, so they shouldn't care about violating that policy

Re: Data exfiltration in Keepa Price Tracker

#26
post #24
post #20

> Unless of course you don’t consider the information collected here personal. I don’t. The author even goes out of their way to point out that these requests aren’t generated by the user and so there’s no latent interest information there. I agree that they should cover this behavior in the privacy policy explicitly, but there’s a tone of moral outrage in this piece that seems unearned.

Note : I am the author of this article. I’m really unsure how you would come to this conclusion. Even if you only read the summary at the beginning or only the conclusions section at the end, you should notice that Keepa is doing both. It will extract data from your Amazon visits (personal information) and do its own scraping (merely wasting your bandwidth if implemented correctly which I am unconvinced of).

Thanks for engaging here. Maybe my reading comprehension is poor, but here’s the full quote that I was objecting to. It comes after a long pull quote where Keepa promises to not log the requests that do contain latent interest behavior:

> This refers to some pieces of the Keepa functionality but it once again completely omits the data collection outlined here. It’s reassuring to know that they don’t log product identifiers when showing product history, but they don’t need to if on another channel their extension sends far more detailed data to the server. This makes the first sentence, formatted as bold text, a clear lie. Unless of course you don’t consider the information collected here personal. I’m not a lawyer, maybe in the legal sense it isn’t.

When I was reading, I thought that “data collection outlined here” referred to the scraping behavior you reverse engineered, since the pull quote covered the user-generated request. I agree that they should include the additional scraping behavior here for clarity (we’re arguing about it after all). I disagree that it constitutes as a “clear lie”, since I don’t think that data is personal.

Re: Data exfiltration in Keepa Price Tracker

#27

Wow, they've built a distributed Amazon listing scraping system – essentially a botnet. As someone who has done a lot of web scraping and had to route around a lot of blocking (we have business contracts to allow scraping, but they don't stop over-eager sysadmins), this feels like a dream come true. But I'd never actually want to use this for scraping and I'm not sure any informed user would agree to use this.

How do you get contracts to allow scraping? What kind of cost are we talking about?

Some companies want you to list their products on your page (usually with some kind of affiliate deal attached) but don't have a tech team to implement a feed or an API. In that case you end up in a situation where you have to scrape the data yourself with permission.

Re: Data exfiltration in Keepa Price Tracker

#29

I use Keepa basic and it has saved me a ton of money. I always just assumed it was scraping the prices from pages I visit, but I didn't know it would automatically fetch Amazon pages in the background. Might just sign out of Amazon, and use a separate browser to purchase from it. Either way, I have some thinking to do on if I should "keepa" it or not (sorry really bad joke). Maybe I should purposely turn a blind eye…

Isn't this always the trade-off? While I do appreciate useful software, it gets tiring that it's almost always at the expense of a little bit of privacy or tracking. Seems like the death of a thousand cuts of our anonymity online. Although, I don't really harbor illusions that we (at least Americans) haven't been tracked since the invention of the credit card. I guess I'm a little jaded at this point as there doesn't seem to be anything I, personally, can do about it and I get a touch of FOMO when I hear about the capabilities of the latest and greatest apps. I understand that data collection is inherently necessary for AI, I just don't like who's in charge of it and making the innovations.
Post reply on HN