Live data from Hacker News

NSA uses Google cookies to pinpoint targets for hacking

washingtonpost.com

141–150 of 178 posts

Re: NSA uses Google cookies to pinpoint targets for hacking

#141
post #7

There are two primary issues here: the prevalence of Google Analytics and the unencrypted nature of the majority of websites. Google Analytics is on a substantial proportion of the Internet. 65% of the top 10k sites, 63.9% of the top 100k, and 50.5% of the top million[1]. My own partial results from a research project I'm doing using Common Crawl estimates approximately 39.7% of the 535 million pages processed so far…

I wish you luck with your research, but I'm not sure your look at Google Analytics is relevant to this discussion or the Washington Post article that started this thread.

First, as pointed out, you "cannot track a unique visitor across the web using GA cookies" because of the way they're designed: https://news.ycombinator.com/reply?id=6889120&whence=item%3f...

Second, the NSA doc as excerpted in the WashPost article talks only about Google's PREF cookie, which is set only when you go to say Google.com, not when you go to a non-Google property. It's a first-party cookie used for things like saving language preferences when you're not logged in, not for advertising across other properites. (That's what the Doubleclick cookie is for.)

Re: NSA uses Google cookies to pinpoint targets for hacking

#142
post #89
post #63

Earlier quoted context omitted.

To answer your reply to my reply: > Given the results, I am quite surprise you would say "look to be first parties serving content for the page". I believe every single domain name you listed (except the Google and Twitter domains, like I said) is a domain owned by or a CDN used by the Guardian or hosts an app run by the Guardian - prove me wrong: > facebook-web-clients.appspot.com > guardian-notifications.appspot.co…

> "prove me wrong" This is a terrible answer: you are suggesting that Disconnect knows exactly which 3rd-party is legit when visiting a web page, and somehow you can vouch that none of these hostnames is a threat to privacy (this is what your defense of this implies). `static-serve.appspot.com` is no different than `ajax.googleapis.com` (you didn't list this one, why ?): they are 3rd-party hostnames, some are CDN whi…

> You are suggesting that Disconnect knows exactly which 3rd-party is legit when visiting a web page.

Yes! You now know how Disconnect works - Disconnect's filter list is based on weekly crawl data that identifies what the most prevalent third parties on the web are.

> `static-serve.appspot.com` is no different than `ajax.googleapis.com` (you didn't list this one, why?)

You think that URL might belong to Google, which I already called an exception 2x?

> I will note that you completely disregarded the other results which are even more embarrassing to explain (like `simplereach.cc`: "SimpleReach tracks every social action on each piece of published content to deliver detailed insights and clear metrics around social behavior.")

I examined and debunked the entirety of the first example on your page, so I'm not inclined to waste any more time on your so-called "science".

Re: NSA uses Google cookies to pinpoint targets for hacking

#143
post #135

Earlier quoted context omitted.

[Replying to aroch, who's too nested.] > And how exactly is it trickery if users have to opt-in to the program and they're told what the program does? Ghostery seems to rely on vague messaging (last I looked, they don't actually say anywhere in their extension that they sell the data you share to ad co's and data brokers) and UX "optimization" (what quesera dubbed the "reconfigure-on-update dance", for example) to ge…

In the second paragraph (though really, its just a statement...) on the preferences page -- no need to navigate to another page, and they tell it to you in plain english. Once again, you have to opt-in, so if you opt-in without knowing what it does it's your own fault and you're being a dumb user: When you enable GhostRank, Ghostery collects anonymous data about the trackers you've encountered and the sites on which…

To be fair, if those recordings are time-stamped, it probably does leak information about user browsing habits.

Anonymizing data is hard.

Re: NSA uses Google cookies to pinpoint targets for hacking

#144
post #86
post #82

Earlier quoted context omitted.

See the commit message at https://github.com/disconnectme/disconnect/commit/691897e21d... . The unencrypted list is at https://github.com/disconnectme/disconnect/blob/b27abbf033c6... . And formatted as JSON at https://disconnect.me/services-plaintext.json . The encrypted list is also trivial to decrypt with the SJCL code in https://github.com/disconnectme/disconnect/blob/master/firef... .

Thanks! I was setting up a proxy for devices that can't use Disconnect or Adblock. I thought of adapting the Disconnect list to the proxy's block list, but a cursory reading of the source only showed the URL of the encrypted list.

Cool, ping me if you need help (byoogle everywhere).

Re: NSA uses Google cookies to pinpoint targets for hacking

#145
post #135

Earlier quoted context omitted.

In the second paragraph (though really, its just a statement...) on the preferences page -- no need to navigate to another page, and they tell it to you in plain english. Once again, you have to opt-in, so if you opt-in without knowing what it does it's your own fault and you're being a dumb user: When you enable GhostRank, Ghostery collects anonymous data about the trackers you've encountered and the sites on which…

To be fair, if those recordings are time-stamped, it probably does leak information about user browsing habits. Anonymizing data is hard.

Assuming they're adhering to privacy standards, the anonymization should be reported in aggregates and not like "{UUID} at {TIMESTAMP} reported {TRACKERS} at {URI}"

Re: NSA uses Google cookies to pinpoint targets for hacking

#146

Earlier quoted context omitted.

What was not blocked: == Disconnect: * s3.amazonaws.com * cloudfront.net * echoenabled.com * troveread.com * trove.com == Ghostery: * s3.amazonaws.com * echoenabled.com * platform.twitter.com == HTTP Switchboard * echoenabled.com For HTTP Switchboard, I could easily identify glancing at the matrix that what was requested was a CSS file, I then proceeded to block with one click anything coming from `echoenabled.com`.…

[You think maybe you should start identifying that you're promoting your own product with these comments?] The first thing I tried when I wrote Disconnect was to block every third-party domain. Within an hour, I realized that I broke the whole web. So I built a crawler to identify and categorize the most prevalent third-party services instead. The domains you list under Disconnect would all be categorized as content…

> "[You think maybe you should start identifying that you're promoting your own product with these comments?]"

Your going personal. I am talking about the extensions, not you. If for-profit companies are going to claim to care about user privacy, expect this claim to be taken to task, especially in the current era.

Above I am providing hard data, not an opinion.

> "The domains you list under Disconnect would all be categorized as content by our crawler"

The page works fine if whitelisting only the page domain. If someone want the comments, then it's a matter of whitelisting `echoenabled.com`. The rest doesn't appear so important, so I personally rather not ping them. But the point is, I am of the opinion that people need to have the ability to know exactly where their browser connects, then they can agree/disagree/not care. I don't see how one can make an informed decision without proper information.

Now regarding:

http://p.typekit.net/p.gif?a=219379&f=175.10294.10295.10296....

There is no way this 1x1 pixel gif would break a web page. And yet it's not blocked by Disconnect as reported in another comment below. I also reported how adobetag.com is reportedly blocked by Disconnect and yet a script from adobetag.com was downloaded by Disconnect.

Can't you appreciate why I am rather skeptical? Going personal rather than provide a credible answer is not going to dispel this skepticism.

Re: NSA uses Google cookies to pinpoint targets for hacking

#147

Earlier quoted context omitted.

Why do we still have referrers? They don't allow us to do anything that we wouldn't be able to do without them. If Mozilla and Google made a statement today saying, "We'll be removing referrers from cross site requests in 6 months time for Chrome and Firefox.", the tiny tiny proportion of sites that are using them for real functionality will have plenty of time to update. Of course, as a web developer, it's useful to…

For lots of us using basic CDN services, we enable referrer checks to ensure that folks aren't hotlinking images or direct linking downloads from other sites. These CDNs allow basic blocking based on referrers. You usually set it to only permit when there is a referrer from your own domain as well as blank referrers (if the CDN supports it) since most privacy conscious folks will disable referrer rather than fake it.…

There is a trivial solution to this. Introduce a new HTTP response header 6 months before phasing out the Referer header. This header would be optionally delivered with content and would specify which third party domains are allowed to access the content. Perhaps Content-Security-Policy could be extended for this purpose.

Re: NSA uses Google cookies to pinpoint targets for hacking

#148

Earlier quoted context omitted.

Why do we still have referrers? They don't allow us to do anything that we wouldn't be able to do without them. If Mozilla and Google made a statement today saying, "We'll be removing referrers from cross site requests in 6 months time for Chrome and Firefox.", the tiny tiny proportion of sites that are using them for real functionality will have plenty of time to update. Of course, as a web developer, it's useful to…

You ask, "Why do we still have referrers?", but then you answer your question: "as a web developer, it's useful to be able to see where people came from." You are of course correct that we don't have a "right" to this information. But I've discovered, many times, through the referrers in my logs, links to my pages from some very interesting places that I might not have discovered otherwise (because the link informati…

Defaults are important. Most users of the web don't know referrers exist.

It's irrelevant how useful you find the information. You'd probably find it useful to know the name and email address of everyone that visits your site too... So?

Re: NSA uses Google cookies to pinpoint targets for hacking

#149
post #105

Earlier quoted context omitted.

Google is already killing referrer on search results (by redirecting), to force people to pay for Analytics

I'd pay for an intermediate level of GA in a heartbeat. Right now, once you hit 10M pageviews a month you either have to sample or pay $150k/year for Premium. I don't need support, an account manager, four-hour turnaround on data, an SLA, etc. I just need more pageviews sometimes.

Hate to burst your bubble, but even the premium GA uses samples. They don't give you a firehose of real data.

Re: NSA uses Google cookies to pinpoint targets for hacking

#150

Earlier quoted context omitted.

It's only hostile to sites that try and steal bandwidth resources by hotlinking/leeching images or direct linking downloads. It's not about milking visitors. It's about preventing unethical behavior by other sites. I've spent 10s of thousands of dollars hosting free and open source software for millions of people over the years and I make sure to prevent bandwidth theft from other sites that cut into my ability to pr…

So it sounds like there's already a solution to your problem that doesn't require leaking privacy all over the internet, since presumably the one-time links are on request from a page on your website, and tells you nothing but someone on your website wanted something from your website. How is this a bad thing? How would removing referrals harm you in any way?

What if the page containing a one-time link was cached, but the resource itself was not? The "secure link" solution doesn't seem to work in all cases.
Post reply on HN