Live data from Hacker News

Mozilla research: Browsing histories are unique enough to identify users

zdnet.com

91–100 of 131 posts

Re: Mozilla research: Browsing histories are unique enough to identify users

#91

Earlier quoted context omitted.

Yes, but the primary identifier is IP address. The detailed profiles built with fingerprinting and other data are attached to the small set of IP addresses a person uses over the course of her lifetime. Most internet users have limited choice when it comes to internet access. A user cannot change her IP address with the same ease as she can change her software fingerprints. If a company is trying to sell online ad se…

Given the prevalence of CGN, especially in mobile / cellular internet, and the reality that mobile is first for a large number, the use of IPs as a primary key feels less likely these days than a decade ago. > A user cannot change her IP address with the same ease as she can change her software fingerprints. I dunno. It’s a lot easier for my less techie friends to reboot their router and get a new IP than it is to ta…

How static/deterministic are the CGNAT translations though? It is conceivable that when client A connects to a Facebook service with IP X and port P that the source IP and port observed by Facebook is always the same.

In any case your ISP is probably logging all your DNS queries and all their dynamic NAT translations to a database, so couple REMOTE_ADDR with REMOTE_PORT and a timestamp and you can almost certainly be identified.

Re: Mozilla research: Browsing histories are unique enough to identify users

#92
post #41

To this me and a friend started sketching on a VPN/HTTP proxy that will have a set of say 100 outgoing IPs, look at the domains being connected to and distribute request destinations over IPs. So e.g. Google would always see the same IP, which would be different from the one Facebook sees. While access times cross-references and identification is still theoretically possible, it should be an entirely different game.…

This sounds like a less secure Tor to me.

Re: Mozilla research: Browsing histories are unique enough to identify users

#93

Earlier quoted context omitted.

Isn’t a postal code about 50-100 houses? It really narrows things down. That particular variable really reduces things.

On average it's about 15 houses, so postcode + date of birth is indeed around 95% accurate.

[deleted]

Re: Mozilla research: Browsing histories are unique enough to identify users

#94

Here in the UK, date of birth and post code is enough to identify something like 95% of people. Anonymised data sets are not really possible once you have more than a few varriables. Most people don't know this.

Isn’t a postal code about 50-100 houses? It really narrows things down. That particular variable really reduces things.

There are 2 reasons postcode matters.

First, postcode is something you give out pretty willingly. If you put your postcode and dob into an insurance quote website, they would no longer be insuring based on a pool of people like you. They'd literally just see how many claims you had. And also what ethnicity and sexuality and 50 other personal, irrelevant criteria they want.

The second is that postcode is only a narrow or broad measure depending on what you're using it for. If you want to do a study on asthma rates vs road traffic, postcode is just right, anything more general and you're comparing side streets and motorways. So it makes sense for that data to be available. But wait, as the data user, I only need one more data set (say voter registration, already available) and I can literally look up you're medical history before deciding whether to hire you.

This is the issue here: data HAS to be specific to be useful. But ifs its specific its dangerous. AND data is much more specific than you realise because a few innocent sounding data points are unique to you when combined.

Re: Mozilla research: Browsing histories are unique enough to identify users

#95

Here in the UK, date of birth and post code is enough to identify something like 95% of people. Anonymised data sets are not really possible once you have more than a few varriables. Most people don't know this.

Isn’t a postal code about 50-100 houses? It really narrows things down. That particular variable really reduces things.

When you consider the birthday paradox, and consider demographics, it probably doesn't.

If you assume that DOB's are evenly and randomly distributed over the last 100 years (1 of ~36500 values) then the probability of none of 100 people sharing a DOB is only ~87%. If you tuned for demographics the true stats would be much worse.

That said, you probably have anti-clustering aspects - parents obviously can't share birth dates with their children, and siblings can't either (unless twins). But! couples tend to be of a similar age...so, tricky.

Re: Mozilla research: Browsing histories are unique enough to identify users

#96
post #69

Earlier quoted context omitted.

Needs to be specified that this was an opt-in study that you had to agree to.

Is no one in this thread going to read this article? Seriously, it isn't that long. RTFM >> The new experiment got underway between July 16 and August 13, 2019, when Mozilla prompted Firefox users to take part of this experiment. >> Mozilla researchers said that more than 52,000 users agreed to take part and agreed to provide anonymous browsing data.

Welcome to HN, where the majority reads nothing but the headline and discus how they feel about the headline.

Re: Mozilla research: Browsing histories are unique enough to identify users

#97
As counterstrategy you can use tools like http://trackmenot.io/

"TrackMeNot runs as a low-priority background process that periodically issues randomized search-queries to popular search engines, e.g., AOL, Yahoo!, Google, and Bing. It hides users' actual search trails in a cloud of 'ghost' queries, significantly increasing the difficulty of aggregating such data into accurate or identifying user profiles. "

Re: Mozilla research: Browsing histories are unique enough to identify users

#99

Earlier quoted context omitted.

Is no one in this thread going to read this article? Seriously, it isn't that long. RTFM >> The new experiment got underway between July 16 and August 13, 2019, when Mozilla prompted Firefox users to take part of this experiment. >> Mozilla researchers said that more than 52,000 users agreed to take part and agreed to provide anonymous browsing data.

Welcome to HN, where the majority reads nothing but the headline and discus how they feel about the headline.

That's nothing to do with HN, that's just how the internet works.

Re: Mozilla research: Browsing histories are unique enough to identify users

#100
post #24

I feel inclined to say "... well yeah, obviously". Not in the "obvious in retrospect" way, but because browsers have been progressively blocking history-sniffing tactics for years precisely because advertisers were using it to identify visitors. Did this research... establish better numbers around it or something?

The subtitle is "Just 50-150 of our favorite sites are enough." More numbers further on in the article, although not that many.
Post reply on HN