Live data from Hacker News

Large-Scale Online Deanonymization with LLMs

simonlermen.substack.com

221–230 of 258 posts

Re: Large-Scale Online Deanonymization with LLMs

#221

many people tend to overlook how little information is needed for successful de-anonymization. i like to introduce students to de-anonymization with an old paper "Robust De-anonymization of Large Sparse Datasets" published in the ancient history of 2008 ( https://www.cs.cornell.edu/~shmat/shmat_oak08netflix.pdf ): " We apply our de-anonymization methodology to the Netflix Prize dataset, which contains anonymous movie…

A silver lining of the ai apocolypse is that users may be able to use the technology to maintain their anonymity via llm paraphrasing.

Re: Large-Scale Online Deanonymization with LLMs

#222

A related past submission comes to mind: Show HN: Using stylometry to find HN users with alternate accounts https://news.ycombinator.com/item?id=33755016 - Nov 2022, 519 comments

This HN stylometry tool is still online: https://antirez.com/hnstyle (though I assume its dataset is not kept updated since mid-2025).

20250415 https://news.ycombinator.com/item?id=43705632 Reproducing Hacker News writing style fingerprinting (325 points, 159 comments)

Re: Large-Scale Online Deanonymization with LLMs

#223
post #221

many people tend to overlook how little information is needed for successful de-anonymization. i like to introduce students to de-anonymization with an old paper "Robust De-anonymization of Large Sparse Datasets" published in the ancient history of 2008 ( https://www.cs.cornell.edu/~shmat/shmat_oak08netflix.pdf ): " We apply our de-anonymization methodology to the Netflix Prize dataset, which contains anonymous movie…

A silver lining of the ai apocolypse is that users may be able to use the technology to maintain their anonymity via llm paraphrasing.

My guess is that a statistical analysis of other things such as access patterns, timestamps, content you engage with, etc, could de-anonymize you regardless of the phrasing you use, so LLMs won't save you.

Re: Large-Scale Online Deanonymization with LLMs

#224

many people tend to overlook how little information is needed for successful de-anonymization. i like to introduce students to de-anonymization with an old paper "Robust De-anonymization of Large Sparse Datasets" published in the ancient history of 2008 ( https://www.cs.cornell.edu/~shmat/shmat_oak08netflix.pdf ): " We apply our de-anonymization methodology to the Netflix Prize dataset, which contains anonymous movie…

MIT showed this in 13 after the government was caught illegally spying on Americans with “just metadata”: https://www.nature.com/articles/srep01376

Re: Large-Scale Online Deanonymization with LLMs

#225
post #101

Earlier quoted context omitted.

I think he's wrong and I'm willing to say that. The ability for people to move beyond the fundamental attribution error is well known and takes major resources to correct that. For anyone that posts a comment, assuming you want to have easy attribution later is that you must future proof your words. That is not possible and it is extremely suppressive to express yourself. For example: "Ellen Page is fantastic in the…

> That is not possible and it is extremely suppressive to express yourself. Also for the fact that you cannot predict how future powers will view past comments - for instance, certain benign political views 20 years ago could become "terroristic speech" tomorrow. I operate by a simple, general rule - I don't often say anything online I wouldn't say directly to someone's face in real life.

> I operate by a simple, general rule - I don't often say anything online I wouldn't say directly to someone's face in real life.

I think this isn't enough for the digital age, simply because "comments you'd say to someone's face" can compromise you on the internet.

Some dirty joke, gossip or whatever you tell a friend, if posted online, could come back to bite you in the ass in the dystopian future, lose you your job, or worse.

Re: Large-Scale Online Deanonymization with LLMs

#226
post #182

Earlier quoted context omitted.

I can be a rude prick online sometimes, but I can be in real life too - basically though the reason I do this is I never want it to be some huge surprise IRL if someone sees what I write online and be like, "wow, I didn't know that about him." I'm pretty much what I am online and IRL the same. For some reason this seems to matter for me, at least in the past when people have tried to like, send employers stuff I may…

As someone who gets dopamine hits from downvotes on HN, I approve of your behavior! >just be yourself basically Yea, it is boring when everyone is the same. I would like a rude but interesting world (even if I might not survive long in one), than a nice, boring one.

"Just be yourself" seems to me a lot like the rightfully discredited "if you don't have anything to hide...".

Everybody has something to hide. Everybody has said things they regret, or meant to be heard by some people but not others.

Re: Large-Scale Online Deanonymization with LLMs

#227
Maybe it's time to finally track down this person: http://voidnull.sdf.org

This page is anonymous

20190119 https://news.ycombinator.com/item?id=20220048 (149 points, 51 comments)

20130501 https://news.ycombinator.com/item?id=5638988 (453 points, 243 comments)

https://news.ycombinator.com/threads?id=voidnull

https://antirez.com/hnstyle?username=voidnull

Re: Large-Scale Online Deanonymization with LLMs

#229
post #189

Earlier quoted context omitted.

well, how about "abortion legal" to "abortion murder"... possible to see this coming, but I know doctors in NY who are now afraid to travel to Texas. How about DEI initiatives as good things in 2024 and a mark of evil in 2025? Lots of people were fired because in 2024 their boss told them to work on DEI and they did what their boss told them to do. Turns out this was a capital offense.

> because in 2024 their boss told them I am not commenting on your specific example of DEI but I want to make the general point that you are always responsible for what you do, irregardless of whether you were told to do it by your boss, or commanding officer, or whatever. So again, I don't care about the specific example you used but if something is 'in fashion' and you go along with it, including at work, then you…

But working on DEI on your boss' orders in 2024 wasn't reprobable, anymore than bringing your boss a cup of coffee to their desk was.

The point is that the shift in what is considered "a capital crime" is arbitrary, this is not the Nuremberg trials. You cannot protect yourself by being a decent person, whatever you do today can be a crime tomorrow, and AI can assist those looking for your flaws.

Re: Large-Scale Online Deanonymization with LLMs

#230
post #221

many people tend to overlook how little information is needed for successful de-anonymization. i like to introduce students to de-anonymization with an old paper "Robust De-anonymization of Large Sparse Datasets" published in the ancient history of 2008 ( https://www.cs.cornell.edu/~shmat/shmat_oak08netflix.pdf ): " We apply our de-anonymization methodology to the Netflix Prize dataset, which contains anonymous movie…

A silver lining of the ai apocolypse is that users may be able to use the technology to maintain their anonymity via llm paraphrasing.

as the_af says, stylometry is only one technique in a bag of techniques used for de-anonymization. a big one to be sure, but nowhere near the only one.
Post reply on HN