Live data from Hacker News

Large-Scale Online Deanonymization with LLMs

simonlermen.substack.com

1–10 of 258 posts

Re: Large-Scale Online Deanonymization with LLMs

#3
i haven't read the full study, but its been on my mind for a while.

https://en.wikipedia.org/wiki/Stylometry

The best course of action to combat this correlation/profiling, seems to be usage of a local llm that rewrites the text while keeping meaning untouched.

Ideally built into a browser like Firefox/Brave.

Re: Large-Scale Online Deanonymization with LLMs

#4

so if they put their linkedin account on their HN account, we can figure out who they are.... genius stuff, AI really is changing the landscape all right

To be clear, we are making a clear concession here that the people weren't truly anonymous. But we did use an LLM to remove any identifying information from HN making them quasi-anonymous, this is more described in the appendix Table 2.

We do also make a more real world like test in section 2. There we use the anthropic interviewer dataset which Anthropic redacted, from the redacted interviews our agent identified 9/125 people based on clues.

The blog post might be more approachable for a quick take: https://simonlermen.substack.com/p/large-scale-online-deanon...

Re: Large-Scale Online Deanonymization with LLMs

#5

so if they put their linkedin account on their HN account, we can figure out who they are.... genius stuff, AI really is changing the landscape all right

That's what I'm wondering, since my linkedin profile is indeed linked to in my HN profile.

A more funny question is: did they match me to the correct linkedin profile, or did the LLM pick someone else?

Re: Large-Scale Online Deanonymization with LLMs

#6
post #3

i haven't read the full study, but its been on my mind for a while. https://en.wikipedia.org/wiki/Stylometry The best course of action to combat this correlation/profiling, seems to be usage of a local llm that rewrites the text while keeping meaning untouched. Ideally built into a browser like Firefox/Brave.

We don't use (much) stylometry, so this won't help. This is totally something you could try, but we use interests and clues. Semantic information you reveal about yourself.

The blog post might be more approachable if you want to get a quick take: https://simonlermen.substack.com/p/large-scale-online-deanon...

Re: Large-Scale Online Deanonymization with LLMs

#9
post #3

i haven't read the full study, but its been on my mind for a while. https://en.wikipedia.org/wiki/Stylometry The best course of action to combat this correlation/profiling, seems to be usage of a local llm that rewrites the text while keeping meaning untouched. Ideally built into a browser like Firefox/Brave.

I don't think this is working any more, but there was a stylometic analysis of HN users a few years ago, and it was extremely effective (at least, for myself and people who felt the need to post in the comments): https://news.ycombinator.com/item?id=33755016

Re: Large-Scale Online Deanonymization with LLMs

#10
post #5

so if they put their linkedin account on their HN account, we can figure out who they are.... genius stuff, AI really is changing the landscape all right

That's what I'm wondering, since my linkedin profile is indeed linked to in my HN profile. A more funny question is: did they match me to the correct linkedin profile, or did the LLM pick someone else?

[deleted]
Post reply on HN