Live data from Hacker News

Large-Scale Online Deanonymization with LLMs

simonlermen.substack.com

11–20 of 258 posts

Re: Large-Scale Online Deanonymization with LLMs

#11
post #3

i haven't read the full study, but its been on my mind for a while. https://en.wikipedia.org/wiki/Stylometry The best course of action to combat this correlation/profiling, seems to be usage of a local llm that rewrites the text while keeping meaning untouched. Ideally built into a browser like Firefox/Brave.

We don't use (much) stylometry, so this won't help. This is totally something you could try, but we use interests and clues. Semantic information you reveal about yourself. The blog post might be more approachable if you want to get a quick take: https://simonlermen.substack.com/p/large-scale-online-deanon...

Thanks for the providing the details, where I've been just lazy about reading the paper now :))

I'm not a fan of your proposed changes, as they further lock down platforms.

I'd like to see better tools for users to engage with. Maybe if someone is in their Firefox anonymous (or private tab) profile they should be warned when writing about locations, jobs, politics, etc. Even there a small local LLM model would be useful, not foolproof, but an extra layet of checks. Paired with protection about stylometry :D

Re: Large-Scale Online Deanonymization with LLMs

#12
Additionally, you can open up copilot.microsoft.com or w/e and ask it to summarize any reddit users (and presumably HN) posts. Not just the content, but their emotional state (without prompting).

[0] Note: last I tried this was months ago, things may have changed.

Re: Large-Scale Online Deanonymization with LLMs

#13
post #3

i haven't read the full study, but its been on my mind for a while. https://en.wikipedia.org/wiki/Stylometry The best course of action to combat this correlation/profiling, seems to be usage of a local llm that rewrites the text while keeping meaning untouched. Ideally built into a browser like Firefox/Brave.

There is also a practical issue here that people usually don't write a lot on linkedin, most people just have structured biographical information. We use very limited stylometry in section 6 for matching reddit users who we synthetically split according to time.

Re: Large-Scale Online Deanonymization with LLMs

#15
post #3

i haven't read the full study, but its been on my mind for a while. https://en.wikipedia.org/wiki/Stylometry The best course of action to combat this correlation/profiling, seems to be usage of a local llm that rewrites the text while keeping meaning untouched. Ideally built into a browser like Firefox/Brave.

[flagged]

Re: Large-Scale Online Deanonymization with LLMs

#17
post #3

i haven't read the full study, but its been on my mind for a while. https://en.wikipedia.org/wiki/Stylometry The best course of action to combat this correlation/profiling, seems to be usage of a local llm that rewrites the text while keeping meaning untouched. Ideally built into a browser like Firefox/Brave.

We don't use (much) stylometry, so this won't help. This is totally something you could try, but we use interests and clues. Semantic information you reveal about yourself. The blog post might be more approachable if you want to get a quick take: https://simonlermen.substack.com/p/large-scale-online-deanon...

[deleted]

Re: Large-Scale Online Deanonymization with LLMs

#18
post #11

Earlier quoted context omitted.

We don't use (much) stylometry, so this won't help. This is totally something you could try, but we use interests and clues. Semantic information you reveal about yourself. The blog post might be more approachable if you want to get a quick take: https://simonlermen.substack.com/p/large-scale-online-deanon...

Thanks for the providing the details, where I've been just lazy about reading the paper now :)) I'm not a fan of your proposed changes, as they further lock down platforms. I'd like to see better tools for users to engage with. Maybe if someone is in their Firefox anonymous (or private tab) profile they should be warned when writing about locations, jobs, politics, etc. Even there a small local LLM model would be use…

Mitigations are pretty difficult, I understand it is kind of cool that some websites have really open APIs where you can just read everything. There are some cool apps that used HN data in the past. But I think there should at least be consideration that LLMs are then going to read everything and potentially discover things. Users might have thought this is protected by obscurity, who would read their 5 year old comments?

Re: Large-Scale Online Deanonymization with LLMs

#19
post #12

Additionally, you can open up copilot.microsoft.com or w/e and ask it to summarize any reddit users (and presumably HN) posts. Not just the content, but their emotional state (without prompting). [0] Note: last I tried this was months ago, things may have changed.

I just retried this with my reddit account (game dev stuff)

Last block of text from copilot :/

-----------

If you want, I can also break down:

Their posting style (tone, frequency, community engagement)

How their work compares to other indie city builders

What seems to resonate most with Reddit users

Just tell me what angle you want to explore next.

Post reply on HN