Live data from Hacker News

Large-Scale Online Deanonymization with LLMs

simonlermen.substack.com

161–170 of 258 posts

Re: Large-Scale Online Deanonymization with LLMs

#161

I post under my real name here, pretty much the only place I post. It keeps me honest and straight in what I say when I choose to say it. I tried talking to my children about leaving as clean of a footprint on the internet as one can in anticipation of future people/systems taking that into consideration. I don't know what it will be but I would expect some adversarial stuff. Trying to keep clean is what I'd prefer f…

I think as the younger generations come of age they simply will not care about that sort of thing. Like it or not, it's part of the culture and might just be accepted as the norm.

When the younger generation comes of age the new younger generation will have a different culture and norm what is acceptable.

People got in trouble for things they posted years ago where they didn‘t care but others did

Re: Large-Scale Online Deanonymization with LLMs

#162
Maybe I missed something, but I see little evidence that there is a concerning ability to deanonymize. Many people post under a pseudonym but then link to their GitHub etc. In fact by construction the HN dataset _only_ consists of people who are comfortable with their real identity being linked to it.

The real question is whether someone who is pseudonymous and actually attempting to remain so can be deanonymized.

Re: Large-Scale Online Deanonymization with LLMs

#163

I'm not sure the practical implications are as dramatic as the paper suggests. Most adversaries who would want to deanonymize people at scale (governments, corporations) already have access to far more direct methods. The people most at risk from this are probably activists and whistleblowers in jurisdictions where those direct methods aren't available, not average users.

While you're right as in, it's nothing new given a trail of info, here they didn't need to do classical feature engineering, but purely LLM (agentic) flow. But yes, given how much information is self exposed online I am not surprised this is made easier with LLMs. But the interesting application is identifying users with multiple usernames on HN or reddit.

Re: Large-Scale Online Deanonymization with LLMs

#164
I bet we're about to see reduction of online public communications. Count how many times you had a desire to share your knowledge or correct someone online (aka somebody is WRONG on the internet). People would stop doing that, just to not train some big-corp model using their knowledge. Artists already not happy about that, but there are many other types of expertise people will stop sharing.

Re: Large-Scale Online Deanonymization with LLMs

#165
post #101

Earlier quoted context omitted.

I think he's wrong and I'm willing to say that. The ability for people to move beyond the fundamental attribution error is well known and takes major resources to correct that. For anyone that posts a comment, assuming you want to have easy attribution later is that you must future proof your words. That is not possible and it is extremely suppressive to express yourself. For example: "Ellen Page is fantastic in the…

> That is not possible and it is extremely suppressive to express yourself. Also for the fact that you cannot predict how future powers will view past comments - for instance, certain benign political views 20 years ago could become "terroristic speech" tomorrow. I operate by a simple, general rule - I don't often say anything online I wouldn't say directly to someone's face in real life.

I think the problem with this, especially amongst younger people, is having spent so much time online, they don't know where to draw this line anymore.

Re: Large-Scale Online Deanonymization with LLMs

#166
everyone in the comments is talking about stylometry and rewriting your posts with LLMs. the paper barely uses stylometry. the attack surface is semantic: your interests, your city, the conference you mentioned once 2 years ago. you can't rewrite your way out of having said you work in fintech in austin and own a golden retriever.

Re: Large-Scale Online Deanonymization with LLMs

#168
post #166

everyone in the comments is talking about stylometry and rewriting your posts with LLMs. the paper barely uses stylometry. the attack surface is semantic: your interests, your city, the conference you mentioned once 2 years ago. you can't rewrite your way out of having said you work in fintech in austin and own a golden retriever.

you can intentionally add false biographical information. what if you had a bot posting responses in subreddits for cities across the world on your account

Re: Large-Scale Online Deanonymization with LLMs

#169

Despite being pseudonymous, I don’t take great pains to hide who I am. I am in my 50s and live on the West coast. I don’t have socials and I don’t post anywhere else. Have at it! If you are semi-retired, you’re free from the threat of cancellation. As long as you aren’t posting about crimes, there’s limits to what anyone can legally do to you. (Still, it’s good to be prudent and limit sharing.)

Kind of short sighted only consider social cancellation. People in power change, laws get applied retroactively. History is full of people who get purged from stuff that was fine when it was written

Re: Large-Scale Online Deanonymization with LLMs

#170
post #101

Earlier quoted context omitted.

I think he's wrong and I'm willing to say that. The ability for people to move beyond the fundamental attribution error is well known and takes major resources to correct that. For anyone that posts a comment, assuming you want to have easy attribution later is that you must future proof your words. That is not possible and it is extremely suppressive to express yourself. For example: "Ellen Page is fantastic in the…

> That is not possible and it is extremely suppressive to express yourself. Also for the fact that you cannot predict how future powers will view past comments - for instance, certain benign political views 20 years ago could become "terroristic speech" tomorrow. I operate by a simple, general rule - I don't often say anything online I wouldn't say directly to someone's face in real life.

This is very import: you don't know how the cancelation culture will be in 20 years.

I like to use the example of a guy who did a blackface in a party back in 2000's. Although reprehensible, was not commom-sense racism back then. Today society sees it as completely unacceptable.

Eventually that guy became prime minister of Canada and things went pretty bad when that photo surfaced decades later.

Is it far to judge someone's actions by the lens of a different culture? When the popular opinion comes, they won't care about historical context.

Post reply on HN