Live data from Hacker News

LLMs can unmask pseudonymous users at scale with surprising accuracy

arstechnica.com

61–70 of 78 posts

Re: LLMs can unmask pseudonymous users at scale with surprising accuracy

#61

Only if said users happen to commit OPSEC failures themselves. LLMs aren't magic... If someone can figure out who I am or what city I live in just by this username or my comments (with proof), I'll personally send you 500,000 JPY. I'm quite confident that's not going to happen though. The paper referenced in the article does not even explain their exact testing methodology (such as the tools or exact prompts used) be…

[dead]

Re: LLMs can unmask pseudonymous users at scale with surprising accuracy

#63
post #62

So tell an LLM what you would like the post to say, and then post the output? LLM as the sickness and the cure...

This is the first thing that comes to mind. However I wonder if not only the “general” vocabulary can be anonymized but also the underlying concepts and references, because they point to a particular place too.

Re: LLMs can unmask pseudonymous users at scale with surprising accuracy

#64
post #44

Earlier quoted context omitted.

My understanding is that the GDPR “right to be forgotten” applies to personal data. Are publicly available comments considered personal data?

From ico.org.uk: “ It is important to note that opinions and inferences are also personal data, maybe special category data, if they directly or indirectly relate to that individual” From gdpr-info.eu: “ Subjective information such as opinions, judgements or estimates can be personal data.” So yes. HN is in violation of the GDPR. I had already filed a complaint about this policy at my local GDPR authority.

If you are posting public comments, then these comments are available publicly... like, what did you expect!?

Re: LLMs can unmask pseudonymous users at scale with surprising accuracy

#66
post #52

I thought this would be more about stylometry but it's mostly about users literally posting the same identifiable information across multiple services, including in one example their age, dog name, profession. It's all classic dox profiling techniques. Even the things like spelling differences being regional signals and commonality to specific things being discussed. It's why one has to think about what is being post…

I’m curious if an LLM-based defense for this could be made. Like a browser plugin that warns you if you type identifiable information (like occupation) into a text field, and highlights turns of phrase that are “unusual” enough to be identifiable.

or something which just inserts random untrue details about you every now and again, like they do in Alaska, where I live.

Re: LLMs can unmask pseudonymous users at scale with surprising accuracy

#69
post #44

Earlier quoted context omitted.

From ico.org.uk: “ It is important to note that opinions and inferences are also personal data, maybe special category data, if they directly or indirectly relate to that individual” From gdpr-info.eu: “ Subjective information such as opinions, judgements or estimates can be personal data.” So yes. HN is in violation of the GDPR. I had already filed a complaint about this policy at my local GDPR authority.

If you are posting public comments, then these comments are available publicly... like, what did you expect!?

Yes they are. The GDPR doesn’t say you can’t post it.

Under article 17 of the GDPR, EU citizens have the right to be forgotten, in which case this data needs to be deleted.

Post reply on HN