Live data from Hacker News

Large-Scale Online Deanonymization with LLMs

simonlermen.substack.com

71–80 of 258 posts

Re: Large-Scale Online Deanonymization with LLMs

#71

Earlier quoted context omitted.

the most important data for LLM is that Microsoft in general and GitHub in particular can never be trusted with your data. I've been trying to delete my GitHub account for many months

> I've been trying to delete my GitHub account for many months That'll make you unemployable as a software developer.

Luckily I don't want to be employable as a software developer

Re: Large-Scale Online Deanonymization with LLMs

#72
post #3

i haven't read the full study, but its been on my mind for a while. https://en.wikipedia.org/wiki/Stylometry The best course of action to combat this correlation/profiling, seems to be usage of a local llm that rewrites the text while keeping meaning untouched. Ideally built into a browser like Firefox/Brave.

[flagged]

I don't really understand the argument your proposing.

Is it impressions in a stylistic sense (flurishes to the language used), which is a what I'm arguing the LLM usage for.

Or is it impression in the subjective sense of what an author would instill through his message. Feelings, imagry, and such.

Or the impression given to the reader? "This person gives me the impression that they know what they talk about", or "don't know what they talk about?"

I don't know which argument your proposing, but I'd like to make an observation of the LLM usage. I don't know what model the perplexity response is based on, but some of them are "eager to please" by default in conversation("you're absolutely right" and all the other memes). If you "preload" it with a contrarian approach (make a brutally honest critique of this comment in reply to this other comment) it will gladly do a 180 https://chatgpt.com/s/t_699f3b13826c8191b701d0cc84923e71

Re: Large-Scale Online Deanonymization with LLMs

#73

Earlier quoted context omitted.

the most important data for LLM is that Microsoft in general and GitHub in particular can never be trusted with your data. I've been trying to delete my GitHub account for many months

> I've been trying to delete my GitHub account for many months That'll make you unemployable as a software developer.

Software developer for 20 years here, never had a problem getting jobs without a github

Maybe that will change in the future. Then again I'm pretty sure my next job won't be software. I have no interest in building software in the AI era.

Re: Large-Scale Online Deanonymization with LLMs

#75

As people will point out, the OSINT techniques described are nothing new - typically, in the past, you could de-anonymize based on writing style or niche topics/interests. Totally deanonymization can occur if any of these accounts link to profiles containing pictures of their faces, which can then be web-searched to link to a real identity. It's astounding how many people re-use handles on stuff like porn sites linke…

If LLMs can identify a person across websites, I can ask LLM to read up his posts and write like him impersonating him and then this feeds back into the tools identifying him. I can probabilistically malign a person this way.

Re: Large-Scale Online Deanonymization with LLMs

#76

Earlier quoted context omitted.

I think the implication is this will become trivial and trivially automated, no human investigator needed. I bet there will be plugins in one year's time to right click on a post and get a full report on who the author is.

Wouldn't it also become trivial to pretend to be another author?

it may become more trivial to llm your comments/blog/whatever into a different "voice", but there is so much that can be used for de-anonymization that the llm-assisted technique dont address.

for example, you may change the content of your comments, but if you only ever comment on the same topic, the topic itself is a signal. when you post (both day and time), frequency of posts, topics of interest, usernames (e.g. themes or patterns), and much more.

Re: Large-Scale Online Deanonymization with LLMs

#77

As people will point out, the OSINT techniques described are nothing new - typically, in the past, you could de-anonymize based on writing style or niche topics/interests. Totally deanonymization can occur if any of these accounts link to profiles containing pictures of their faces, which can then be web-searched to link to a real identity. It's astounding how many people re-use handles on stuff like porn sites linke…

If LLMs can identify a person across websites, I can ask LLM to read up his posts and write like him impersonating him and then this feeds back into the tools identifying him. I can probabilistically malign a person this way.

stylometry is only one aspect of de-anonymization. what you describe is certainly a threat that we will have to deal with, but there is a lot more to credible impersonation than just being able to mimic a writing style

Re: Large-Scale Online Deanonymization with LLMs

#79

As people will point out, the OSINT techniques described are nothing new - typically, in the past, you could de-anonymize based on writing style or niche topics/interests. Totally deanonymization can occur if any of these accounts link to profiles containing pictures of their faces, which can then be web-searched to link to a real identity. It's astounding how many people re-use handles on stuff like porn sites linke…

If LLMs can identify a person across websites, I can ask LLM to read up his posts and write like him impersonating him and then this feeds back into the tools identifying him. I can probabilistically malign a person this way.

How to conduct a psy-op

https://youtu.be/YTGQXVmrc6g

Re: Large-Scale Online Deanonymization with LLMs

#80

As people will point out, the OSINT techniques described are nothing new - typically, in the past, you could de-anonymize based on writing style or niche topics/interests. Totally deanonymization can occur if any of these accounts link to profiles containing pictures of their faces, which can then be web-searched to link to a real identity. It's astounding how many people re-use handles on stuff like porn sites linke…

If LLMs can identify a person across websites, I can ask LLM to read up his posts and write like him impersonating him and then this feeds back into the tools identifying him. I can probabilistically malign a person this way.

This already is a thing people did at least as far back as I started getting into web privacy, which was ~10 years ago. I have been the target of it before.

LLM's are probably better at it, but I don't know if this is as destructive as people may guess it would be. Probably highly person dependent.

The micro-signals this paper discusses are more difficult to fake.

Post reply on HN