Live data from Hacker News

Large-Scale Online Deanonymization with LLMs

simonlermen.substack.com

231–240 of 258 posts

Re: Large-Scale Online Deanonymization with LLMs

#231

Many years ago (early 2000s) I worked for a firm that would help identify people who were doing "pump and dump" stock scams on Yahoo Finance message boards. Step 1 was to scrape all of their posts into a database. Step 2 was to have a human analyst review all of the posts for clues about who that person was It was amazing that you could easily figure out: - if they were at work or home from when they posted (9am to 5…

I recently decided to play around with this, given... well my profile... and I will say that Gemini was good at zeroing in on who I was, but for whatever reason would refuse to stay my name.

Re: Large-Scale Online Deanonymization with LLMs

#232

Earlier quoted context omitted.

One could just as easily make the opposite argument. Given that your values and priorities may change significantly over the decades, a smart investment now into a solid, stable, and prosocial public identity may reap considerable and wide-ranging benefits in ways you couldn't even predict. This is especially true if you take seriously the idea that it's not what you say but how you say it that matters in the end.

This is actually what I believe as well although I believe that its better to be pseudo-anonymous for me, right now. In the sense that if I ever create any business/idea which can be serious enough that I want to back it up. I might create hackernews post about it. Although that being said, I do sometimes make alts just to publish something if I don't want it under this particular account. I do feel like I can be wro…

You seem polite enough even psuedonymously, so I'd say you're doing a good job so far. :)

>I have had some paranoid thoughts as to what if I get into controversy later on in life because of some things I do in my teen years

I have a relevant anecdote, from back in halcyon 2008. Maybe it will help you when it comes to believing your friend, or at least it will temper your paranoia, which I think is well meaning in small doses.

When I was 13 or 14 years old I got suspended from high school because a friend posted a link to the Anarchist's Cookbook, which I had never heard of, on my Facebook wall. Some of my classmates got very scared and called the headmaster saying I had made a bomb threat against the school.

When the principal pulled me in to talk to me about this, it became very clear I had no idea what they were talking about. We talked for much longer than I think anyone in the room expected, maybe for three hours about existentialism, Zappfe's essay The Last Messiah which I had read the night before, whether I thought I was a victim of bullying (I didn't), what I thought of the school (excellent, a welcome refuge from a very turbulent home), thoughts on Cicero's speeches, the books we were reading in English class at the time.

I got "suspended" for a week and my parents took me to a therapist for several months afterward. I had thought after this for the rest of high school that my chances of ever going to college were totally shot, because a suspension appears on your permanent record. However, when it came time for me to actually apply to colleges, I found out no such record of the week at home ever existed. There appeared to have been a miscommunication all those years ago; I had actually been put on some kind of medical leave.

Now of course going through all of high school thinking that no college in the country will accept you now no matter how hard you do is going to change your incentives a bit. Ironically the very thinkers I had been reading at the time helped me quickly conclude that I wanted to do my level best anyway, even if there was going to be no payoff at the end of the road at all for me. In some ways it let me take more risks than my other classmates. I became the earliest person in my class to take our infamously hard physics course, and I walked out with top marks on both kinematics and electromagnetism. I don't think I would have taken that risk if I thought I had to optimize my GPA.

I trust you to think about this story and come to your own conclusions on how it moves your needle.

Re: Large-Scale Online Deanonymization with LLMs

#233
What's wild to me is that people worry about writing style fingerprinting while casually uploading their literal DNA to consumer genomics companies. 23andMe went bankrupt and suddenly 15 million people's most identifying data imaginable is an asset in a fire sale.

Your writing style can theoretically be masked with an LLM. Your genome can't. And it doesn't just identify you -- it identifies your relatives, your disease risks, your ancestry, things you might not even know about yourself yet. The deanonymization vector here is permanent and irrevocable in a way that no amount of OPSEC can fix after the fact.

The semantic approach in this paper (interests, clues, behavioral patterns) is scary enough. Now imagine combining that with leaked genetic data. You don't even need to match writing styles when you can match someone's 23andMe profile to their health subreddit posts about conditions they're genetically predisposed to.

Re: Large-Scale Online Deanonymization with LLMs

#235

Earlier quoted context omitted.

> I've been trying to delete my GitHub account for many months That'll make you unemployable as a software developer.

Luckily I don't want to be employable as a software developer

Amen comrade

Re: Large-Scale Online Deanonymization with LLMs

#236

many people tend to overlook how little information is needed for successful de-anonymization. i like to introduce students to de-anonymization with an old paper "Robust De-anonymization of Large Sparse Datasets" published in the ancient history of 2008 ( https://www.cs.cornell.edu/~shmat/shmat_oak08netflix.pdf ): " We apply our de-anonymization methodology to the Netflix Prize dataset, which contains anonymous movie…

We don't need everyone to be completely anonymous to state and corporate actors. We just need to make it so that they can't identify and surveil everyone at once, because it would be too expensive.

The US defense budget is about $1T dollars. They can't spend it all on surveillance, but let's say tech companies + gov spends about this amount per year on surveillance in total. If we can raise the cost to surveil the average person to over $10K/yr, they just lose. This is very doable.

Every little precaution you take will raise the cost, probably more than you think. Every open-source project that aims to anonymize and decentralize is an arrow in their knee. They're hoping that you'll get cynical and stop trying because they don't stand a chance otherwise.

Re: Large-Scale Online Deanonymization with LLMs

#237

Earlier quoted context omitted.

Throwaway accounts using "clever" turns of phrase can often be anonymized by double click, right-clicking -> googling their witty pun and seeing their the sole instance elsewhere, on Twitter, Facebook, etc If I see a couple words I dont know in a row, I can infer a posters real name. Id be more specific but any example is doxxing, literally so

I assume one's vocabulary is basically a fingerprint, even if one doesn't use unique turns of phrase. Domain knowledge just leaks in and we aren't conscious of it being identifiable.

It also geographic. There's a bunch of quizzes online where in 10 or 20 questions, it can tell you exactly what area in the US somebody is from. It comes down to the terms you use that you might not even realize are not universal. Highway vs freeway, what you call a sugary carbonated drink, and so on.

OTOH I think a lot of these methods don't matter that much because of plausible deniability. Stylometry and other stuff processes is always probabilistic, and can be dismissed.

Re: Large-Scale Online Deanonymization with LLMs

#238
Information leaks everywhere, as the ability to process it increases, I think ultimately it will lead to a world where there are no secrets, provided one has the resources and intention to look for something.

For a few years now I have been telling people how unprepared the world is for this change. Not understanding how this is possible will lead to people outright deifying AI that has the capability to do things like this. It will seem like omniscience.

I think the main protection we have in a world where you cannot effectively hide, is that anyone who abuses this ability will be operating under the same system. You can use it to your advantage, but not without getting caught.

Re: Large-Scale Online Deanonymization with LLMs

#239

Earlier quoted context omitted.

I assume one's vocabulary is basically a fingerprint, even if one doesn't use unique turns of phrase. Domain knowledge just leaks in and we aren't conscious of it being identifiable.

It also geographic. There's a bunch of quizzes online where in 10 or 20 questions, it can tell you exactly what area in the US somebody is from. It comes down to the terms you use that you might not even realize are not universal. Highway vs freeway, what you call a sugary carbonated drink, and so on. OTOH I think a lot of these methods don't matter that much because of plausible deniability. Stylometry and other stu…

>OTOH I think a lot of these methods don't matter that much because of plausible deniability. Stylometry and other stuff processes is always probabilistic, and can be dismissed.

while all of it is probabilistic, the issue is that the probability can quickly begin to approach 1 when multiple sources of data & varying techniques are combined.

Post reply on HN