Many years ago (early 2000s) I worked for a firm that would help identify people who were doing "pump and dump" stock scams on Yahoo Finance message boards. Step 1 was to scrape all of their posts into a database. Step 2 was to have a human analyst review all of the posts for clues about who that person was It was amazing that you could easily figure out: - if they were at work or home from when they posted (9am to 5…
Large-Scale Online Deanonymization with LLMs
231–240 of 258 posts
Re: Large-Scale Online Deanonymization with LLMs
#232Earlier quoted context omitted.
One could just as easily make the opposite argument. Given that your values and priorities may change significantly over the decades, a smart investment now into a solid, stable, and prosocial public identity may reap considerable and wide-ranging benefits in ways you couldn't even predict. This is especially true if you take seriously the idea that it's not what you say but how you say it that matters in the end.
This is actually what I believe as well although I believe that its better to be pseudo-anonymous for me, right now. In the sense that if I ever create any business/idea which can be serious enough that I want to back it up. I might create hackernews post about it. Although that being said, I do sometimes make alts just to publish something if I don't want it under this particular account. I do feel like I can be wro…
>I have had some paranoid thoughts as to what if I get into controversy later on in life because of some things I do in my teen years
I have a relevant anecdote, from back in halcyon 2008. Maybe it will help you when it comes to believing your friend, or at least it will temper your paranoia, which I think is well meaning in small doses.
When I was 13 or 14 years old I got suspended from high school because a friend posted a link to the Anarchist's Cookbook, which I had never heard of, on my Facebook wall. Some of my classmates got very scared and called the headmaster saying I had made a bomb threat against the school.
When the principal pulled me in to talk to me about this, it became very clear I had no idea what they were talking about. We talked for much longer than I think anyone in the room expected, maybe for three hours about existentialism, Zappfe's essay The Last Messiah which I had read the night before, whether I thought I was a victim of bullying (I didn't), what I thought of the school (excellent, a welcome refuge from a very turbulent home), thoughts on Cicero's speeches, the books we were reading in English class at the time.
I got "suspended" for a week and my parents took me to a therapist for several months afterward. I had thought after this for the rest of high school that my chances of ever going to college were totally shot, because a suspension appears on your permanent record. However, when it came time for me to actually apply to colleges, I found out no such record of the week at home ever existed. There appeared to have been a miscommunication all those years ago; I had actually been put on some kind of medical leave.
Now of course going through all of high school thinking that no college in the country will accept you now no matter how hard you do is going to change your incentives a bit. Ironically the very thinkers I had been reading at the time helped me quickly conclude that I wanted to do my level best anyway, even if there was going to be no payoff at the end of the road at all for me. In some ways it let me take more risks than my other classmates. I became the earliest person in my class to take our infamously hard physics course, and I walked out with top marks on both kinematics and electromagnetism. I don't think I would have taken that risk if I thought I had to optimize my GPA.
I trust you to think about this story and come to your own conclusions on how it moves your needle.
Re: Large-Scale Online Deanonymization with LLMs
#233Your writing style can theoretically be masked with an LLM. Your genome can't. And it doesn't just identify you -- it identifies your relatives, your disease risks, your ancestry, things you might not even know about yourself yet. The deanonymization vector here is permanent and irrevocable in a way that no amount of OPSEC can fix after the fact.
The semantic approach in this paper (interests, clues, behavioral patterns) is scary enough. Now imagine combining that with leaked genetic data. You don't even need to match writing styles when you can match someone's 23andMe profile to their health subreddit posts about conditions they're genetically predisposed to.
Re: Large-Scale Online Deanonymization with LLMs
#234Re: Large-Scale Online Deanonymization with LLMs
#235Re: Large-Scale Online Deanonymization with LLMs
#236many people tend to overlook how little information is needed for successful de-anonymization. i like to introduce students to de-anonymization with an old paper "Robust De-anonymization of Large Sparse Datasets" published in the ancient history of 2008 ( https://www.cs.cornell.edu/~shmat/shmat_oak08netflix.pdf ): " We apply our de-anonymization methodology to the Netflix Prize dataset, which contains anonymous movie…
The US defense budget is about $1T dollars. They can't spend it all on surveillance, but let's say tech companies + gov spends about this amount per year on surveillance in total. If we can raise the cost to surveil the average person to over $10K/yr, they just lose. This is very doable.
Every little precaution you take will raise the cost, probably more than you think. Every open-source project that aims to anonymize and decentralize is an arrow in their knee. They're hoping that you'll get cynical and stop trying because they don't stand a chance otherwise.
Re: Large-Scale Online Deanonymization with LLMs
#237Earlier quoted context omitted.
Throwaway accounts using "clever" turns of phrase can often be anonymized by double click, right-clicking -> googling their witty pun and seeing their the sole instance elsewhere, on Twitter, Facebook, etc If I see a couple words I dont know in a row, I can infer a posters real name. Id be more specific but any example is doxxing, literally so
I assume one's vocabulary is basically a fingerprint, even if one doesn't use unique turns of phrase. Domain knowledge just leaks in and we aren't conscious of it being identifiable.
OTOH I think a lot of these methods don't matter that much because of plausible deniability. Stylometry and other stuff processes is always probabilistic, and can be dismissed.
Re: Large-Scale Online Deanonymization with LLMs
#238For a few years now I have been telling people how unprepared the world is for this change. Not understanding how this is possible will lead to people outright deifying AI that has the capability to do things like this. It will seem like omniscience.
I think the main protection we have in a world where you cannot effectively hide, is that anyone who abuses this ability will be operating under the same system. You can use it to your advantage, but not without getting caught.
Re: Large-Scale Online Deanonymization with LLMs
#239Earlier quoted context omitted.
I assume one's vocabulary is basically a fingerprint, even if one doesn't use unique turns of phrase. Domain knowledge just leaks in and we aren't conscious of it being identifiable.
It also geographic. There's a bunch of quizzes online where in 10 or 20 questions, it can tell you exactly what area in the US somebody is from. It comes down to the terms you use that you might not even realize are not universal. Highway vs freeway, what you call a sugary carbonated drink, and so on. OTOH I think a lot of these methods don't matter that much because of plausible deniability. Stylometry and other stu…
while all of it is probabilistic, the issue is that the probability can quickly begin to approach 1 when multiple sources of data & varying techniques are combined.
Re: Large-Scale Online Deanonymization with LLMs
#240Stylometry can match not only people, but ethnic groups. No LLM required.