Many years ago (early 2000s) I worked for a firm that would help identify people who were doing "pump and dump" stock scams on Yahoo Finance message boards. Step 1 was to scrape all of their posts into a database. Step 2 was to have a human analyst review all of the posts for clues about who that person was It was amazing that you could easily figure out: - if they were at work or home from when they posted (9am to 5…
Did you have a contract with SEC? Just wondering what kind of business would have an interest in that.
Large-Scale Online Deanonymization with LLMs
201–210 of 258 posts
Re: Large-Scale Online Deanonymization with LLMs
#202Earlier quoted context omitted.
I think he's wrong and I'm willing to say that. The ability for people to move beyond the fundamental attribution error is well known and takes major resources to correct that. For anyone that posts a comment, assuming you want to have easy attribution later is that you must future proof your words. That is not possible and it is extremely suppressive to express yourself. For example: "Ellen Page is fantastic in the…
> That is not possible and it is extremely suppressive to express yourself. Also for the fact that you cannot predict how future powers will view past comments - for instance, certain benign political views 20 years ago could become "terroristic speech" tomorrow. I operate by a simple, general rule - I don't often say anything online I wouldn't say directly to someone's face in real life.
I am not going to give examples, because I don't want them to be pinned on me as my views, but I'm sure most of us have enough imagination to come up with them.
Re: Large-Scale Online Deanonymization with LLMs
#203I post under my real name here, pretty much the only place I post. It keeps me honest and straight in what I say when I choose to say it. I tried talking to my children about leaving as clean of a footprint on the internet as one can in anticipation of future people/systems taking that into consideration. I don't know what it will be but I would expect some adversarial stuff. Trying to keep clean is what I'd prefer f…
Here's a different vision for the future:
Let information filtering become each individual's own responsibility. We have LLMs now, and they'll get more efficient, so why not use them locally to filter incoming feeds according to each of our own preferences, but remove all of the filtering/moderation for posting info out. Build systems to decentralize and anonymize the Internet so that people can discover anyone and aren't afraid to post anything. Make it so that everyone can get a message out to the world and nobody can be arrested or assassinated for it. This will put an end to most violent conflict because they'd be replaced by online discourse.
Let the Internet be flooded with trash and gold at the same time. Let each individual decide what info is/isn't valuable to them. Let those individuals self-organize. Let ideas compete freely, so that the best ones may prevail.
Re: Large-Scale Online Deanonymization with LLMs
#204Earlier quoted context omitted.
> I tried talking to my children about leaving as clean of a footprint on the internet as one can in anticipation of future people/systems taking that into consideration. I don’t think you’re wrong, but the fact that people consider it inevitable we’ll all have an immutable social acceptance grade that includes everything from teenage shitposts to things you said after a loved one died, or getting diagnosed with canc…
I think he's wrong and I'm willing to say that. The ability for people to move beyond the fundamental attribution error is well known and takes major resources to correct that. For anyone that posts a comment, assuming you want to have easy attribution later is that you must future proof your words. That is not possible and it is extremely suppressive to express yourself. For example: "Ellen Page is fantastic in the…
Your point may be more valid when it comes to political attitudes, in cases where the issues were known at the time but the Overton window has shifted since.
Re: Large-Scale Online Deanonymization with LLMs
#205many people tend to overlook how little information is needed for successful de-anonymization. i like to introduce students to de-anonymization with an old paper "Robust De-anonymization of Large Sparse Datasets" published in the ancient history of 2008 ( https://www.cs.cornell.edu/~shmat/shmat_oak08netflix.pdf ): " We apply our de-anonymization methodology to the Netflix Prize dataset, which contains anonymous movie…
Well said.
Re: Large-Scale Online Deanonymization with LLMs
#206Earlier quoted context omitted.
> I operate by a simple, general rule - I don't often say anything online I wouldn't say directly to someone's face in real life. More people should keep this same energy. I try to stress this to my kids and it feels like it's falling on deaf ears in regards to my teen. Alas.
I can be a rude prick online sometimes, but I can be in real life too - basically though the reason I do this is I never want it to be some huge surprise IRL if someone sees what I write online and be like, "wow, I didn't know that about him." I'm pretty much what I am online and IRL the same. For some reason this seems to matter for me, at least in the past when people have tried to like, send employers stuff I may…
Because I don't really appreciate flame wars and when that's the case, I like to take some time to find common ground and just have a respectable discussion when possible.
This approach is harder to work irl because those moments are also spontaneous & it does require significantly more discipline to control one's emotion within seconds rather than minutes, but its something that I think I can work upon as well.
But I would say that aside from that, most of my comments are pretty spontaneously written. I frame it as a question of being honest with myself at times, I think I am mostly pretty much the same IRL and online as well.
Another point but such forums also act like a journal to me for my future to read as well. I try to write comments in such sense that in future, I can read them and try to accurately remember what my mind was thinking during the time/days I wrote that comment for self-retrospection as well.
Edit: Although now that I think about it, there are definitely some subtle changes I might have online vs irl but I would still say that I feel like my accounts are pretty authentic fwiw (personally) but I am happy with my authenticity online but there's definitely a level of my thinking which worries about any comment being permanently available though.
Re: Large-Scale Online Deanonymization with LLMs
#207I post under my real name here, pretty much the only place I post. It keeps me honest and straight in what I say when I choose to say it. I tried talking to my children about leaving as clean of a footprint on the internet as one can in anticipation of future people/systems taking that into consideration. I don't know what it will be but I would expect some adversarial stuff. Trying to keep clean is what I'd prefer f…
Do you want culture to be frozen and instant digital communication with anyone else in the world to become a privilege of the few? Because that's where "clean" leads. And all you get is a little bit of temporary safety. Here's a different vision for the future: Let information filtering become each individual's own responsibility. We have LLMs now, and they'll get more efficient, so why not use them locally to filter…
Re: Large-Scale Online Deanonymization with LLMs
#208Earlier quoted context omitted.
I view posting online with a real name like getting a permanent tattoo. My values or priorities may significantly change over decades, especially as a child, so why would I want to jeopardize the reputation of a potential future identity with something I may post today?
One could just as easily make the opposite argument. Given that your values and priorities may change significantly over the decades, a smart investment now into a solid, stable, and prosocial public identity may reap considerable and wide-ranging benefits in ways you couldn't even predict. This is especially true if you take seriously the idea that it's not what you say but how you say it that matters in the end.
In the sense that if I ever create any business/idea which can be serious enough that I want to back it up. I might create hackernews post about it.
Although that being said, I do sometimes make alts just to publish something if I don't want it under this particular account.
I do feel like I can be wrong, I usually am[0] but I think that I want to improve myself and perhaps this account can be a way for people to see me grow perhaps and sometimes fall as well. Life feels like a sin wave with ups and downs.
I have had some paranoid thoughts as to what if I get into controversy later on in life because of some things I do in my teen years but there was a line from a friend that I heard which said, "that anyone with more than 1 brain cell can figure out if a person has improved or not"
I do feel like authenticity is gonna be the differentiator if both code and infra aren't the bottlenecks. Perhaps authenticity can be treated as part of marketing but I feel like its also paradoxical to gain authenticity if you want to do marketing. Imo, a person has to be authentic for the sake of being authentic and only then and then can he also get some marketing benefits.
Authenticity means to share both good and bad (well as much as you can, I don't think one should be completely 100% authentic but rather only keep a few personal things to oneselves and even if they get leaked, then y'know just have the grace to accept it and considering that quote from above, I think most people will understand most things especially when you realize that there are people / (youtubers?) in the world who are part of serious accusations/controversies where I feel like most other controversies should be pretty non-issue fwiw.
Like my idea is being authentic enough to satisfy myself. If I become more authentic but if I feel unsatisfied/worried etc.,then that's wrong too.
[0]: (This is such a good quote from how to win friends that I use it quite often)
Re: Large-Scale Online Deanonymization with LLMs
#209This is exactly why local inference matters. Every query you send to a cloud API is another data point. Your prompts contain your code, your logs, your thought process — arguably more identifying than your HN comments. The paper shows deanonymization from public posts. Imagine what's possible with private API traffic: the questions you ask, the code you paste, the errors you debug. Even if providers don't read it tod…
> Air-gapped local inference isn't paranoia. It's necessary.
I definitely agree, I am seeing new model like qwen-3.5-30A3b (iirc) being able to be run reasonably on normal hardware (You can buy a mac mini whose price hasn't been inflated) and get decent tps while having a decent model overall.
There are some services like proton lumo, the service by signal, kagi's AI which seem to try to be better but long term, my plan is to buy mac-mini for such levels of inference for basic queries.
Of course, in the meanwhile like for example coding, it might not make too big of a difference between using local model or not unless for the most extremely sensitive work (perhaps govt/bank oriented)
Re: Large-Scale Online Deanonymization with LLMs
#210Earlier quoted context omitted.
This is very import: you don't know how the cancelation culture will be in 20 years. I like to use the example of a guy who did a blackface in a party back in 2000's. Although reprehensible, was not commom-sense racism back then. Today society sees it as completely unacceptable. Eventually that guy became prime minister of Canada and things went pretty bad when that photo surfaced decades later. Is it far to judge so…
Only idiots don’t care about it historical context.