Live data from Hacker News

“Should this even be released?” Deep learning tool that may be used for doxxing

github.com

41–50 of 64 posts

Re: “Should this even be released?” Deep learning tool that may be used for doxxing

#41
I'm curious how you can know the tool is 95% accurate, if it's being tested on real world data, such as from reddit etc.?

I can only assume it was tested on a synthetic dataset perhaps.

Also I'm wondering how many unique users are present in the dataset, along with the volume of content for each user.

Re: “Should this even be released?” Deep learning tool that may be used for doxxing

#43
post #14

Earlier quoted context omitted.

>"2. Being able to identify troublemakers in a community (such as a forum or a social networking site). I'm sure a lot of administrators would love to know if that suspicious looking new guy is the alias of a banned troll from a few weeks back (posting through a proxy server)." Orwellian.

So, what do you when a troll persistently and constantly attacks your community website? As in, flames everyone to a crisp, posts as much porn as possible, tries to incite a civil war between a few members that might not be on good terms with each other or the staff and registers hundreds of accounts, some of which stay semi dormant until they strike? Because that can happen very easily online, especially if you get…

ISTM that any site that is capable of controlling spam, is also capable of dealing with trolls. Sure spam is more repetitive than trolls, so it's easier to notice automatically, but there is also much more spam, so chances are that at least some will get through. If you have moderators authorized to zap spam, let them zap trollage as well.

Re: “Should this even be released?” Deep learning tool that may be used for doxxing

#44
post #36

Earlier quoted context omitted.

> This entire post and the debate surrounding it, is frankly stupid. What does 95% accuracy even means?? Generally in supervised machine learning a claim of X% accuracy means that when tested on a large dataset for which the correct result is known and that was not part of the training dataset or validation dataset, it classified X% of that dataset correctly. Typically you gather a big dataset of labeled data and the…

Huh... of course I know definition of accuracy, and how its calculated. However even in supervised learning, accuracy is only used in very limited cases such as multi class classification. For a whole bunch of problems including the one being discussed its a poor and in some cases a biased metric. E.g. consider a heavily unbalanced problem 99% positive labels. By predicting all instances with majority label its possi…

I understand this concern. It's like the old 20 Newsgroups data set for testing classification, where supposedly you're distinguishing the topics of conversation between comp.graphics, sci.electronics, talk.politics.misc, and so on...

...but what the most effective classifiers do is memorize the names and signature blocks of people who posted in each newsgroup.

Re: “Should this even be released?” Deep learning tool that may be used for doxxing

#45

I'm curious how you can know the tool is 95% accurate, if it's being tested on real world data, such as from reddit etc.? I can only assume it was tested on a synthetic dataset perhaps. Also I'm wondering how many unique users are present in the dataset, along with the volume of content for each user.

You can easily turn any dataset with labeled authors into a de-anonymization dataset: split each author's writings in half and give them different IDs. Now you know the true answer for every pairwise combination.

Re: “Should this even be released?” Deep learning tool that may be used for doxxing

#47
post #7
post #5

Yeah, it should be released. I mean sure, there are various 'immoral' uses for it (like doxxing), but there are also many good ones. Such as: 1. Working out who wrote a bunch of anonymous reviews on Amazon or other such sites, which could be used to stop fake reviews. You actually mention this usage in your article. 2. Being able to identify troublemakers in a community (such as a forum or a social networking site).…

Well, if it works, it probably only works for a little while. This is because people who don't want to be identified can run it too, and they can change their prose until they can't be identified.

Whats worse, after a while, and enough trolling in different styles, this tool might ban more and more legitimate attempts to communicate.

Re: “Should this even be released?” Deep learning tool that may be used for doxxing

#48
post #45

I'm curious how you can know the tool is 95% accurate, if it's being tested on real world data, such as from reddit etc.? I can only assume it was tested on a synthetic dataset perhaps. Also I'm wondering how many unique users are present in the dataset, along with the volume of content for each user.

You can easily turn any dataset with labeled authors into a de-anonymization dataset: split each author's writings in half and give them different IDs. Now you know the true answer for every pairwise combination.

Yeah, that's a good point. So with that approach you could even get results of accuracy from data from a single social media source at least I guess.

Re: “Should this even be released?” Deep learning tool that may be used for doxxing

#49
post #19
post #3

If it works, prove it by unmasking Satoshi Nakamoto.

This is exactly the type of thing I'm afraid of when I read the link. Being doxxed due to something I did is one thing, but being doxxed due to something I didn't do is something else altogether. Can you imagine how much it would suck if you woke up the next morning and the entire internet is convinced that you're Satoshi Nakamoto or a pedophile due to a false positive from this program? There is no due process and n…

That's already what happens with false rape/pedophile accusations due to America's love of "trial by media". One false rape accusation and your photo is all over the local media. Your life is ruined.

Re: “Should this even be released?” Deep learning tool that may be used for doxxing

#50
post #5

Yeah, it should be released. I mean sure, there are various 'immoral' uses for it (like doxxing), but there are also many good ones. Such as: 1. Working out who wrote a bunch of anonymous reviews on Amazon or other such sites, which could be used to stop fake reviews. You actually mention this usage in your article. 2. Being able to identify troublemakers in a community (such as a forum or a social networking site).…

>"2. Being able to identify troublemakers in a community (such as a forum or a social networking site). I'm sure a lot of administrators would love to know if that suspicious looking new guy is the alias of a banned troll from a few weeks back (posting through a proxy server)." Orwellian.

What about Russian trolls that have infected the internet, trying to manipulate public opinion or just collect data on regular citizens? Or Correct The Record? Or just about any other mass propaganda campaign? This would be very helpful
Post reply on HN