Live data from Hacker News

“Should this even be released?” Deep learning tool that may be used for doxxing

github.com

11–20 of 64 posts

Re: “Should this even be released?” Deep learning tool that may be used for doxxing

#11
This entire post and the debate surrounding it, is frankly stupid. What does 95% accuracy even means?? Consider face recognition, even when there is good gold standard for matching faces (human judgement, since human are good at recognizing faces), determining accuracy of Face recognition algorithms is still challenging (E.g. Megaface challenge). When it comes to a piece of text written by an author its even more difficult. There are several practical problems too, such as how do you distinguish Quotes and copy-pasted paragraphs from rest of the text.

This sounds like a beginner who created a dataset, with a flawed metric. And is now going around claiming 95% accuracy, using "Deep" learning. And equally clueless commenters are hyping it up.

Why stop at claiming 95%, hell even I can create a "dataset" and a "deep learning" algorithm and get 99.9%.

I am not discounting that there are legitimate stylometric analysis methods, which have been peer reviewed. But please lets not hype "Deep learning for doxxing". This just sullies the real progress being made in deep learning.

Re: “Should this even be released?” Deep learning tool that may be used for doxxing

#12
This is valuable code and should be released, but the product's naming and one-line about already betray it's intended purpose (as the author envisions), providing an additional vector of criticism.

In cases like this, I feel erring on the side of being less explicit tends to help. Leave just enough out to let everyone read between the lines, and put the pieces together.

Don't say it's a "Machine learning algorithm to connect anonymous accounts to real names", say it's a 'speech pattern analyzer', or say it 'allows comparison of speech patterns for likelihood of same author'.

Re: “Should this even be released?” Deep learning tool that may be used for doxxing

#13
post #5

Yeah, it should be released. I mean sure, there are various 'immoral' uses for it (like doxxing), but there are also many good ones. Such as: 1. Working out who wrote a bunch of anonymous reviews on Amazon or other such sites, which could be used to stop fake reviews. You actually mention this usage in your article. 2. Being able to identify troublemakers in a community (such as a forum or a social networking site).…

>"2. Being able to identify troublemakers in a community (such as a forum or a social networking site). I'm sure a lot of administrators would love to know if that suspicious looking new guy is the alias of a banned troll from a few weeks back (posting through a proxy server)." Orwellian.

It would certainly run the risk of false positives.

Re: “Should this even be released?” Deep learning tool that may be used for doxxing

#14
post #5

Yeah, it should be released. I mean sure, there are various 'immoral' uses for it (like doxxing), but there are also many good ones. Such as: 1. Working out who wrote a bunch of anonymous reviews on Amazon or other such sites, which could be used to stop fake reviews. You actually mention this usage in your article. 2. Being able to identify troublemakers in a community (such as a forum or a social networking site).…

>"2. Being able to identify troublemakers in a community (such as a forum or a social networking site). I'm sure a lot of administrators would love to know if that suspicious looking new guy is the alias of a banned troll from a few weeks back (posting through a proxy server)." Orwellian.

So, what do you when a troll persistently and constantly attacks your community website?

As in, flames everyone to a crisp, posts as much porn as possible, tries to incite a civil war between a few members that might not be on good terms with each other or the staff and registers hundreds of accounts, some of which stay semi dormant until they strike?

Because that can happen very easily online, especially if you get the ire of someone with a lot of free time and very few morals. Or if your site ends up at war with a troll site/gets raided by 4chan.

Do you avoid the hassle now, or wait until the situation blows up and half the site is now in the middle of it?

Re: “Should this even be released?” Deep learning tool that may be used for doxxing

#16
post #13

Earlier quoted context omitted.

>"2. Being able to identify troublemakers in a community (such as a forum or a social networking site). I'm sure a lot of administrators would love to know if that suspicious looking new guy is the alias of a banned troll from a few weeks back (posting through a proxy server)." Orwellian.

It would certainly run the risk of false positives.

This is a good point though. Even a 5% chance of it misidentifying someone as a troublemaker would be a problem for an online community site, and it's likely the actual chance is a fair bit higher than that (because it's likely not been tested on a particularly large scale in this context).

Re: “Should this even be released?” Deep learning tool that may be used for doxxing

#17
post #5

Yeah, it should be released. I mean sure, there are various 'immoral' uses for it (like doxxing), but there are also many good ones. Such as: 1. Working out who wrote a bunch of anonymous reviews on Amazon or other such sites, which could be used to stop fake reviews. You actually mention this usage in your article. 2. Being able to identify troublemakers in a community (such as a forum or a social networking site).…

Well not necessarily. 1 and 2 could both be a bet negative for humanity just as easily as they could be a positive.

Yeah, I just realised a few negative uses for 1 right now. Like those cases where a business threatens to sue anyone who leaves negative feedback, and hence such a tool could be used to unmask anonymous reviewers giving a honest evaluation of their products/services.

Or maybe an odd case where it turns out the author of a book or creator of a product finds out someone they know in real life left the negative review and physically attacks them or something.

Re: “Should this even be released?” Deep learning tool that may be used for doxxing

#19
post #3

If it works, prove it by unmasking Satoshi Nakamoto.

This is exactly the type of thing I'm afraid of when I read the link. Being doxxed due to something I did is one thing, but being doxxed due to something I didn't do is something else altogether.

Can you imagine how much it would suck if you woke up the next morning and the entire internet is convinced that you're Satoshi Nakamoto or a pedophile due to a false positive from this program? There is no due process and no chance of appeal; your social life is simply over at that point. All because of a 5% chance of a false positive.

Re: “Should this even be released?” Deep learning tool that may be used for doxxing

#20
post #5

Yeah, it should be released. I mean sure, there are various 'immoral' uses for it (like doxxing), but there are also many good ones. Such as: 1. Working out who wrote a bunch of anonymous reviews on Amazon or other such sites, which could be used to stop fake reviews. You actually mention this usage in your article. 2. Being able to identify troublemakers in a community (such as a forum or a social networking site).…

Essentially killing anonymity on the web is not worth any of these. Facebook alone provides a huge data trove of people's real names linked to posts that this machine could parse. Anyone with any form of professional presence on the web would possibly be vulnerable to any comments long enough to be matched on other sites.

Unfortunately if the tech exists, it will be released, but I don't think the positives outweigh the negatives here. The chilling effects in terms of comments alone would be bad.

Post reply on HN