Live data from Hacker News

“Should this even be released?” Deep learning tool that may be used for doxxing

github.com

51–60 of 64 posts

Re: “Should this even be released?” Deep learning tool that may be used for doxxing

#51
post #14

Earlier quoted context omitted.

So, what do you when a troll persistently and constantly attacks your community website? As in, flames everyone to a crisp, posts as much porn as possible, tries to incite a civil war between a few members that might not be on good terms with each other or the staff and registers hundreds of accounts, some of which stay semi dormant until they strike? Because that can happen very easily online, especially if you get…

Ban behavior, not people. You can convert non-productive people to productive people, most of them only want to be noticed or accepted and past a certain point there is only so much you can do to block anyone anyway. Stylometric analysis would just be another simple hurdle to cross for a persistent person. The idea that this tool would be useful for community management is terrible.

You can convert non-productive people to productive people, most of them only want to be noticed or accepted

I'm a moderator of multiple online spaces.

A few months ago, in one of them, a user got too heated and started flinging insults at someone else. As was standard policy for the place where it was occurring, I issued the user a ban of a few days (enforced cooling-off) and pointed to our guidelines on how to behave.

This user then proceeded, over a period of months, to continually harass me, send me increasingly graphic threats, and try to track me down in real life.

Pray tell, how exactly would you go about "converting" such a person to be productive? I come to you since you are apparently quite the expert on it, or else you wouldn't be giving out advice to just "convert" people.

Re: “Should this even be released?” Deep learning tool that may be used for doxxing

#52
post #19
post #3

If it works, prove it by unmasking Satoshi Nakamoto.

This is exactly the type of thing I'm afraid of when I read the link. Being doxxed due to something I did is one thing, but being doxxed due to something I didn't do is something else altogether. Can you imagine how much it would suck if you woke up the next morning and the entire internet is convinced that you're Satoshi Nakamoto or a pedophile due to a false positive from this program? There is no due process and n…

That's among the true terrors of a global information domain society. It's not that you can prove the goods on anyone, it's that you can plausibly assert them. And proving a negative is very, very difficult.

That said, I'm not sure the genie can be rebottled. It's an area of privacy in which law rather than technology must be applied. Including, say, exceptionally fierce penalties for misuse.

Re: “Should this even be released?” Deep learning tool that may be used for doxxing

#53
post #19

Earlier quoted context omitted.

This is exactly the type of thing I'm afraid of when I read the link. Being doxxed due to something I did is one thing, but being doxxed due to something I didn't do is something else altogether. Can you imagine how much it would suck if you woke up the next morning and the entire internet is convinced that you're Satoshi Nakamoto or a pedophile due to a false positive from this program? There is no due process and n…

That's already what happens with false rape/pedophile accusations due to America's love of "trial by media". One false rape accusation and your photo is all over the local media. Your life is ruined.

I've addressed this in part in my own reply to the parent. But the larger one is that with 1) pervasive information and 2) very cheap analysis or assertion, you're hugely increasing the potential for this type of abuse.

The limit is in attention paid -- the public has a limited capacity to absorb information, and there are a few hundred, perhaps a thousand or so "top celebrities" at any one time.

And some of those can attain a highly significant level of immunity to criticism. Ronald Reagan's presidency was the most scandal-prone in recent memory, and yet his moniker was "the Teflon president". William Jefferson Clinton took far more flack for far less, and Barak Obama takes the hit for complete fabrications. Meantime, a major party presidential candidate advocates overt violence to protesters and various other views ... and is only embraced all the more strongly by his supporters.

The dynamics of this are odd.

Re: “Should this even be released?” Deep learning tool that may be used for doxxing

#55
post #45

I'm curious how you can know the tool is 95% accurate, if it's being tested on real world data, such as from reddit etc.? I can only assume it was tested on a synthetic dataset perhaps. Also I'm wondering how many unique users are present in the dataset, along with the volume of content for each user.

You can easily turn any dataset with labeled authors into a de-anonymization dataset: split each author's writings in half and give them different IDs. Now you know the true answer for every pairwise combination.

The problem with that approach is that it is training the classifier to solve a different problem than what is claimed. It assumes that there is no difference between how people write when they are identifiable and when they are not (since it would only train on identifiable samples that are anonymized after the fact). Further, it would require solving the problem of identifying authors between different media--which would be a huge achievement on its own.

A great test set for anyone trying to do this: look at the Scott Adams sock puppet controversy on Metafilter [1] and see if you can train something on his public writing to match the "PlannedChaos" commenter's posts and Adams' own tweets. It is probably the closest you could get to a "pure" training set in the sense that presumably Adams didn't think he'd get caught. (And if he did, and therefore did alter his stylometrics, then it's even a better challenge.)

[1] http://www.adweek.com/galleycat/scott-adams-caught-defending...

Re: “Should this even be released?” Deep learning tool that may be used for doxxing

#57
Doxxing does makes me feel bad. Fixing the problems that led to for example throwaway Ask HN posts is in the long term a better solution, though may be easier said than than done. (Yes, I mean doing the Ask HN posts non-anonymously instead.)

Re: “Should this even be released?” Deep learning tool that may be used for doxxing

#58
post #12

This is valuable code and should be released, but the product's naming and one-line about already betray it's intended purpose (as the author envisions), providing an additional vector of criticism. In cases like this, I feel erring on the side of being less explicit tends to help. Leave just enough out to let everyone read between the lines, and put the pieces together. Don't say it's a "Machine learning algorithm t…

Good point, but personally I think it's a breath of fresh air that not only has the author elected to forego the doublespeak on describing this tool, but they also sparked a discussion on the ethics of using it.

Re: “Should this even be released?” Deep learning tool that may be used for doxxing

#59
post #18

There already exists a counter tool to help against this kind of privacy invasion https://github.com/psal/anonymouth

I just tried to get anonymouth working, but unfortunately even after fixing the invalid code issues/errors preventing compilation, it crashes after you fill out about 6 screens worth questions about where various types of text content is located: >>>>>>>>>>>>>>>>>>>>>>> LOGGING STACK TRACE It's too bad the flow isn't more along the lines of "Give me some docs from one author, and the other document you want to test.…

From this trace it looks like the file jsan_resources/koppel_function_words.txt is missing/empty/invalid.

Re: “Should this even be released?” Deep learning tool that may be used for doxxing

#60

Earlier quoted context omitted.

Ban behavior, not people. You can convert non-productive people to productive people, most of them only want to be noticed or accepted and past a certain point there is only so much you can do to block anyone anyway. Stylometric analysis would just be another simple hurdle to cross for a persistent person. The idea that this tool would be useful for community management is terrible.

You can convert non-productive people to productive people, most of them only want to be noticed or accepted I'm a moderator of multiple online spaces. A few months ago, in one of them, a user got too heated and started flinging insults at someone else. As was standard policy for the place where it was occurring, I issued the user a ban of a few days (enforced cooling-off) and pointed to our guidelines on how to beha…

[deleted]
Post reply on HN