Live data from Hacker News

1.5 TB of Dark Net Market scrapes

gwern.net

1–10 of 78 posts

Re: 1.5 TB of Dark Net Market scrapes

#3
What amazing work! I am very interested in doing research with Tor and a dataset like this could make my job a heck of a lot easier. I have a legal question though: Are your scrapes text only? Before I work with this dataset, I want to make sure that there's no possibility it contains illegal images (child porn).

Re: 1.5 TB of Dark Net Market scrapes

#5
post #3

What amazing work! I am very interested in doing research with Tor and a dataset like this could make my job a heck of a lot easier. I have a legal question though: Are your scrapes text only? Before I work with this dataset, I want to make sure that there's no possibility it contains illegal images (child porn).

Lower down on the page, he says he did scrape at least one site with such images, although he specifically only took text. Can't verify that this was the case for all scraped sites.

Re: 1.5 TB of Dark Net Market scrapes

#7
post #3

What amazing work! I am very interested in doing research with Tor and a dataset like this could make my job a heck of a lot easier. I have a legal question though: Are your scrapes text only? Before I work with this dataset, I want to make sure that there's no possibility it contains illegal images (child porn).

How about ascii art?

Actually, this is an interesting topic. Poisoning a dataset. CP would work for private security investigators, and to poison against government investigators you could use leaked classified secrets.

Could you work around this by operating on the files on VPS you don't own, streaming a very low-res ('Basilisk'-proof - https://en.wikipedia.org/wiki/BLIT_(short_story) ) remote desktop image.

Re: 1.5 TB of Dark Net Market scrapes

#8
post #7
post #3

What amazing work! I am very interested in doing research with Tor and a dataset like this could make my job a heck of a lot easier. I have a legal question though: Are your scrapes text only? Before I work with this dataset, I want to make sure that there's no possibility it contains illegal images (child porn).

How about ascii art? Actually, this is an interesting topic. Poisoning a dataset. CP would work for private security investigators, and to poison against government investigators you could use leaked classified secrets. Could you work around this by operating on the files on VPS you don't own, streaming a very low-res ('Basilisk'-proof - https://en.wikipedia.org/wiki/BLIT_(short_story) ) remote desktop image.

Possession laws are pretty strict and hard to decode. I wouldn't want to be the test case in court. The idea of "poisoning" a dataset is an interesting theoretical. But in practice, I just want to judge the likelihood that the dataset is poisoned by the presence of images. If it is then there's not much I can do with it.

Re: 1.5 TB of Dark Net Market scrapes

#10
post #8
post #7

Earlier quoted context omitted.

How about ascii art? Actually, this is an interesting topic. Poisoning a dataset. CP would work for private security investigators, and to poison against government investigators you could use leaked classified secrets. Could you work around this by operating on the files on VPS you don't own, streaming a very low-res ('Basilisk'-proof - https://en.wikipedia.org/wiki/BLIT_(short_story) ) remote desktop image.

Possession laws are pretty strict and hard to decode. I wouldn't want to be the test case in court. The idea of "poisoning" a dataset is an interesting theoretical. But in practice, I just want to judge the likelihood that the dataset is poisoned by the presence of images. If it is then there's not much I can do with it.

Yes, this absolutely needs to be clarified by Gwern. This is a very dangerous thing to link researchers to if it contains any illegal content.
Post reply on HN