Live data from Hacker News

Web Scraping in 2016

franciskim.co

401–402 of 402 posts

Re: Web Scraping in 2016

#401

Earlier quoted context omitted.

You seem to be assuming 1. I'm scraping the resized galleries. 2. I don't have the Hath perk that makes the galleries full sized. 3. I don't have a phash-based fuzzy image deduplication system on top of all this (see https://github.com/fake-name/IntraArchiveDeduplicator ). It's main purpose is to deduplicate manga ( https://github.com/fake-name/MangaCMS ).

Jesus, your projects are massive. Does your job involve working on these or are these just side things?

It's all entirely hobby things.

Re: Web Scraping in 2016

#402
post #396

Earlier quoted context omitted.

That's a separate project: - https://github.com/fake-name/ExHentai-Archival - https://github.com/fake-name/PatreonArchiver - https://github.com/fake-name/xA-Scraper - https://github.com/fake-name/DanbooruScraper Or... well, 4 separate projects. Whoops? At one point, a friend and I were looking at trying to basically replicate the google deep-dream neural net thing, only with a training set of porn. It turns out getti…

Oh my god. Can you share any results?

The project never went anywhere, unfortunately, and I haven't had time to look at it recently.

I have huge, uh, "datasets" around still, though.

Post reply on HN