Wikipedia is struggling with voracious AI bot crawlers
31–40 of 105 posts
Re: Wikipedia is struggling with voracious AI bot crawlers
#32I thought all of Wikipedia can be downloaded directly if that's the goal? [0] Why scrape? [0] https://en.wikipedia.org/wiki/Wikipedia:Database_download
Re: Wikipedia is struggling with voracious AI bot crawlers
#33People got to make bots pay. That's the only way to get rid of this world wide DDOSing backed up by multi billions companies. There are captcha to block bots or at least make them pay money to solve them, some people in Linux community also made tools to combat that, i think something that use a little cpu energy. And in the same time, you offer an api, less expensive than the cost to crawl it, and everyone win. Mult…
Re: Wikipedia is struggling with voracious AI bot crawlers
#34This has to be one of strangest targets to crawl, since they themselves make database dumps available for download ( https://en.wikipedia.org/wiki/Wikipedia:Database_download ) and if that wasn't enough, there are 3rd party dumps as well ( https://library.kiwix.org/#lang=eng&category=wikipedia ) that you could use if the official ones aren't good enough for some reason. Why would you crawl the web interface when the…
I've written some unfathomably bad web crawlers in the past. Indeed, web crawlers might be the most natural magnet for bad coding and eye-twitchingly questionable architectural practices I know of. While it likely isn't the major factor here I can attest that there are coders who see pages-articles-multistream.xml.bz2 and then reach for a wget + HTML parser combo. If you don't live and breath Wikipedia it is going to…
Re: Wikipedia is struggling with voracious AI bot crawlers
#35This has to be one of strangest targets to crawl, since they themselves make database dumps available for download ( https://en.wikipedia.org/wiki/Wikipedia:Database_download ) and if that wasn't enough, there are 3rd party dumps as well ( https://library.kiwix.org/#lang=eng&category=wikipedia ) that you could use if the official ones aren't good enough for some reason. Why would you crawl the web interface when the…
I've written some unfathomably bad web crawlers in the past. Indeed, web crawlers might be the most natural magnet for bad coding and eye-twitchingly questionable architectural practices I know of. While it likely isn't the major factor here I can attest that there are coders who see pages-articles-multistream.xml.bz2 and then reach for a wget + HTML parser combo. If you don't live and breath Wikipedia it is going to…
Re: Wikipedia is struggling with voracious AI bot crawlers
#36I thought all of Wikipedia can be downloaded directly if that's the goal? [0] Why scrape? [0] https://en.wikipedia.org/wiki/Wikipedia:Database_download
Re: Wikipedia is struggling with voracious AI bot crawlers
#37I also like Anna's (Creative Commons) framing of the problem being money + attribution + reciprocity.
Re: Wikipedia is struggling with voracious AI bot crawlers
#38Re: Wikipedia is struggling with voracious AI bot crawlers
#39Wikipedia spends 1% of its budget on hosting fees. It can spend a bit more given the rest of their corruptions.
So perhaps price comes with the greatness?
Re: Wikipedia is struggling with voracious AI bot crawlers
#40Wikipedia spends 1% of its budget on hosting fees. It can spend a bit more given the rest of their corruptions.