This has to be one of strangest targets to crawl, since they themselves make database dumps available for download ( https://en.wikipedia.org/wiki/Wikipedia:Database_download ) and if that wasn't enough, there are 3rd party dumps as well ( https://library.kiwix.org/#lang=eng&category=wikipedia ) that you could use if the official ones aren't good enough for some reason. Why would you crawl the web interface when the…
Wikipedia is struggling with voracious AI bot crawlers
51–60 of 105 posts
Re: Wikipedia is struggling with voracious AI bot crawlers
#52The bigger problem is that the LLMs are so good that their users no longer feel the need to visit these sites directly. It looks like the business model of most of our clients is becoming obsolete. My paycheck is downstream of that, and I don't see a fix for it.
Re: Wikipedia is struggling with voracious AI bot crawlers
#53Re: Wikipedia is struggling with voracious AI bot crawlers
#54Re: Wikipedia is struggling with voracious AI bot crawlers
#55what's the best way to stop the bots? cloudflare?
Why should we stop the bots? Wikipedia supposedly wants the world to have this free information, a bit is just another way of supporting that goal.
Re: Wikipedia is struggling with voracious AI bot crawlers
#56Wikipedia spends 1% of its budget on hosting fees. It can spend a bit more given the rest of their corruptions.
BTW, I do not know why you are getting downvoted, this is a real concern that someone needs to tackle one day.
Previous discussion: (2022) https://news.ycombinator.com/item?id=32840097
Re: Wikipedia is struggling with voracious AI bot crawlers
#57Earlier quoted context omitted.
This is an interesting model in general: free for humans, pay for automation. How do you enforce that though? Captchas sounds like a waste.
Why not just rate limit every user to realistic human rates. You just punish anyone behaving like a bot.
Re: Wikipedia is struggling with voracious AI bot crawlers
#58This has to be one of strangest targets to crawl, since they themselves make database dumps available for download ( https://en.wikipedia.org/wiki/Wikipedia:Database_download ) and if that wasn't enough, there are 3rd party dumps as well ( https://library.kiwix.org/#lang=eng&category=wikipedia ) that you could use if the official ones aren't good enough for some reason. Why would you crawl the web interface when the…
Re: Wikipedia is struggling with voracious AI bot crawlers
#59what's the best way to stop the bots? cloudflare?
my guess is the gnome anime girl anti bot captcha
It's an interesting project, I wish there would be better ways to do that, but I guess we are on war with crawlers for a while already.
Re: Wikipedia is struggling with voracious AI bot crawlers
#60This has to be one of strangest targets to crawl, since they themselves make database dumps available for download ( https://en.wikipedia.org/wiki/Wikipedia:Database_download ) and if that wasn't enough, there are 3rd party dumps as well ( https://library.kiwix.org/#lang=eng&category=wikipedia ) that you could use if the official ones aren't good enough for some reason. Why would you crawl the web interface when the…
This is what you get when an AI generates your code and your prompts are vague.