Nepenthes is a tarpit to catch AI web crawlers
41–50 of 290 posts
Re: Nepenthes is a tarpit to catch AI web crawlers
#42We finally have a viable mouse trap for LLM scrapers for them to continuously scrape garbage forever, depleting the host of their resources whilst the LLM is fed garbage which the result will be unusable to the trainer, accelerating model collapse.
It is like a never ending fast food restaurant for LLMs forced to eat garbage input and will destroy the quality of the model when used later.
Hope to see this sort of defense used widely to protect websites from LLM scrapers.
Re: Nepenthes is a tarpit to catch AI web crawlers
#43There are already “infinite” websites like these on the Internet. Crawlers (both AI and regular search) have a set number of pages they want to crawl per domain. This number is usually determined by the popularity of the domain. Unknown websites will get very few crawls per day whereas popular sites millions. Source: I am the CEO of SerpApi.
> There are already “infinite” websites like these on the Internet. Cool. And how much of the software driving these websites is FOSS and I can download and run it for my own (popular enough to be crawled more than daily by multiple scrapers) website?
It’s useless to do this though as all crawlers have a way to handle this. It’s very crawler 101.
Re: Nepenthes is a tarpit to catch AI web crawlers
#44There are already “infinite” websites like these on the Internet. Crawlers (both AI and regular search) have a set number of pages they want to crawl per domain. This number is usually determined by the popularity of the domain. Unknown websites will get very few crawls per day whereas popular sites millions. Source: I am the CEO of SerpApi.
we're hosting some pretty unknown very domain specific sites and are getting hammered by Claude and others who, compared to old-school search engine bots also get caught up in the weeds and request the same pages all over.
They also seem to not care about response time of the page they are fetching, because when they are caught in the weeds and hit some super bad performing edge-cases, they do not seem to throttle at all and continue to request at 30+ requests per second even when a page takes more than a second to be returned.
We can of course handle this and make them go away, but in the end, this behavior will only hurt them both because they will face more and more opposition by web masters and because they are wasting their resources.
For decades, our solution for search engine bots was basically an empty robots.txt and have the bots deal with our sites. Bots behaved reasonably and intelligently enough that this was a working strategy.
Now in light of the current AI bots which from an outsider observer's viewpoint look like they were cobbled together with the least effort possible, this strategy is no longer viable and we would have to resort to provide a meticulously crafted robots.txt to help each hacked-up AI bot individually to not get lost in the weeds.
Or, you know, we just blanket ban them.
Re: Nepenthes is a tarpit to catch AI web crawlers
#45Re: Nepenthes is a tarpit to catch AI web crawlers
#46Re: Nepenthes is a tarpit to catch AI web crawlers
#47I have a very vague concept for this, with a different implementation. Some, uh, sites (forums?) have content that the AI crawlers would like to consume, and, from what I have heard, the crawlers can irresponsibly hammer the traffic of said sites into oblivion. What if, for the sites which are paywalled, the signup, which invariably comes with a long click-through EULA, had a legal trap within it, forbidding ingestio…
That said, I am not a lawyer and this may not be true in all jurisdictions.
Re: Nepenthes is a tarpit to catch AI web crawlers
#48[flagged]
If AI companies want to sue webmasters for that then by all means, they can waste their money and get laughed out of court.
Re: Nepenthes is a tarpit to catch AI web crawlers
#49Re: Nepenthes is a tarpit to catch AI web crawlers
#50Good. We finally have a viable mouse trap for LLM scrapers for them to continuously scrape garbage forever, depleting the host of their resources whilst the LLM is fed garbage which the result will be unusable to the trainer, accelerating model collapse. It is like a never ending fast food restaurant for LLMs forced to eat garbage input and will destroy the quality of the model when used later. Hope to see this sort…
and all of us will benefit from this.