Live data from Hacker News

Nepenthes is a tarpit to catch AI web crawlers

zadzmo.org

11–20 of 290 posts

Re: Nepenthes is a tarpit to catch AI web crawlers

#11
post #5

Unless this concept becomes a mass phenomenon with many implementations, isn’t this pretty easy to filter out? And furthermore, since this antagonizes billion-dollar companies that can spin up teams doing nothing but browse Github and HN for software like this to prevent polluting their datalakes, I wonder whether this is a very efficient approach.

If it means it makes your own content safe when you deploy it on a corner of your website: mission accomplished!

Re: Nepenthes is a tarpit to catch AI web crawlers

#12
There are already “infinite” websites like these on the Internet.

Crawlers (both AI and regular search) have a set number of pages they want to crawl per domain. This number is usually determined by the popularity of the domain.

Unknown websites will get very few crawls per day whereas popular sites millions.

Source: I am the CEO of SerpApi.

Re: Nepenthes is a tarpit to catch AI web crawlers

#13
post #5

Unless this concept becomes a mass phenomenon with many implementations, isn’t this pretty easy to filter out? And furthermore, since this antagonizes billion-dollar companies that can spin up teams doing nothing but browse Github and HN for software like this to prevent polluting their datalakes, I wonder whether this is a very efficient approach.

It would be more efficient for them to spin up a team to study this robots.txt thing. They've ignored that low hanging fruit, so they won't do the more sophisticated thing any time soon.

Re: Nepenthes is a tarpit to catch AI web crawlers

#14

I have a very vague concept for this, with a different implementation. Some, uh, sites (forums?) have content that the AI crawlers would like to consume, and, from what I have heard, the crawlers can irresponsibly hammer the traffic of said sites into oblivion. What if, for the sites which are paywalled, the signup, which invariably comes with a long click-through EULA, had a legal trap within it, forbidding ingestio…

Legal traps are not a thing.

Re: Nepenthes is a tarpit to catch AI web crawlers

#15
post #3

This keeps generating new pages to keep the crawler occupied. Looks like this would tarpit any web crawler.

It would indeed. Note the warning: "There is not currently a way to differentiate between web crawlers that are indexing sites for search purposes, vs crawlers that are training AI models. ANY SITE THIS SOFTWARE IS APPLIED TO WILL LIKELY DISAPPEAR FROM ALL SEARCH RESULTS."

[deleted]

Re: Nepenthes is a tarpit to catch AI web crawlers

#16
post #4

Earlier quoted context omitted.

It's actually a great idea to spread malware without leaving traces too, it makes content inspection to be very difficult, view-source: to be broken and most of debugging tools, saving to .har, etc.

how is view source broken

It waits for the whole page to load

Re: Nepenthes is a tarpit to catch AI web crawlers

#17

I have a very vague concept for this, with a different implementation. Some, uh, sites (forums?) have content that the AI crawlers would like to consume, and, from what I have heard, the crawlers can irresponsibly hammer the traffic of said sites into oblivion. What if, for the sites which are paywalled, the signup, which invariably comes with a long click-through EULA, had a legal trap within it, forbidding ingestio…

Legal traps are not a thing.

Laws don't apply to billionaires

Re: Nepenthes is a tarpit to catch AI web crawlers

#18

There are already “infinite” websites like these on the Internet. Crawlers (both AI and regular search) have a set number of pages they want to crawl per domain. This number is usually determined by the popularity of the domain. Unknown websites will get very few crawls per day whereas popular sites millions. Source: I am the CEO of SerpApi.

> There are already “infinite” websites like these on the Internet.

Cool. And how much of the software driving these websites is FOSS and I can download and run it for my own (popular enough to be crawled more than daily by multiple scrapers) website?

Re: Nepenthes is a tarpit to catch AI web crawlers

#20

I have a very vague concept for this, with a different implementation. Some, uh, sites (forums?) have content that the AI crawlers would like to consume, and, from what I have heard, the crawlers can irresponsibly hammer the traffic of said sites into oblivion. What if, for the sites which are paywalled, the signup, which invariably comes with a long click-through EULA, had a legal trap within it, forbidding ingestio…

This doesn't work like you think it does but even if it did, do you have the money to sustain several years long legal battle against OpenAI?
Post reply on HN