[flagged]
This is a really bad take, it's not like this server is hacking clients which connect to it. It's providing perfectly valid HTTP responses that just happen to be slow and full of markov gibberish, any harm which comes of that is self inflicted by assuming that websites must provide valuable data as a matter of course. If AI companies want to sue webmasters for that then by all means, they can waste their money and ge…
Nepenthes is a tarpit to catch AI web crawlers
51–60 of 290 posts
Re: Nepenthes is a tarpit to catch AI web crawlers
#52Unless this concept becomes a mass phenomenon with many implementations, isn’t this pretty easy to filter out? And furthermore, since this antagonizes billion-dollar companies that can spin up teams doing nothing but browse Github and HN for software like this to prevent polluting their datalakes, I wonder whether this is a very efficient approach.
It would be more efficient for them to spin up a team to study this robots.txt thing. They've ignored that low hanging fruit, so they won't do the more sophisticated thing any time soon.
Re: Nepenthes is a tarpit to catch AI web crawlers
#53Bug, or feature, this? Could be a way to keep your site public yet unfindable.
Re: Nepenthes is a tarpit to catch AI web crawlers
#54Basically a single HTTP Request to ChatGPT API can trigger 5000 HTTP requests by ChatGPT crawler to a website.
The vulnerability is/was thoroughly ignored by OpenAI/Microsoft/BugCrowd but I really wonder what would happen when ChatGPT crawler interacts with this tarpit several times per second. As ChatGPT crawler is using various Azure IP ranges I actually think the tarpit would crash first.
The vulnerability reporting experience with OpenAI / BugCrowd was really horrific. It's always difficult to get attention for DOS/DDOS vulnerabilities and companies always act like they are not a problem. But if their system goes dark and the CEO calls then suddenly they accept it as a security vulnerability.
I spent a week trying to reach OpenAI/Microsoft to get this fixed, but I gave up and just published the writeup.
I don't recommend you to exploit this vulnerability due to legal reasons.
[1] https://github.com/bf/security-advisories/blob/main/2025-01-...
Re: Nepenthes is a tarpit to catch AI web crawlers
#55Earlier quoted context omitted.
This is a really bad take, it's not like this server is hacking clients which connect to it. It's providing perfectly valid HTTP responses that just happen to be slow and full of markov gibberish, any harm which comes of that is self inflicted by assuming that websites must provide valuable data as a matter of course. If AI companies want to sue webmasters for that then by all means, they can waste their money and ge…
[flagged]
Re: Nepenthes is a tarpit to catch AI web crawlers
#56Re: Nepenthes is a tarpit to catch AI web crawlers
#57[flagged]
This is a really bad take, it's not like this server is hacking clients which connect to it. It's providing perfectly valid HTTP responses that just happen to be slow and full of markov gibberish, any harm which comes of that is self inflicted by assuming that websites must provide valuable data as a matter of course. If AI companies want to sue webmasters for that then by all means, they can waste their money and ge…
heads i win, tails you lose. we own all your content, and you better behave.
i can bet this is incentive-speak.
Re: Nepenthes is a tarpit to catch AI web crawlers
#58[flagged]
Re: Nepenthes is a tarpit to catch AI web crawlers
#59Earlier quoted context omitted.
This is a really bad take, it's not like this server is hacking clients which connect to it. It's providing perfectly valid HTTP responses that just happen to be slow and full of markov gibberish, any harm which comes of that is self inflicted by assuming that websites must provide valuable data as a matter of course. If AI companies want to sue webmasters for that then by all means, they can waste their money and ge…
[flagged]
> You can choose to gatekeep your content, and by doing so, make it unscrapeable, and legally protected.
so... robots.txt, which the AI parasites ignore?
> Also, consider that relatively small, cheap llms are able to parse the difference between meaningful content and Markovian jabber such as this software produces.
okay, so it's not damaging, and there you've refuted your entire argument