Live data from Hacker News

Nepenthes is a tarpit to catch AI web crawlers

zadzmo.org

51–60 of 290 posts

Re: Nepenthes is a tarpit to catch AI web crawlers

#51
post #48

[flagged]

This is a really bad take, it's not like this server is hacking clients which connect to it. It's providing perfectly valid HTTP responses that just happen to be slow and full of markov gibberish, any harm which comes of that is self inflicted by assuming that websites must provide valuable data as a matter of course. If AI companies want to sue webmasters for that then by all means, they can waste their money and ge…

[flagged]

Re: Nepenthes is a tarpit to catch AI web crawlers

#52
post #13
post #5

Unless this concept becomes a mass phenomenon with many implementations, isn’t this pretty easy to filter out? And furthermore, since this antagonizes billion-dollar companies that can spin up teams doing nothing but browse Github and HN for software like this to prevent polluting their datalakes, I wonder whether this is a very efficient approach.

It would be more efficient for them to spin up a team to study this robots.txt thing. They've ignored that low hanging fruit, so they won't do the more sophisticated thing any time soon.

You can't make money out of studying robots.txt, but you can avoid costs skipping bad web sites.

Re: Nepenthes is a tarpit to catch AI web crawlers

#54
Haha, this would be an amazing way to test the ChatGPT crawler reflective DDOS vulnerability [1] I published last week.

Basically a single HTTP Request to ChatGPT API can trigger 5000 HTTP requests by ChatGPT crawler to a website.

The vulnerability is/was thoroughly ignored by OpenAI/Microsoft/BugCrowd but I really wonder what would happen when ChatGPT crawler interacts with this tarpit several times per second. As ChatGPT crawler is using various Azure IP ranges I actually think the tarpit would crash first.

The vulnerability reporting experience with OpenAI / BugCrowd was really horrific. It's always difficult to get attention for DOS/DDOS vulnerabilities and companies always act like they are not a problem. But if their system goes dark and the CEO calls then suddenly they accept it as a security vulnerability.

I spent a week trying to reach OpenAI/Microsoft to get this fixed, but I gave up and just published the writeup.

I don't recommend you to exploit this vulnerability due to legal reasons.

[1] https://github.com/bf/security-advisories/blob/main/2025-01-...

Re: Nepenthes is a tarpit to catch AI web crawlers

#55
post #48

Earlier quoted context omitted.

This is a really bad take, it's not like this server is hacking clients which connect to it. It's providing perfectly valid HTTP responses that just happen to be slow and full of markov gibberish, any harm which comes of that is self inflicted by assuming that websites must provide valuable data as a matter of course. If AI companies want to sue webmasters for that then by all means, they can waste their money and ge…

[flagged]

[deleted]

Re: Nepenthes is a tarpit to catch AI web crawlers

#57
post #48

[flagged]

This is a really bad take, it's not like this server is hacking clients which connect to it. It's providing perfectly valid HTTP responses that just happen to be slow and full of markov gibberish, any harm which comes of that is self inflicted by assuming that websites must provide valuable data as a matter of course. If AI companies want to sue webmasters for that then by all means, they can waste their money and ge…

yea, it comes across as an extremely entitled mobster take.

heads i win, tails you lose. we own all your content, and you better behave.

i can bet this is incentive-speak.

Re: Nepenthes is a tarpit to catch AI web crawlers

#59
post #48

Earlier quoted context omitted.

This is a really bad take, it's not like this server is hacking clients which connect to it. It's providing perfectly valid HTTP responses that just happen to be slow and full of markov gibberish, any harm which comes of that is self inflicted by assuming that websites must provide valuable data as a matter of course. If AI companies want to sue webmasters for that then by all means, they can waste their money and ge…

[flagged]

> If you want to protect your content, use the technical mechanisms that are available,

> You can choose to gatekeep your content, and by doing so, make it unscrapeable, and legally protected.

so... robots.txt, which the AI parasites ignore?

> Also, consider that relatively small, cheap llms are able to parse the difference between meaningful content and Markovian jabber such as this software produces.

okay, so it's not damaging, and there you've refuted your entire argument

Post reply on HN