Earlier quoted context omitted.
The first is incorrect, these scrapers are usually distributed across many IPs, in my experience. I usually refer to them as "disturbed, non-identifying crawlers (DNCs)" when I want to be maximally explicit. (The worst I've seen is some crawler/botnet making exactly one request per IP -_-)
I think the second is incorrect too. DDoS is a DDoS no matter what the intent is.
Miasma: A tool to trap AI web scrapers in an endless poison pit
201–210 of 276 posts
Re: Miasma: A tool to trap AI web scrapers in an endless poison pit
#202I dunno... it feels like the same approach as those people who tell you gleeful stories of how they kept a phone spammer on a call for 45 minutes: "That'll teach 'em, ha ha!" Do these types of techniques really work? I’m not convinced. Also, inserting hidden or misleading links is specifically a no-no for Google Search [0], who have this to say: We detect policy-violating practices both through automated systems and,…
> gleeful stories of how they kept a phone spammer on a call for 45 minutes: "That'll teach 'em, ha ha!" Do these types of techniques really work? I’m not convinced. It’s one of the best time investments I’ve ever made. They just don’t call me anymore. I think they have two lists: the “do not call” list, and the “unprofitable to call” list. You want to be on the latter list.
Re: Miasma: A tool to trap AI web scrapers in an endless poison pit
#203Re: Miasma: A tool to trap AI web scrapers in an endless poison pit
#204Earlier quoted context omitted.
I agree theft isn't a good analogy, but there is something similar going on. I put my words out into the world as a form of sharing. I enjoy reading things others write and share freely, so I write so others might enjoy the things I write. But now the things I write and share freely are being used to put money in the bank accounts of the worst people on the planet. They are using my work in a way I don't want it to b…
>but there is something similar going on [...] No, what you're basically describing is "I shared something but then I didn't like how it ended up being used". If you put stuff out in public for anyone to use, then find out it's used in a way you don't like, it's your right to stop sharing, but it's not "similar" to stealing beyond "I hate stealing"
> If you put stuff out in public for anyone to use, then find out it's used in a way you don't like, it's your right to stop sharing
Yes. The entire point of Copyright and the reason it was invented is to ensure people will keep sharing things. Because otherwise people will just stop publishing things, which is a detriment to all. (Including AI companies, who now don't get new training data)
We have collectively decided that we will give authors some power to say "I don't like how my work is being used" to ensure they don't just "stop sharing".
Fair Use is an exception to that, where the public good does outweigh an individual author's objections. But critically, not such that authors stop publishing. Hence the 4th "factor" in US copyright law (which is one of the most expansive on fair use), where the "effect of the use upon the potential market for or value of the copyrighted work" is evaluated. Fair use isn't supposed to obliterate the value of the original work, or people will stop publishing again.
This is what makes AI training's status so contentious. In terms of direct copyright it is a very weak case. It is incredibly hard to prove a direct 1:1 copy from AI training data into the model and into the output, you have to argue about the architecture of LLMs, and it's incapability of separating copyrightable expressions from uncopyrightable facts.
Yet in spirit, AI training clearly violates copyright. The explicit stated purpose is to copy the works for training data, oft without any compensation or even permission, in order to create a machine that will annihilate the market for all works used.
People already are pulling back on the amount of works they share.
Re: Miasma: A tool to trap AI web scrapers in an endless poison pit
#205Earlier quoted context omitted.
Wow, how did you manually hand-write 6 million web pages? That is impressive. It would take me a while to even montonically count that high.
You're trying to use a quite unfunny "sarcasm" to move the goalpost to the strawman (they never claimed they handcrafted these pages) and quickly gloss ove the fact it's 20 years of work so why not?
> Be kind. Don't be snarky. Converse curiously; don't cross-examine. Edit out swipes.
> Please respond to the strongest plausible interpretation of what someone says, not a weaker one that's easier to criticize. Assume good faith.
I am a friend, not a foe, and so are your other fellow HN posters.
Re: Miasma: A tool to trap AI web scrapers in an endless poison pit
#206> I definitely get this. The thing that gives me hope is that you only need to poison a very small % of content to damage AI models pretty significantly. It helps combat the mass scraping, because a significant chunk of the data they get will be useless, and its very difficult to filter it by hand
It'd be great if the code returned by this project is code that doesn't work. Imagine if all these models are being trained with code that looks OK but in the end it just bullshit. I'd be amazing.
Re: Miasma: A tool to trap AI web scrapers in an endless poison pit
#207Earlier quoted context omitted.
> gleeful stories of how they kept a phone spammer on a call for 45 minutes: "That'll teach 'em, ha ha!" Do these types of techniques really work? I’m not convinced. It’s one of the best time investments I’ve ever made. They just don’t call me anymore. I think they have two lists: the “do not call” list, and the “unprofitable to call” list. You want to be on the latter list.
From my experience, it's the opposite. The more you fuck with them the more they call. It's better not to answer. But I just can't help myself.
Re: Miasma: A tool to trap AI web scrapers in an endless poison pit
#208Earlier quoted context omitted.
Unintentional denial-of-service attacks from AI scrapers are definitely a problem, I just don't know if "theft" is the right way to classify them. They shouldn't get lumped in with intellectual property concerns, which are a different matter. AI scrapers are a tragedy of the commons problem kind of like Kessler syndrome: a few bad actors can ruin low Earth orbit for everyone via space pollution, which is definitely a…
Theft isn't far off, it seems closer to me than using the word for IP violations. When a crawler aggressively crawls your site, they're permanently depriving you the use of those resources for their intended purpose. Arguably, it looks a lot like conversion.
is this why media networks are buying social ai apps
Re: Miasma: A tool to trap AI web scrapers in an endless poison pit
#209Earlier quoted context omitted.
> gleeful stories of how they kept a phone spammer on a call for 45 minutes: "That'll teach 'em, ha ha!" Do these types of techniques really work? I’m not convinced. It’s one of the best time investments I’ve ever made. They just don’t call me anymore. I think they have two lists: the “do not call” list, and the “unprofitable to call” list. You want to be on the latter list.
From my experience, it's the opposite. The more you fuck with them the more they call. It's better not to answer. But I just can't help myself.
Re: Miasma: A tool to trap AI web scrapers in an endless poison pit
#210Earlier quoted context omitted.
>I dunno... it feels like the same approach as those people who tell you gleeful stories of how they kept a phone spammer on a call for 45 minutes: "That'll teach 'em, ha ha!" Do these types of techniques really work? I’m not convinced If you are automating it, I don't see why not. Kitboga, a you-tuber kept scam callers in AI call-center loops tying up there resources so they cant use them on unsuspecting victims.[0]…
more and more scammers are automating their side as well so soon the loop will be just bots talking to bots