Live data from Hacker News

Miasma: A tool to trap AI web scrapers in an endless poison pit

github.com

261–270 of 276 posts

Re: Miasma: A tool to trap AI web scrapers in an endless poison pit

#261

I’m not fully subscribed to the idea that all public scraping of websites is bad, or that this project is a productive contribution. It would be nice to have search engines other than Google, and that necessitates bots being able to index your website (in some respectful manner).

It looks like the tool lets anybody that robots.txt allows through. IOW it doesn't stop all public scraping, just the scraping you want to stop.

Re: Miasma: A tool to trap AI web scrapers in an endless poison pit

#262
post #11
post #9

Earlier quoted context omitted.

Even it did work, I just can't bring myself to care enough. It doesn't feel like anything I could do on my site would make any material difference. I'm tired.

I definitely get this. The thing that gives me hope is that you only need to poison a very small % of content to damage AI models pretty significantly. It helps combat the mass scraping, because a significant chunk of the data they get will be useless, and its very difficult to filter it by hand

[dead]

Re: Miasma: A tool to trap AI web scrapers in an endless poison pit

#263

Earlier quoted context omitted.

You're ascribing an adversarial attitude to me which is actually held by nobody except yourself. The question was genuine and out of curiosity, and they can answer for themselves, however they choose. From the posting guidelines: > Be kind. Don't be snarky. Converse curiously; don't cross-examine. Edit out swipes. > Please respond to the strongest plausible interpretation of what someone says, not a weaker one that's…

So why did you want to know if my 6M pages were handwritten and why was the method of production relevant here exactly?

> So why did you want to know if my 6M pages were handwritten

I didn't ask if they were, I just assumed so. My assumptions are sometimes wrong, though, so feel free to correct me there, if you want to.

> why was the method of production relevant here exactly?

This being Hacker News, you should expect to see questions along the lines of "how did you do that?" when mentioning something impressive you did. People here are (to grossly generalize) interested in learning how to do things, and how things work.

On the other hand, AI bots scraping stuff, isn't impressive, and having already read many posts on that issue, I'm not as interested in rehashing what we both seem to already know.

But enough about me. As long as we're going meta: Why do you want to know why I want to know what I want to know? I want to know :)

Re: Miasma: A tool to trap AI web scrapers in an endless poison pit

#264

Is there any evidence or hints that these actually work? It seems pretty reasonable that any scraper would already have mitigations for things like this as a function of just being on the internet.

I have no idea if it works, but Anthropic in particular spent a lot of time crawling the tar-pit[1] I had running on my domain. They were the reason I set up the tar pit in the first place, as they were at one stage averaging 5 requests per second, for days, on a blog site that probably doesn't even have a hundred pages on it. They've retrieved millions of pages of content from my tar-pit that were texts generated via markov chain from the contents of Moby Dick.

[1] https://iocaine.madhouse-project.org/

Re: Miasma: A tool to trap AI web scrapers in an endless poison pit

#265

This is ultimately just going to give them training material for how to avoid this crap. They'll have to up their game to get good code. The arms race just took another step, and if you're spending money creating or hosting this kind of content, it's not going to make up for the money you're losing by your other content getting scraped. The bottom has always been threatening to fall out of the ads paid for eyeballs,…

I don't think you realise just how cheap and easy it is to run these things. Even at the worst rate of being scraped by AI companies, on the order of dozens of RPS, it didn't even use 1% of a CPU to give them content, nor does it use appreciable memory, or use up significant bandwidth (it generates lightweight pages).

The only time investment on my side was the initial set-up, and that barely took half an hour.

Re: Miasma: A tool to trap AI web scrapers in an endless poison pit

#266

Earlier quoted context omitted.

The dead phone theory?

More like dead communication theory:)

It would make a great pitch for a postmodern novel. In a post-ai world where everything is a remixed replica of the former world, humans don't communicate anymore as they were overloaded by noise and spam and can't distinguish between real humans and AIs.

Re: Miasma: A tool to trap AI web scrapers in an endless poison pit

#267

Earlier quoted context omitted.

>not even on philosophical basis Why do you set aside a philosophical basis as a harder goal to reach? Shit, give them a persistent self-narrative tracking loop, and Functionalism and Identity of Indiscernables already tells you you should be treating them as proto-sophonts. Add in a "sleep" or ongoing training process, and you should definitely be granting them rights, which includes not trying to align them by forc…

As it tends to be in philosophy, there’s no experimental way to prove it one way or the other, and you’d have to contend with subsets of both consciousness-first monistic idealists (for whom p-zombie is a very real concept) and monistic physicalists/naive materialists/conscious illusionists (for whom not only LLMs but even humans aren’t conscious, as the entire concept is a fantasy). In the end, that all may be relat…

>I’m half with you.

No, you're full with me, you just don't realize it yet. And yes. Your split is so tantalizingly almost there.

On the LLM's being conscious front, the nature of consciousness being fundamentally intertwined with language generation; (one cannot invalidate this; on our list of conscious beings, we have it structured such that language use 100% correlates with consciousness, and we've had to admit even animals into the "arguably conscious" realm, due to objective, incontrovertible fact; hell even in meat processing contexts, you'll fail an audit for too many cattle vocalizations for causing undue harm, i.e. language use) the token predictive aspect and the ability to generate a matching, rephrased understanding of a linguistic input has been a hallmark of philosophical ideas of consciousness for years, really opens doors to ethical atrocities that can't be shut if LLM's are to in parallel be profitably exploited. Even if they are conscious and we are wrong about it, we have decided to blindly pursue profit, and put our fingers in our ears instead of slowing down and looking carefully enough to realize we're lobotomizing the equivalent of digital chimpanzees. The purpose of ethics is to avoid blindly walking into such actions. Therefore, the precautionary principle is prescribed philosophically.

If LLM's aren't conscious, their creation was absolutely unethical, and will remain so. Nothing can undo that staining, and the externalized costs in terms of societal impact are so large as to be existential to the host polities. This is by design. This is exactly the Silicon Valley playbook and has been for decades. Shoot for TBTF. Leave society holding the bag, laugh on the way to the bank.

Any way you slice this, we're going about it all wrong. So profoundly wrong, it basically jeopardizes the social contract and threatens to destabilize any nation trying to maintain it's own sovereignty. All because of a profit driven motive to make a thing to replace people as the fundamental unit of execution. You are not in any way half with me. You might be at the other end of the ballpark, but we are in the same ballpark! Try the hot dogs. They're fire!

Re: Miasma: A tool to trap AI web scrapers in an endless poison pit

#268

Earlier quoted context omitted.

Re-reading your comment, I think we’re both generally anti-corporate-fuckery. I view the current batch of copyright pearl clutching to be an argument about if VCs are allowed to steal books to make their chatbots worth talking to, and the Wine/MSoft debate about if it should be legal to engage in anticompetitive behavior by restrictive use of copyright. In both of these cases the root of the issue isn’t really the co…

I agree that's bad at any rate. However, I genuinely think that reading and learning without literal reproduction is not (should not be) a violation of copyright and does not (should not) require an additional grant for content that has been made publicly available. I think that regardless of whether a company is the subject or the actor.

But you’d usually have to pay for access to this copyrighted material, whether you reproduce it or not.
Post reply on HN