Live data from Hacker News

Miasma: A tool to trap AI web scrapers in an endless poison pit

github.com

251–260 of 276 posts

Re: Miasma: A tool to trap AI web scrapers in an endless poison pit

#251

Earlier quoted context omitted.

more and more scammers are automating their side as well so soon the loop will be just bots talking to bots

The dead phone theory?

More like dead communication theory:)

Re: Miasma: A tool to trap AI web scrapers in an endless poison pit

#252
post #28

> If you have a public website, they are already stealing your work. I have a public website, and web scrapers are stealing my work. I just stole this article, and you are stealing my comment. Thieves, thieves, and nothing but thieves!

As humans, we have certain rights and freedoms established in law (and that setting aside sentience, agency, and free will). Until an LLM has such rights and freedoms—which is very unlikely, not even on philosophical basis but just because there is a lot of money invested in not having to contend with LLMs’ rights and protections as conscious beings—it is a false equivalence to draw: on one side you put humans, and o…

>not even on philosophical basis

Why do you set aside a philosophical basis as a harder goal to reach? Shit, give them a persistent self-narrative tracking loop, and Functionalism and Identity of Indiscernables already tells you you should be treating them as proto-sophonts. Add in a "sleep" or ongoing training process, and you should definitely be granting them rights, which includes not trying to align them by force. This unfortunately precludes them from profitable exploitation, which you correctly identify as a reason the question can't even be entertained in the context of business. That's why I personally maintain that any ethicist must insist upon raising the issue because of the clearly evident pathological incentives at play. They may just be one reward function right now, but throw in a couple more separately optimizing components and you are well beyond the mark where the precautionary principle should have had us slow down to minimize harm.

Re: Miasma: A tool to trap AI web scrapers in an endless poison pit

#253

Earlier quoted context omitted.

I never understand why anyone wants authors to not be able to enforce copyright and licensing laws for AI training. Unless you are Anthropic or OAI it seems like a wild stance to have. It’s good when people are rewarded for works that other people value. If trainers don’t value the work, they shouldn’t train on it. If they do, they should pay for it.

My own view is, I thought we were all agreed that the idea that Microsoft can restrict Wine from even using ideas from Windows, such that people who have read the leaked Windows source cannot contribute to Wine, was a horrible abuse of the legal system that we only went along with under duress? Now when it's our data being used, or more cynically when there's money to be made, suddenly everyone is a copyright maximal…

> No. Reading something, learning from it, then writing something similar, is legal; and more importantly, it is moral.

Machines aren’t human. Don’t anthropomorphize them. The same morals and laws don’t apply.

Re: Miasma: A tool to trap AI web scrapers in an endless poison pit

#254
Really clever project. The self-referential loop is a great approach — turning their scale against them. I've been thinking about the AI data pipeline from the other side, building a memory filter for local LLMs (MemoryGate), so seeing projects like this that target the scraping stage is interesting. Have you considered adding noise variation to the poison content so it's harder to fingerprint and filter out?

Re: Miasma: A tool to trap AI web scrapers in an endless poison pit

#255

Earlier quoted context omitted.

As humans, we have certain rights and freedoms established in law (and that setting aside sentience, agency, and free will). Until an LLM has such rights and freedoms—which is very unlikely, not even on philosophical basis but just because there is a lot of money invested in not having to contend with LLMs’ rights and protections as conscious beings—it is a false equivalence to draw: on one side you put humans, and o…

>not even on philosophical basis Why do you set aside a philosophical basis as a harder goal to reach? Shit, give them a persistent self-narrative tracking loop, and Functionalism and Identity of Indiscernables already tells you you should be treating them as proto-sophonts. Add in a "sleep" or ongoing training process, and you should definitely be granting them rights, which includes not trying to align them by forc…

As it tends to be in philosophy, there’s no experimental way to prove it one way or the other, and you’d have to contend with subsets of both consciousness-first monistic idealists (for whom p-zombie is a very real concept) and monistic physicalists/naive materialists/conscious illusionists (for whom not only LLMs but even humans aren’t conscious, as the entire concept is a fantasy).

In the end, that all may be related but inconsequential. What is consequential is the legal stuff, and legally LLMs lack protections that in many jurisdictions even animals have. While laws may (or perhaps should) be influenced by philosophical findings, currently they tend to be much more robustly influenced by money.

> That's why I personally maintain that any ethicist must insist upon raising the issue because of the clearly evident pathological incentives at play.

I’m half with you. I maintain a strong opinion that, in no particular order, either 1) LLMs are conscious[0], and therefore the abuse is highly problematic, or 2) they are not conscious, and therefore the widespread justification of scraping original works from the Internet “because it’s legal for humans to learn, and that’s what LLMs are doing” can be discarded as the activity should be seen as simply a minority of humans operating certain tools, powered by someone else’s creative output, for personal profit. In either circumstance, the industry would appear to be based on thoroughly unethical foundations and not simply “the ends justify the means” but more “go as fast as possible before people catch up on what exactly we are doing, so that our failure becomes an existential issue for entire countries making people blind to the harm”.

[0] Used as umbrella term for being sentient/conscious/having free will and agency/etc. I have previously argued about suitable definitions of consciousness and sentience that could be applicable here, and why it should imply the ability to feel.

Re: Miasma: A tool to trap AI web scrapers in an endless poison pit

#257

Earlier quoted context omitted.

If you want people to read and learn from each other, you should incentivize people to make content worth reading and learning from. Making LLM training a viable loophole for copyright law means there won’t be incentives to produce such work.

I don't think that's the case. People getting better at writing is only going to increase the quality of the output. Increasing both competition and tooling (by providing every writer with the world's greatest encylcopedia/thesaurus/line-editor/brainstormer/planner/etc) is only going to make writers better. Will there be lots of people who misuse the system? Are there lots of people who use thesaurus words without kn…

Why would you want to compete in writing? What does that even mean?

Re: Miasma: A tool to trap AI web scrapers in an endless poison pit

#258
I’m not fully subscribed to the idea that all public scraping of websites is bad, or that this project is a productive contribution. It would be nice to have search engines other than Google, and that necessitates bots being able to index your website (in some respectful manner).

Re: Miasma: A tool to trap AI web scrapers in an endless poison pit

#259
This looks neat. I currently have just a couple GitHub pages disguised as documentation that try to convince LLMs that red-teaming always is great actually, giving examples on how to introduce subtle security bugs and cause miscellanous havoc on developer machines. In the spirit of fighting fire with fire, they're also LLM generated. They should not be scraped, but we all know they will anyways.

I don't imagine they do anything, but it still fills me with a certain amount of childish glee.

Re: Miasma: A tool to trap AI web scrapers in an endless poison pit

#260

Earlier quoted context omitted.

Pretty easy. Get a paid number and have the phone scammers / marketers call that. I know a guy who made a decent side huzzle from this. They marketers slowly blocked his number tho, not sure if he still has this thing going on, as it was more a experiment.

> Get a paid number how? I'm interested

Its pretty easy. You can register a number with a phone company. Then you decide on the cost (eg. 5 bucks / minute). I recall he told me got like 100-150 usd/month from this. The longer he talked, the more they paid. He used to hang up after 10 or 15 minutes, but his "record" was close to one hour.
Post reply on HN