Live data from Hacker News

Miasma: A tool to trap AI web scrapers in an endless poison pit

github.com

241–250 of 276 posts

Re: Miasma: A tool to trap AI web scrapers in an endless poison pit

#243

Earlier quoted context omitted.

You're trying to use a quite unfunny "sarcasm" to move the goalpost to the strawman (they never claimed they handcrafted these pages) and quickly gloss ove the fact it's 20 years of work so why not?

You're ascribing an adversarial attitude to me which is actually held by nobody except yourself. The question was genuine and out of curiosity, and they can answer for themselves, however they choose. From the posting guidelines: > Be kind. Don't be snarky. Converse curiously; don't cross-examine. Edit out swipes. > Please respond to the strongest plausible interpretation of what someone says, not a weaker one that's…

So why did you want to know if my 6M pages were handwritten and why was the method of production relevant here exactly?

Re: Miasma: A tool to trap AI web scrapers in an endless poison pit

#244

Earlier quoted context omitted.

My own view is, I thought we were all agreed that the idea that Microsoft can restrict Wine from even using ideas from Windows, such that people who have read the leaked Windows source cannot contribute to Wine, was a horrible abuse of the legal system that we only went along with under duress? Now when it's our data being used, or more cynically when there's money to be made, suddenly everyone is a copyright maximal…

Re-reading your comment, I think we’re both generally anti-corporate-fuckery. I view the current batch of copyright pearl clutching to be an argument about if VCs are allowed to steal books to make their chatbots worth talking to, and the Wine/MSoft debate about if it should be legal to engage in anticompetitive behavior by restrictive use of copyright. In both of these cases the root of the issue isn’t really the co…

I agree that's bad at any rate. However, I genuinely think that reading and learning without literal reproduction is not (should not be) a violation of copyright and does not (should not) require an additional grant for content that has been made publicly available. I think that regardless of whether a company is the subject or the actor.

Re: Miasma: A tool to trap AI web scrapers in an endless poison pit

#245
I love the idea but this will only end up harming your SME in the long run. It would also further entrench the large corps.

The only way something like this would be remotely plausible as a concept would be for enough data providers with overlapping authority on given topics to implement it.

Sadly SMEs have no choice but to go with the flow and allow AI scrapers in. If they don’t, they won’t be as visible in AI generations at the top of the SERPs and they won’t get the visits, which will mean they don’t make the money required to stay afloat.

The fish that attempts to swim against the current ultimately dies and has its corpse carried where the current was going, anyway. Without the sway which comes with size your only option is to go with the flow and drop a little dirty protest every now and then.

Re: Miasma: A tool to trap AI web scrapers in an endless poison pit

#246

Earlier quoted context omitted.

>I dunno... it feels like the same approach as those people who tell you gleeful stories of how they kept a phone spammer on a call for 45 minutes: "That'll teach 'em, ha ha!" Do these types of techniques really work? I’m not convinced If you are automating it, I don't see why not. Kitboga, a you-tuber kept scam callers in AI call-center loops tying up there resources so they cant use them on unsuspecting victims.[0]…

Pretty easy. Get a paid number and have the phone scammers / marketers call that. I know a guy who made a decent side huzzle from this. They marketers slowly blocked his number tho, not sure if he still has this thing going on, as it was more a experiment.

> Get a paid number

how? I'm interested

Re: Miasma: A tool to trap AI web scrapers in an endless poison pit

#247
post #123
post #92

Earlier quoted context omitted.

One would assume legit spiders obey robots.txt.

This, to me, is the strongest argument to offer these slop generators. It provides an incentive to follow the robots.txt.

Exactly. You disobey robots file => we'll make your crawl gain a net negative.

Re: Miasma: A tool to trap AI web scrapers in an endless poison pit

#248
post #91

Earlier quoted context omitted.

About two years ago, I made up reference to a nonexistent python library and put code "using" it in just 5 GitHub repos. Several months later the free ChatGPT picked it up. So IMO it works.

Via websearch? Or training?

[deleted]

Re: Miasma: A tool to trap AI web scrapers in an endless poison pit

#249

Earlier quoted context omitted.

Pretty easy. Get a paid number and have the phone scammers / marketers call that. I know a guy who made a decent side huzzle from this. They marketers slowly blocked his number tho, not sure if he still has this thing going on, as it was more a experiment.

Was he picking up the phone and telling them to call him back on the other number?

IIRC he did something like that, ask them to call back in "10 minutes after my meeting, and call my personal number, not my corporate phone, as it is tracked". On other occasions he filled in this number to some online forms that he "was asked to fill before continuing".

Re: Miasma: A tool to trap AI web scrapers in an endless poison pit

#250
post #138

Earlier quoted context omitted.

Been there recently. Rate limit on nginx and anti-syn flood on pf solved it.

I'm being hit with 300 req/s 24/7 from hundreds of thousands of unique IP's from residential proxies. I can't rate limit any further without hurting the real users.

Yeah real users are just as hosed as the small sites. I get blocked simply because of the netblock I browse from.
Post reply on HN