Perplexity is using stealth, undeclared crawlers to evade no-crawl directives
611–620 of 799 posts
Re: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives
#612Respone from Perpelexity to Tech Crunch... >Perplexity spokesperson Jesse Dwyer dismissed Cloudflare’s blog post as a “sales pitch,” adding in an email to TechCrunch that the screenshots in the post “show that no content was accessed.” In a follow-up email, Dwyer claimed the bot named in the Cloudflare blog “isn’t even ours.”
Either way, the CDNs profit big time from the AI scraping hype and the current copyright anarchy in the US
Re: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives
#613Earlier quoted context omitted.
I don’t subscribe to technological inevitabilism. Cloudflare banning bad actors has at least made scraping more expensive, and changes the economics of it - more sophisticated deception is necessarily more expensive. If the cost is high enough to force entry, scrapers might be willing to pay for access. But I can imagine more extreme measures. e.g. old web of trust style request signing[0]. I don’t see any easy way f…
> Cloudflare banning bad actors has at least made scraping more expensive, and changes the economics of it - more sophisticated deception is necessarily more expensive. If the cost is high enough to force entry, scrapers might be willing to pay for access. I think this might actually point at the end state. Scraping bots will eventually get good enough to emulate a person well enough to be indistinguishable (are we t…
Interested to see some LLM-adverserial equivalent of MPAA dots![1]
Re: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives
#614Earlier quoted context omitted.
> > We conducted an experiment by querying Perplexity AI with questions about these domains, and discovered Perplexity was still providing detailed information regarding the exact content hosted on each of these restricted domains. This response was unexpected, as we had taken all necessary precautions to prevent this data from being retrievable by their crawlers. Right, I'm confused why CloudFlare is confused. You a…
> You asked the web-enabled AI to look at the domains. Right, and the domain was configured to disallow crawlers, but Perplexity crawled it anyway. I am really struggling to see how this is hard to understand. If you mean to say "I don't think there is anything wrong with ignoring robots.txt" then just say that . Don't pretend they didn't make it clear what they're objecting to, because they spell it out repeatedly.
No, they did not. Crawling = recursive fetching, which wasn't what was happening here.
But also, I don't think there is anything wrong with ignoring robots.txt. In fact, I believe it is discriminatory and people should ignore it. See: https://wiki.archiveteam.org/index.php/Robots.txt
Re: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives
#615Earlier quoted context omitted.
Your comment and the above comment of course show different cases. An agent making a request on the explicit behalf of someone else is probably something most of us agree is reasonable. "What are the current stories on Hacker News?" -- the agent is just doing the same request to the same website that I would have done anyways. But the sort of non-explicit just-in-case crawling that Perplexity might do for a general q…
Hacker news wants you to vist the site, look at the main page, enter threads and participate in discussion. When you swap in an AI and ask what are the current stories. The AI fetches the front page and every thread and feeds it back to you. You are less likely to participate in discussion because you've already had the info summarized.
And yet people install ad blockers and defend their freedom to not participate in this because they don't want to be annoyed by ads.
They claim that since they are free to not buy an advertised product, why would they be forced to see ads for it. But Foo news claims that they are also free to not waste bandwidth to serve their free website to people who declare (by using an ad blocker or the modern alternative: AI aummarizera) they won't participate in the funding of the service
Re: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives
#616I find this problem quite difficult to solve: 1. If I as a human request a website, then I should be shown the content. Everyone agrees. 2. If I as the human request the software on my computer to modify the content before displaying it, for example by installing an ad-blocker into my user agent, then that's my choice and the website should not be notified about it. Most users agree, some websites try to nag you into…
People are usually fine with the latter but not the former, even though they come down to the same thing.
I think this is because people don't want LLMs to train on their content, and they don't differentiate between accessing a website to show it to the user, versus accessing it to train.
Re: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives
#617I wonder if DRM is useful for this. The problem: I want people to access my site, but not Google, not bots, not crawlers and certainly not for use by AI. I don't really know anything about DRM except it is used to take down sites that violate it. Perhaps it is possible for cloudflare (or anyone else) to file a take down notice with Perplexity. That might at least confuse them. Corporations use this to protect their c…
Re: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives
#618Earlier quoted context omitted.
Your comment and the above comment of course show different cases. An agent making a request on the explicit behalf of someone else is probably something most of us agree is reasonable. "What are the current stories on Hacker News?" -- the agent is just doing the same request to the same website that I would have done anyways. But the sort of non-explicit just-in-case crawling that Perplexity might do for a general q…
As a person who has a couple of sites out there, and witnesses AI crawlers coming and fetching pages from these sites, I have a question: What prevents these companies from keeping a copy of that particular page, which I specifically disallowed for bot scraping, and feed it to their next training cycle? Pinky promises? Ethics? Laws? Technical limitations? Leeroy Jenkins?
If you give them a URL that does not appear in Google, ask them to visit that URL specifically, and then notice the content from that URL in the training data, it's proof that they're doing this, which would be quite damaging to them.
Re: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives
#619Earlier quoted context omitted.
My point is that to invoke the "they're DOSing" excuse, you actually have to provide evidence it's happening in this specific instance, rather than vaguely gesturing at some class of entities (AI companies) and concluding that because some AI companies are DOSing, all AI companies are DOSing. Otherwise it's like youtube blocking all youtube-dl users for "DOSing" (some fraction of users arguably are), and then justify…
I tell you of an instance where the biggest ai company is DOS’ing and your reply is that I haven’t proven all of them are doing it? Why do I waste my time on this stuff
Re: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives
#620Earlier quoted context omitted.
Thanks for sharing your experience. A little off-topic but I'd like to start hosting some personal content, guides/tutorials, etc. Do you still see authentic human traffic on your domains, is it easy to discern? I feel like I missed the bus on running a blog pre-AI.
if you do analytics, it is not so hard, but then you need to store user data (if not directly, then worse, with a third party), which should be viewed as a liability. I see ~2/3 human traffic, ~1/3 bot traffic (I just parse user agent strings and count whitelisted browsers as human), but my main landing page is all dynamic-populated webgl. I just asked Gemini what it sees on website, and it states "The page appears t…
> I made a stateful Internet implementation in Python earlier for proof-of-concept
Is there a repo or some other form of public access? I'd like to see this.