I find this problem quite difficult to solve: 1. If I as a human request a website, then I should be shown the content. Everyone agrees. 2. If I as the human request the software on my computer to modify the content before displaying it, for example by installing an ad-blocker into my user agent, then that's my choice and the website should not be notified about it. Most users agree, some websites try to nag you into…
Perplexity is using stealth, undeclared crawlers to evade no-crawl directives
551–560 of 799 posts
Re: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives
#552Except when their agents happily click the "I"m not a robot" checkbox.
Re: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives
#553Earlier quoted context omitted.
> but I think it's a reasonable stance to not want other companies to profit from my hard work Imagine someone at another company reads your site, and it informs a strategic decision they make at the company to make money around the niche activity you're talking about. And they make lots of money they wouldn't have otherwise. That's totally legal and totally ethical as well. The reality is, if you do hard work and ma…
On a more human level, I think it's bleak that someone who makes a blog just to share stuff for fun is going to have most of his traffic be scrapers that distill, distort, and reheat whatever he's writing before serving it to potential readers.
If someone writes valuable stuff on a blog almost nobody finds, that's a tragedy.
If LLM's can process the information and provide it to people in conversations where it will be most helpful, where they never would have found it otherwise, then that's amazing!
If all you're trying to do is help people with the information you've discovered, why do you care if it's delivered via your own site or via LLM? You just want it out there helping people.
Re: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives
#554Earlier quoted context omitted.
Your comment and the above comment of course show different cases. An agent making a request on the explicit behalf of someone else is probably something most of us agree is reasonable. "What are the current stories on Hacker News?" -- the agent is just doing the same request to the same website that I would have done anyways. But the sort of non-explicit just-in-case crawling that Perplexity might do for a general q…
As a person who has a couple of sites out there, and witnesses AI crawlers coming and fetching pages from these sites, I have a question: What prevents these companies from keeping a copy of that particular page, which I specifically disallowed for bot scraping, and feed it to their next training cycle? Pinky promises? Ethics? Laws? Technical limitations? Leeroy Jenkins?
What prevents anyone else? robots.txt is a request, not an access policy.
Re: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives
#555I had to check that this did come out of CloudFlare.
Re: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives
#556Re: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives
#557Earlier quoted context omitted.
To stretch the analogy to the breaking point: If you send 10,000 personal shoppers all at once to the same store just to check prices, the store's going to be rightfully annoyed that they aren't making sales because legit buyers can't get in.
Your comment and the above comment of course show different cases. An agent making a request on the explicit behalf of someone else is probably something most of us agree is reasonable. "What are the current stories on Hacker News?" -- the agent is just doing the same request to the same website that I would have done anyways. But the sort of non-explicit just-in-case crawling that Perplexity might do for a general q…
When you swap in an AI and ask what are the current stories. The AI fetches the front page and every thread and feeds it back to you. You are less likely to participate in discussion because you've already had the info summarized.
Re: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives
#558Earlier quoted context omitted.
That's somewhat antisocial, but perfectly legal in the US. It's called PayPal Honey, for example, and has been running for 13 years now.
Since when does PayPal Honey replace ads on websites? > PayPal Honey is a browser extension that automatically finds and applies coupon codes at checkout with a single click.
Re: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives
#559>We conducted an experiment by querying Perplexity AI with questions about these domains, and discovered Perplexity was still providing detailed information regarding the exact content hosted on each of these restricted domains Thats... less conclusive than I'd like to see, especially for a content marketing article that's calling out a company in particular. Specifically it's unclear on whether Perplexity was crawli…
They do end up looking bad out of Cloudflare's report, who are the "good guys" in this story - btw Cloudflare's been very pushy lately with their we'll save the web, content independence day marketspeak. But deep in the back of my head, Cloudflare's goodwill elevates Perplexity cunning habilities (assuming they're the culprit since no real evidence, only heresay is in the OP), both companies look like titans fighting, which ends up being positive for Perplexity, at least in the inflated perception of their firepower... if that makes any sense.
Re: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives
#560We learned to dislike "bubbles" in the past decades but bubbles make sense and are natural, obviously if you're not alone in it.
When it becomes awfully busy with machines and machine content humans will learn to reconnect.