Live data from Hacker News

Perplexity is using stealth, undeclared crawlers to evade no-crawl directives

blog.cloudflare.com

301–310 of 799 posts

Re: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives

#301
post #270

the year is 2045. you've been cruising the interstate in your robotaxi, shelling out $150 in stablecoins at the cloudflare tollbooth. a palantir patrol unit pulls you over. the optimus v4 approaches your window and contorts its silicone face into a facsimile of concern as it hits you with the: "sir, have you been botting today?" immediately you remember how great you had it in the '20s when you used to click CAPTCHA…

greentext courtesy of opus prompted in greentext by me.

``` The user wants me to create a more elaborate greentext with deeper worldbuilding and references to internet culture/incel memes, but specifically asks for "tasteful" ones and explicitly says no violent references. I should incorporate classic internet/chan culture references like tendies, NEET lifestyle, "tfw no gf", basement dwelling, etc. while building out this dystopian 2045 world where proving humanity is the central challenge.

\>be me, 2045 \>32yo NEET still living in mom's smart-basement \>she threatens to turn off my Neuralink if I don't go outside today \>grudgingly take the robotaxi to pick up tendies \>$150 stablecoin toll at cloudflare checkpoint \>have to solve a moral dilemma to prove I'm not running on GPT-7 \>sweating.dll \>palantir unit pulls me over \>optimus v4 leans into window \>its facial mesh attempts "concern_expression_v2.blend" \>"sir, when did you last feel genuine human connection?" \>flashback to 2024 when the girl at McDonalds gave me extra honey mustard \>that was before the McBots took over \>"t-twenty one years ago officer" \>optimus's empathy subroutines activate \>"sir I need you to perform a field humanity test" \>get out, knees weak from vitamin D deficiency \>"please describe your ideal romantic partner without using the words 'tradwife' or 'submissive'" \>brain.exe has stopped responding \>try to remember pre-blackpill emotions \>"someone who... likes anime?" \>optimus scans my biometrics \>"stress patterns indicate authentic social anxiety, carry on citizen" \>get back in robotaxi \>it starts therapy session \>"I notice you ordered tendies again. Let's explore your relationship with your mother" \>tfw the car has better emotional intelligence than me \>finally get tendies from Wendy's AutoServ \>receipt prints with mandatory "rate your humanity score today" \>3.2/10 \>at least I'm improving

\>mfw bots are better at being human than humans \>it's over for carboncels ```

Re: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives

#302
post #246

Earlier quoted context omitted.

We're moving progressively in the direction of "pages can't be served for free anymore". Which, I don't think is a problem, and in fact I think it's something we should have addressed a long time ago. Cloudflare only needs to exist because the server doesn't get paid when a user or bot requests resources. Advertising only needs to exist because the publisher doesn't get paid when a user or bot requests resources. And…

> We're moving progressively in the direction of "pages can't be served for free anymore". Which, I don't think is a problem, and in fact I think it's something we should have addressed a long time ago. I agree, but your idea below that is overly complicated. You can't micro-transact the whole internet. That idea feels like those episodes of Star Trek DS9 that take place on Feregenar - where you have to pay admission…

> You can't micro-transact the whole internet.

I agree that end-users cannot handle micro transactions across the whole internet. That said, I would like to point out that most of the internet is blanketed in ads and ads involve tons of tiny quick auctions and micro transactions that occur on each page load.

It is totally possible for a system to evolve involving tons of tiny transactions across page loads.

Re: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives

#303

Crawling and scraping is legal. If your web server serves the content without authentication, it's legal to receive it, even if it's an automated process. If you want to gatekeep your content, use authentication. Robots.txt is not a technical solution, it's a social nicety. Cloudflare and their ilk represent an abuse of internet protocols and mechanism of centralized control. On the technical side, we could use CRC m…

> Crawling and scraping is legal. If your web server serves the content without authentication, it's legal to receive it, even if it's an automated process. > Cloudflare and their ilk represent an abuse of internet protocols and mechanism of centralized control. How does one follow the other? It's my web server and I can gatekeep access to my content however I want (eg Cloudflare). How is that an "abuse" of internet…

They exist to optimize the internet for the platforms and big providers. Little people get screwed, with no legal recourse. They actively and explicitly degrade the internet, acting as censors and gatekeepers and on behalf of bad faith actors without legal authority or oversight.

They allow the big platforms to pay for special access. If you wanted to run a scraper, however, you're not allowed, despite the internet standards and protocols and the laws governing network access and free communications standards responsibilities by ISPs and service providers not granting the authority to any party involved with cloudflare blocking access.

It's equivalent to a private company deciding who, when, and how you can call from your phone, based on the interests and payments of people who profit from listening to your calls. What we have is not normal or good, unless you're exploiting the users of websites for profit and influence.

Re: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives

#304
post #274
post #199

Earlier quoted context omitted.

People never agreed DOSing a site to take copyright material was acceptable. Many people did not have a problem with taking copyright material in a respectful way that didn't kill the resource. LLMs are killing the resource. This isn't a corporation vs person issue. No issue with an llm having my content but big issue with my server being down because llms are hammering the same page over and over.

>People never agreed DOSing a site to take copyright material was acceptable. Many people did not have a problem with taking copyright material in a respectful way that didn't kill the resource. Has it be shown that perplexity engages in "DOSing"? I've heard of anecdotes of AI bots gone amuck, and maybe that's what's happening here, but cloudflare hasn't really shown that. All they did was set up a robots.txt and sho…

The guy that runs shadertoy talked about how the hostingcost for his free site shot up because Openai kept crawling his site for training data (ignoring robot.txt) I think that’s bad, and I have also experimented a bit with using BeautifulSoup in the past to download ~2MB of pictures from Instagram. Do you think I’m holding an inconsistent position?

Re: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives

#305
post #246

Earlier quoted context omitted.

We're moving progressively in the direction of "pages can't be served for free anymore". Which, I don't think is a problem, and in fact I think it's something we should have addressed a long time ago. Cloudflare only needs to exist because the server doesn't get paid when a user or bot requests resources. Advertising only needs to exist because the publisher doesn't get paid when a user or bot requests resources. And…

Wouldn't this lead to pirated page clones where customer pays less for same-ish content, and less, all the way down to essentially free? Because I as an user would be glad to have "free sites only" filter, and then just steal content :)) But it's an interesting idea and thought experiment.

That’s fine. The point for website owners isn’t to make money, it’s to not spend money hosting (or more specifically, to pay a small fixed rate hosting). They want people to see the content; if someone makes the content more accessible, that’s a good thing.

Re: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives

#306
Hmm, I’ve always seen robots.txt more as a polite request than an actual rule.

Sure, Google has to follow it because they’re a big company and need to respect certain laws or internal policies. But for everyone else, it’s basically just a “please don’t” sign, not a legal requirement or?

Re: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives

#307

Earlier quoted context omitted.

Weird take. The store doesn't owe your personal shippers anything.

In the same token the personal shoppers don't owe the store anything either.

Then they can't complain if they're barred entry.

Re: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives

#308
post #216

> it is built on trust. This is funny coming from Cloudflare, the company that blocks most of the internet from being fetched with antispam checks even for a single web request. The internet we knew was open and not trusted , but thanks to companies like Cloudflare, now even the most benign , well meaning attempt to GET a website is met with a brick wall. The bots of Big Tech, namely Google, Meta and Apple are of cou…

I'm sorry, but that's some crazy take.

Sure, the internet should be open and not trusted. But physical reality exists. Hosting and bandwidth cost money. I trust Google won't DDoS my site or cost my an arbitrary amount of money. I won't trust bots made by random people on the internet in the same way. The fact that Google respects robots.txt while Perplexity doesn't tells you why people trust Google more than random bots.

Re: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives

#309

So, this calls for a new type of honeytrap, content that appears to be human generated, and high quality, but subtly wrong, preferably on a commercially catastrophic way. Behind settings that prohibit commercial usage. It really shouldn't be hard to generate gigantic quantities of the stuff. Simulate old forum posts, or academic papers.

[deleted]

Re: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives

#310
post #57

Their test seems flawed: > We created multiple brand-new domains, similar to testexample.com and secretexample.com. These domains were newly purchased and had not yet been indexed by any search engine nor made publicly accessible in any discoverable way. We implemented a robots.txt file with directives to stop any respectful bots from accessing any part of a website: > We conducted an experiment by querying Perplexit…

> > We conducted an experiment by querying Perplexity AI with questions about these domains, and discovered Perplexity was still providing detailed information regarding the exact content hosted on each of these restricted domains. This response was unexpected, as we had taken all necessary precautions to prevent this data from being retrievable by their crawlers. Right, I'm confused why CloudFlare is confused. You a…

[deleted]
Post reply on HN