Live data from Hacker News

Perplexity is using stealth, undeclared crawlers to evade no-crawl directives

blog.cloudflare.com

531–540 of 799 posts

Re: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives

#531
post #491

Earlier quoted context omitted.

Yes, because there's always the option for a camera pointed at the screen and a robot arm moving the mouse. AI is hoping to solve much harder problems.

Won't work with biometric attestation. For example, banks in China require periodic facial recognition to continue the banking session.

What's stopping these companies from offloading the scraping onto their users?

"Either pay us $50/month or install our extension, and when prompted, solve any captchas or authenticate with your ID (as applicable) on the given website so we can train on the content.

Re: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives

#532
post #525

Earlier quoted context omitted.

> but I think it's a reasonable stance to not want other companies to profit from my hard work Imagine someone at another company reads your site, and it informs a strategic decision they make at the company to make money around the niche activity you're talking about. And they make lots of money they wouldn't have otherwise. That's totally legal and totally ethical as well. The reality is, if you do hard work and ma…

For me, the point is that the person who has put in the work then has some rights to decide how that information is accessed and re-used. I think it is a reaosnable position for someone to hold that they want individuals to be able to freely use some content they produced, but not for a company to use and profit from that same content. I think just saying "It's public now" lacks any nuance. Ultimately these AI tools…

> the person who has put in the work then has some rights to decide how that information is accessed and re-used

You do, but you give up those rights when you make the work public.

You think an author has any control over who their book gets lent to once somebody buys a copy? You think they get a share of profits when a CEO reads their book and they make a better decision? Of course not.

What you're asking for is unreasonable. It's not workable. Knowledge can't be owned. Once you put it out there, it's out there. We have copyright and patent protections in specific circumstances, but that's all. You don't own facts, no matter how much hard work and research they took to figure out.

Re: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives

#533

I find this problem quite difficult to solve: 1. If I as a human request a website, then I should be shown the content. Everyone agrees. 2. If I as the human request the software on my computer to modify the content before displaying it, for example by installing an ad-blocker into my user agent, then that's my choice and the website should not be notified about it. Most users agree, some websites try to nag you into…

The solution to 3 seems fairly straightforward: user requests content and passes it to llm to summarise.

Re: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives

#534
post #105

Seems a win. CF being internet police is a problem too but someone credible publicly shaming a company for shady scraping is good. Even if it just creates conversation Somehow this needs to go back to search era where all players at least attempt to behave. This scrapping Ddos stuff and I don’t care if it kills your site (while “borrowing” content) is unethical bullshit

Shaming doea not work in the era of "no shame".

Re: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives

#535
post #335

Earlier quoted context omitted.

> The no-crawl directives don't really make sense when I'm doing research and want my tool of choice to be able to pull from any relevant source. If you are the source I think they could make plenty of sense. As an example, I run a website where I've spent a lot of time documenting the history of a somewhat niche activity. Much of this information isn't available online anywhere else. As it happens I'm happy to let b…

> but I think it's a reasonable stance to not want other companies to profit from my hard work Imagine someone at another company reads your site, and it informs a strategic decision they make at the company to make money around the niche activity you're talking about. And they make lots of money they wouldn't have otherwise. That's totally legal and totally ethical as well. The reality is, if you do hard work and ma…

On a more human level, I think it's bleak that someone who makes a blog just to share stuff for fun is going to have most of his traffic be scrapers that distill, distort, and reheat whatever he's writing before serving it to potential readers.

Re: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives

#536
post #156

Earlier quoted context omitted.

This has already started with people using special tags also people making content just for llms.

There is a standard for making content just for LLMs: https://llmstxt.org

From their example I don’t see any value in this on top of making and actually human friendly site.

> Converting complex HTML pages with navigation, ads, and JavaScript into LLM-friendly plain text is both difficult and imprecise.

None of these conditions should apply for websites with purpose of providing information.

Re: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives

#537
post #491

Earlier quoted context omitted.

Yes, because there's always the option for a camera pointed at the screen and a robot arm moving the mouse. AI is hoping to solve much harder problems.

Won't work with biometric attestation. For example, banks in China require periodic facial recognition to continue the banking session.

yea but those are not open sites, try imposing that on an open site you'd want to actually attract human traffic to

Re: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives

#538
post #342

Earlier quoted context omitted.

do you think that every well-meaning GET request should be treated the same way as a distributed attack ? The latter is the reason why people use CF not the former.

How does one tell a "well-meaning" request from an attack?

By the volume, distribution, and parameters (get and post body) of the requests.

Re: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives

#539

Earlier quoted context omitted.

Maybe we should just institutionalize and explicitly legalize the Internet Archive and Archive Team. Then, I can download a complete and halfway current crawl of domain X from the IA and that way, no additional costs are incurred for domain X. But of course, most website publishers would hate that. Because they don't want people to access their content, they want people to look at the ads that pay them. That's why to…

I have mixed feelings on this. Many websites (especially the bigger ones) are just businesses. They pay people to produce content, hopefully make enough ad revenue to make a profit, and repeat. Anything that reproduces their content and steals their views has a direct effect on their income and their ability to stay in business. Maybe IA should have a way for websites to register to collect payment for lost views or…

If magazines and newspapers were once able to be funded by native ads, so can websites. The spying industry doesn't want you to know this, but ads work without spying too - just look at all the IRL billboards still around.

Re: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives

#540

the internet needs micropayments (probably millipayments). if crawlers want to pay me a penny a page, crawl me 24-7 plz if I am willing to pay a penny a page, i and the people like me won't have to put up with clickwrap nonsense free access doesn't have to be shut off (ok, it will be, but it doesn't have to be, and doesn't that tell you something?) reddit could charge stiffer fees, but refund quality content to encou…

[dead]
Post reply on HN