the service is actually very convenient no matter faang likes it or not.
Perplexity is using stealth, undeclared crawlers to evade no-crawl directives
91–100 of 799 posts
Re: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives
#92Earlier quoted context omitted.
>I do not really get why user-agent blocking measures are despised for browsers but celebrated for agents? AI broke the brains of many people. The internet isn't a monolith, but prior to the AI boom you'd be hard pressed to find people who were pro-copyright (except maybe a few who wanted to use it to force companies to comply with copyleft obligations), pro user-agent restrictions, or anti-scraping. Now such positio…
I think the intelligent conclusion would be that the people you are looking at have more nuanced beliefs than you initially thought. Talking about broken brains is often just mediocre projecting
You don't seem to reject my claim that for many, principles took a backseat to "does this help or hurt evil corporations". If that's what passes as "nuance" to you, then sure.
>Talking about broken brains is often just mediocre projecting
To be clear, that part is metaphorical/hyperbolic and not meant to be taken literally. Obviously I'm not diagnosing people who switched sides with a psychiatric condition.
Re: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives
#93I find this problem quite difficult to solve: 1. If I as a human request a website, then I should be shown the content. Everyone agrees. 2. If I as the human request the software on my computer to modify the content before displaying it, for example by installing an ad-blocker into my user agent, then that's my choice and the website should not be notified about it. Most users agree, some websites try to nag you into…
The problem is with the LLM then training on that data _once_ and then storing it forever and regurgitating it N times in the future without ever crediting the original author.
So far, humans themselves did this, but only for relatively simple information (ratio of rice and water in specific $recipe). You're not gonna send a link to your friend just to see the ratio, you probably remember it off the top of your head.
Unfortunately, the top of an LLMs head is pretty big, and they are fitting almost the entire website's content in there for most websites.
The threshold beyond which it becomes irreproducible for human consumers, and therefore, copyrightable (lot of copyright law has "reasonable" term which refers to this same concept) has now shifted up many many times higher.
Now, IMO:
So far, for stuff that won't fit in someone's head, people were using citations (academia, for example). LLMs should also use citations. That solves the ethical problem pretty much. That the ad ecosystem chose views as the monetisation point and is thus hurt by this is not anyone else's problem. The ad ecosystem can innovate and adjust to the new reality in their own time and with their own effort. I promise most people won't be waiting. Maybe google can charge per LLM citation. Cost Per Citation, you even maintain the acronym :)
Re: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives
#94Re: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives
#95god help us if they ever manage to build anything more than shitty chatbots
Re: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives
#96Earlier quoted context omitted.
It doesn’t matter what you do with it afterwards. Crawling is defined by recursively following links. If a user asks software about a specific page and it fetches it, then a human is operating that software, it’s not a crawler. You can’t just redefine “crawler” to mean “software that does things I don’t like”. It very specifically refers to software that recursively follows links.
Technically correct (the best kind of correct), but if I set a thousand users on to a website to each download a single page and then feed the information they retrieve from that one page into my AI model, then are those thousand users not performing the same function as a crawler, even though they are (technically) not one? If it looks like a duck, quacks like a duck and surfs a website like a duck, then perhaps we…
No.
robots.txt is designed to stop recursive fetching. It is not designed to stop AI companies from getting your content. Devising scenarios in which AI companies get your content without recursively fetching it is irrelevant to robots.txt because robots.txt is about recursively fetching.
If you try to use robots.txt to stop AI companies from accessing your content, then you will be disappointed because robots.txt is not designed to do that. It’s using the wrong tool for the job.
Re: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives
#97I find this problem quite difficult to solve: 1. If I as a human request a website, then I should be shown the content. Everyone agrees. 2. If I as the human request the software on my computer to modify the content before displaying it, for example by installing an ad-blocker into my user agent, then that's my choice and the website should not be notified about it. Most users agree, some websites try to nag you into…
Their reasons vary. Some don’t want their businesses perception of quality to be taken out of their control (delivering cold food, marking up items, poor substitutions). Some would prefer their staff service and build relationships with customers directly, instead of disinterested and frequently quite demanding runners. Some just straight up disagree with the practice of third party delivery.
I think that it’s pretty unambiguously reasonable to choose to not allow an unrelated business to operate inside of your physical storefront. I also think that maps onto digital services.
Re: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives
#98I find this problem quite difficult to solve: 1. If I as a human request a website, then I should be shown the content. Everyone agrees. 2. If I as the human request the software on my computer to modify the content before displaying it, for example by installing an ad-blocker into my user agent, then that's my choice and the website should not be notified about it. Most users agree, some websites try to nag you into…
All of these scenarios assume you have an unconditional right to access the content on a website in whatever way you want. Do you think you do? Or is there a balance between the owner's rights, who bears the content production and hosting/serving costs, and the rights of the end user who wishes to benefit from that content? If you say that you have the right, and that right should be legally protected, to do whatever…
The reason people are up in arms is because rights they previously enjoyed are being stripped away by the current platforms. The content owner's rights aren't as valid as my own in the current world; they trump mine 10 to 1. If I "buy" a song and the content owner decides that my country is politically unfriendly, they just delete it and don't refund me. If I request to view their content and they start by wasting my bandwidth sending me an ad I haven't consented to, how can I even "not engage"? The damage is done, and there's no recourse.
Re: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives
#99Re: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives
#100I find this problem quite difficult to solve: 1. If I as a human request a website, then I should be shown the content. Everyone agrees. 2. If I as the human request the software on my computer to modify the content before displaying it, for example by installing an ad-blocker into my user agent, then that's my choice and the website should not be notified about it. Most users agree, some websites try to nag you into…
Correct, it’s user hostile to dictate which software is allowed to see content.