Live data from Hacker News

Perplexity Response to Cloudflare

twitter.com

41–48 of 48 posts

Re: Perplexity Response to Cloudflare

#41

Earlier quoted context omitted.

I use Perplexity regularly for research because it does a good job accessing, preprocessing and citing relevant resources. Which do you think is better: the service respects my desire for it to do a good job and ignore site owners blocking agent access because "don't like automated agents", or the service respects said site owners' - what I consider unreasonable - desires and not do a good job for me? Expand to the i…

I can totally see your point. It's a bit like that fight of news agencies against the free snippets and aggregations on 3rd party websites. The Internet is supposed to be open after all. But it also feels like essentially "pirating" the webpages while erasing their brand. Maybe it's even a tolerable transitive situation, but you can't even argue it's beneficial in the same way as game piracy could be according to som…

> In the long term, we need an incentive for the content creators to willingly allow such processing. Otherwise, a lot of high quality content will eventually become members-only with DRM-like anti agent protections.

I partially agree with this. Yes, some incentive is OK, for some cases. I wouldn't be OK with a mandatory header/message for example showing up in my output, unless there's some very direct relevance to the content. But there could be some kind of tipper markup/code embedded in the site metadata that my agent abstracts away as content rating feedback options, and tips automatically made on my behalf if I have it configured and selected the "useful" option. Of course source citation should also be a mandatory part of the output, for that branding and also in case there's desire to go beyond the output.

However, there will also always be content authors out there who share quality content freely with no expectation of any kind of return. The "problem" is that such content usually isn't SEO-optimized, and so likely won't be in the top results. There will be little lost if those optimizing for return start blocking their content as they'll also be automatically deranked, by virtue of content access issues, and the non-optimized content will then rise to the surface.

TL;DR: suggested configurable creator-tipping system abstracted behind feedback options, and the likely case that those who block access will be deranked in favor of those maintaining open access.

Re: Perplexity Response to Cloudflare

#42
post #31

Perplexity has really convinced me about this. There is a clear difference between automated bots scraping data at bulk for later use, and automated bots working on behalf of users on direct requests. I can see a reasonable argument that some of the first type of automation could be tolerable for websites with strict limits, the second type I think by default should not be tolerated at all. Perplexity's value proposi…

> Perplexity's value proposition appears to be "we're going to take the stuff off your website, and present it to our users. We're not going to show them your ads, we're not going to offer them your premium services or referrals to other products, we're going to strip out the value from your content and take it for our users". This, exactly this is a primary reason why I use Perplexity. I want the valued content, wit…

Yes and the result will be one of two options: option a (more likely) the underlying sites will literally just disappear, their business model no longer works and the content that you want (but apparently not enough to respect the authors) will cease to exist. It will most likely be replaced with AI slop replicas of the content you wanted. Or option b (much less likely) the content you want will move behind premium services where AI companies will have to negotiate subscriptions you will have entered the cable TV bundle era of the internet.

Re: Perplexity Response to Cloudflare

#43
post #42

Earlier quoted context omitted.

> Perplexity's value proposition appears to be "we're going to take the stuff off your website, and present it to our users. We're not going to show them your ads, we're not going to offer them your premium services or referrals to other products, we're going to strip out the value from your content and take it for our users". This, exactly this is a primary reason why I use Perplexity. I want the valued content, wit…

Yes and the result will be one of two options: option a (more likely) the underlying sites will literally just disappear, their business model no longer works and the content that you want (but apparently not enough to respect the authors) will cease to exist. It will most likely be replaced with AI slop replicas of the content you wanted. Or option b (much less likely) the content you want will move behind premium s…

Oh there's also part c where, once a or b happens, it'll clear the way for quality non-SEO content by creators who just want to share with 0 expectation of any return, which will finally see light once the return-optimized stuff has died or been walled away. The internet will be back to how it was before and meant to be with content shared for fun and/or interest, not profit, taking the front scene.

Re: Perplexity Response to Cloudflare

#44

Cloudflare did explain a proper solution: "Separate bots for separate activities". E.g. here: one bot for scraping/indexing, and one for non-persistent user-driven retrieval. Website owners have a right to block both if they wish. Isn't it obvious that bypassing a bot block is a violation of the owners right to decide whom to admit? Perplexity's almost seems to believe that "robots.txt was only made for scraping bots…

> bypassing a bot block is a violation of the owners right to decide whom to admit? There is only a violation if the bot finds a way around a login block. Same for human. But whatever is on the public web is... public. For all.

So it's ok to block someone "because you didn't include a session token I gave you in exchange for knowing the password" but it's not ok to block someone "because you didn't stick to manually-operated user agents as I told you via robots.txt"? What about not letting someone play level 42 "because you didn't complete level 41"?

A web server providing a response to your request is akin to a restaurant server doing the same. Except for specific situations related to civil rights, they are free to not deal with you for any reason.

Re: Perplexity Response to Cloudflare

#45

Earlier quoted context omitted.

I use Perplexity regularly for research because it does a good job accessing, preprocessing and citing relevant resources. Which do you think is better: the service respects my desire for it to do a good job and ignore site owners blocking agent access because "don't like automated agents", or the service respects said site owners' - what I consider unreasonable - desires and not do a good job for me? Expand to the i…

I can totally see your point. It's a bit like that fight of news agencies against the free snippets and aggregations on 3rd party websites. The Internet is supposed to be open after all. But it also feels like essentially "pirating" the webpages while erasing their brand. Maybe it's even a tolerable transitive situation, but you can't even argue it's beneficial in the same way as game piracy could be according to som…

This kind of makes sense for chatgpt and others. But perplexity links to your content directly. I end up clicking more perplexity sources than search results in practice. I don't know how well that generalises, but the traffic is not just going away.

Re: Perplexity Response to Cloudflare

#46

Earlier quoted context omitted.

> bypassing a bot block is a violation of the owners right to decide whom to admit? There is only a violation if the bot finds a way around a login block. Same for human. But whatever is on the public web is... public. For all.

So it's ok to block someone "because you didn't include a session token I gave you in exchange for knowing the password" but it's not ok to block someone "because you didn't stick to manually-operated user agents as I told you via robots.txt"? What about not letting someone play level 42 "because you didn't complete level 41"? A web server providing a response to your request is akin to a restaurant server doing the…

Typically when something is behind a login, it denotes a private space intended for a particular set of persons given explicit access. It's senseless to block people from using agents if the same people would otherwise have access, unless there is an abuse of that access, ie. action which is to the detriment of the space. And though some of that does happen, it obviously isn't the full story. I have a Perplexica instance running locally that I sometimes use (but often don't as Perplexity does a much better job). Should that also be blocked?

Hmm maybe a civil case could be potentially made here too, re disability. By blocking LLM use, sites are reducing the ability of select users to reasonably interact with the content. Just could become a thing in a few years if this nonsense continues.

Re: Perplexity Response to Cloudflare

#47
post #8
post #7

Why did Perplexity use ChatGPT to write this? That's a competitor. Or are they just so bad at writing that their own style looks like it?

Perplexity uses third-party models. They aren't a frontier AI lab like OpenAI, Anthropic, or DeepSeek.

The main thing I use ChatGPT for is web searches.

Re: Perplexity Response to Cloudflare

#48

While agents act on behalf of the user, they won't see nor click any ads; they won't sign-up to any newsletter; they won't buy the website owner a coffee. They don't act as humans just because humans triggered them. They simply take what they need and walk away.

And Cloudflare is the police entity to enforce this, right?
Post reply on HN