Live data from Hacker News

Perplexity Response to Cloudflare

twitter.com

31–40 of 48 posts

Re: Perplexity Response to Cloudflare

#31
Perplexity has really convinced me about this. There is a clear difference between automated bots scraping data at bulk for later use, and automated bots working on behalf of users on direct requests. I can see a reasonable argument that some of the first type of automation could be tolerable for websites with strict limits, the second type I think by default should not be tolerated at all.

Perplexity's value proposition appears to be "we're going to take the stuff off your website, and present it to our users. We're not going to show them your ads, we're not going to offer them your premium services or referrals to other products, we're going to strip out the value from your content and take it for our users".

You can argue all you want about whether that's 5k impressions a day or 1m impressions a day. It should be 0 impressions a day. It is literally just free-riding.

Also, they're meant to be a professional company taking VC money to build a business, why are writing whiny posts like a teenager? The impression I get with a lot of these companies is that their business is losing money hand over fist, they have no idea how they're going to make it work and they look absolutely panicked as a result. They come across like a company I would want to be nowhere near.

Re: Perplexity Response to Cloudflare

#32
post #31

Perplexity has really convinced me about this. There is a clear difference between automated bots scraping data at bulk for later use, and automated bots working on behalf of users on direct requests. I can see a reasonable argument that some of the first type of automation could be tolerable for websites with strict limits, the second type I think by default should not be tolerated at all. Perplexity's value proposi…

google was also accused in its early days of free riding on other people's content - google news still remains controversial. Also, in its early days, google did not have to deal with a large counter-party like cloudflare which is now a gatekeeper of sorts.

The problem I see for chatgpt/perplexity and the like is this: for good responses to many questions, they have to index the web real-time. ie, they become a search-engine. However, they cannot share revenue with the content-providers since they dont have an ad-model. I wonder how this would be resolved - perhaps thru content licensing with the large publishers.

Re: Perplexity Response to Cloudflare

#33

Earlier quoted context omitted.

They're already doing that https://docs.perplexity.ai/guides/bots There's PerplexityBot and Perplexity‑User .

And then once they see that the website operator blocked the perplexity-user, apparently instead of respecting that, they not only ignore robots.txt, but actively try to bypass the security measures established with the explicit purpose of limiting their access. If this was about bypassing DRM rather than AI-WAF, it would be plainly illegal. To me this invalidates their whole claim that Cloudflare fails to tell the d…

I use Perplexity regularly for research because it does a good job accessing, preprocessing and citing relevant resources. Which do you think is better: the service respects my desire for it to do a good job and ignore site owners blocking agent access because "don't like automated agents", or the service respects said site owners' - what I consider unreasonable - desires and not do a good job for me? Expand to the inevitably increasing LLM-for-research user base.

Re: Perplexity Response to Cloudflare

#34

Cloudflare did explain a proper solution: "Separate bots for separate activities". E.g. here: one bot for scraping/indexing, and one for non-persistent user-driven retrieval. Website owners have a right to block both if they wish. Isn't it obvious that bypassing a bot block is a violation of the owners right to decide whom to admit? Perplexity's almost seems to believe that "robots.txt was only made for scraping bots…

> bypassing a bot block is a violation of the owners right to decide whom to admit?

There is only a violation if the bot finds a way around a login block. Same for human. But whatever is on the public web is... public. For all.

Re: Perplexity Response to Cloudflare

#35

While agents act on behalf of the user, they won't see nor click any ads; they won't sign-up to any newsletter; they won't buy the website owner a coffee. They don't act as humans just because humans triggered them. They simply take what they need and walk away.

Funnily enough this would also mean some humans are not human, like me. I exhibit that behavior exactly as well. Maybe I'm an agent acting on behalf of myself, whatever that means.

Re: Perplexity Response to Cloudflare

#36
post #22
post #21

Earlier quoted context omitted.

Do you have an ad-blocker? If so, should website owners be able to disable your ad-blocker via a setting they send to you? It's their content, after all.

They may, and do, refuse to send you the page. That is a more realistic parallel. Without advertising the web would be largely unsupportable financially without per site subscriptions.

Without advertising, the web would be a far healthier place today, free of click-bait and ad-infested, low effort, low quality content. Bad business models wouldn't've survived the light of day.

Re: Perplexity Response to Cloudflare

#37
post #31

Perplexity has really convinced me about this. There is a clear difference between automated bots scraping data at bulk for later use, and automated bots working on behalf of users on direct requests. I can see a reasonable argument that some of the first type of automation could be tolerable for websites with strict limits, the second type I think by default should not be tolerated at all. Perplexity's value proposi…

> Perplexity's value proposition appears to be "we're going to take the stuff off your website, and present it to our users. We're not going to show them your ads, we're not going to offer them your premium services or referrals to other products, we're going to strip out the value from your content and take it for our users".

This, exactly this is a primary reason why I use Perplexity. I want the valued content, without the unnecessary distractions that I'll never consciously touch anyway (there have been accidental clicks now and then, because some site designers really want people to click that ad and go all out to embed it into the content, and it only leads to great annoyance and sometimes a promise never to visit that site again).

Re: Perplexity Response to Cloudflare

#39

Earlier quoted context omitted.

And then once they see that the website operator blocked the perplexity-user, apparently instead of respecting that, they not only ignore robots.txt, but actively try to bypass the security measures established with the explicit purpose of limiting their access. If this was about bypassing DRM rather than AI-WAF, it would be plainly illegal. To me this invalidates their whole claim that Cloudflare fails to tell the d…

I use Perplexity regularly for research because it does a good job accessing, preprocessing and citing relevant resources. Which do you think is better: the service respects my desire for it to do a good job and ignore site owners blocking agent access because "don't like automated agents", or the service respects said site owners' - what I consider unreasonable - desires and not do a good job for me? Expand to the i…

I can totally see your point. It's a bit like that fight of news agencies against the free snippets and aggregations on 3rd party websites. The Internet is supposed to be open after all.

But it also feels like essentially "pirating" the webpages while erasing their brand. Maybe it's even a tolerable transitive situation, but you can't even argue it's beneficial in the same way as game piracy could be according to some. In the long term, we need an incentive for the content creators to willingly allow such processing. Otherwise, a lot of high quality content will eventually become members-only with DRM-like anti agent protections.

The incentive doesn't have to be monetary. I could for example imagine some website owners allow AI agents that commit to upfront verbatim repeating some sort of mandatory headers/messages/acknowledgements from the content authors, before copying or summarizing, and are known to stick to this commitment.

You can also bypass the problem already now by accessing and copying the content manually, and then putting it in the context of a tool like NotebookLM. Nobody's hurt, because you have actually seen the source by yourself, and that's all the website owners can reasonably demand.

TL;DR: why even post quality content in open if the audience won't see your ads, your donation button, or even your name. What do you think?

Re: Perplexity Response to Cloudflare

#40

> When companies like Cloudflare mischaracterize user-driven AI assistants as malicious bots, they're arguing that any automated tool serving users should be suspect—a position that would criminalize email clients and web browsers, or any other service a would-be gatekeeper decided they don’t like. I wonder if Perplexity or others mix the traffic of the two types so they’re indistinguishable, specifically to make thi…

@viraptor above mentions that they actually do try first with an explicit perplexity-agent: https://news.ycombinator.com/item?id=44797682 . So there's no ambiguity. The worst they could accuse Cloudflare, is that they don't give website owners an easy way to only block scrapers while allowing user-driven agents (do they?).
Post reply on HN