Live data from Hacker News

Perplexity is using stealth, undeclared crawlers to evade no-crawl directives

blog.cloudflare.com

151–160 of 799 posts

Re: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives

#151
post #144

Those Challenges can be bypassed too using various browser automation. With the Comet-like tool, Perplexity can advance its crawling activity with much more human-like behaviour.

If they can trick the ad networks then go for it. If the ad networks can detect it and exclude those visits we should be able to.

Re: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives

#153
post #139

Earlier quoted context omitted.

But I can send my personal shopper and you'll be none the wiser.

It's all about scale. The impact of your personal shopper is insignificant unless you manage to scale it up into a business where everyone has a personal shopper by default.

Well then. Seems like you would be a fool to not allow personal shoppers then.

The point is the web is changing, and people use a different type of browser now. Ans that browser happens to be LLMs.

Anybody complaining about the new browser has just not got it yet, or has and is trying to keep things the old way because they don’t know how or won’t change with the times. We have seen it before, Kodak, blockbuster, whatever.

Grow up cloud flare, some is your business models don’t make sense any more.

Re: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives

#154

Earlier quoted context omitted.

But I can send my personal shopper and you'll be none the wiser.

It’s possible to violate all sorts of social norms. Societies that celebrate people that do so are on the far opposite end of the spectrum from high trust ones. They are rather unpleasant.

Just the Silicon Valley ethos extended to it's logical conclusions. These companies take advantage of public space, utilities and goodwill at industrial scale to "move fast and break things" and then everyone else has to deal with the ensuing consequences. Like how cities are awash in those fucking electric scooters now.

Mind you I'm not saying electric scooters are a bad idea, I have one and I quite enjoy it. I'm saying we didn't need five fucking startups all competing to provide them at the lowest cost possible just for 2/3s of them to end up in fucking landfills when the VC funding ran out.

Re: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives

#155

Earlier quoted context omitted.

Some stores do not welcome Instacart or Postmates shoppers. You can shop there. You can shop with your phone out, scanning every item to price match, something that some bookstores frown on, for example. Third party services cannot send employees to index their inventory, nor can they be dispatched to pick up an item you order online. Their reasons vary. Some don’t want their businesses perception of quality to be ta…

But I can send my personal shopper and you'll be none the wiser.

True, and I would ask, what is your point? Is it that no rule can have 100% perfect enforcement? That all rules have a grey area if you look close enough? Was it just a "gotcha" statement meant to insinuate what the prior commenter said was invalid?

Re: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives

#156
post #123

[flagged]

> companies who want AI to recommend their products need to turn this off before it starts hurting them financially Content marketing, gamified SEO, and obtrusive ads significantly hurt the quality of Google search. For all its flaws, LLMs don’t feel this gamified yet. It’s disappointing that this is probably where we’re headed. But I hope OpenAI and Anthropic realize that this drop in search result quality might be…

This has already started with people using special tags also people making content just for llms.

Re: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives

#157
post #82

Earlier quoted context omitted.

Relevant to this is that Perplexity lies to the user when specifically asked about this. When the user asks if there is a robots.txt file for the domain, it lies and says there is not. If an LLM will not (cannot?) tell the truth about basic things, why do people assume it is a good summarizer of more complex facts?

The article did not test if the issue was specific to robots.txt or if it can not find other files. There is a difference between doing a poor summarization of data, and failing to even be able to get the data to summarize in the first place.

> specific to robots.txt > poor summarization of data

I'm not really addressing the issue raised in the article. I am noting that the LLM, when asked, is either lying to the user or making a statement that it does not know to be true (that there is no robots.txt). This is way beyond poor summarization.

Re: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives

#158

I find this problem quite difficult to solve: 1. If I as a human request a website, then I should be shown the content. Everyone agrees. 2. If I as the human request the software on my computer to modify the content before displaying it, for example by installing an ad-blocker into my user agent, then that's my choice and the website should not be notified about it. Most users agree, some websites try to nag you into…

1. To access a website you need a limited anonymized token that proves you are a human being, issued by a state authority

2. the end

I am firmly convinced that this should be the future in the next decade, since the internet as we know it has been weaponized and ruined by social media, bots, state actors and now AI.

There should exist an internet for humans only, with a single account per domain.

Re: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives

#159
post #141
post #111

Earlier quoted context omitted.

Unexpected underdog argument. What is happening in reality is all companies are racing to (a) scrape, buy and collect as much as they can from others, both individuals and companies while (b) locking down their own data against everyone else who isn’t directly making them money (eg through viewing their ads). Part of me thinks that the open web has a paradox of tolerance issue, leading to a race to the bottom/tragedy…

We need to wake up and understand that all the information already uploaded is more or less a free web material, once taken through the lens of ML-somethings. With all the second, and third-order effects such as the fact that this changes completely the whole motivation, and consequence of open-source perhaps. It is also only a matter of time scrapers once again get through walls by twitter, reddit and alike. This is…

Reddit sold their data already. Twitter made thier own AI.

Re: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives

#160

Earlier quoted context omitted.

Some stores do not welcome Instacart or Postmates shoppers. You can shop there. You can shop with your phone out, scanning every item to price match, something that some bookstores frown on, for example. Third party services cannot send employees to index their inventory, nor can they be dispatched to pick up an item you order online. Their reasons vary. Some don’t want their businesses perception of quality to be ta…

But I can send my personal shopper and you'll be none the wiser.

To stretch the analogy to the breaking point: If you send 10,000 personal shoppers all at once to the same store just to check prices, the store's going to be rightfully annoyed that they aren't making sales because legit buyers can't get in.
Post reply on HN