Earlier quoted context omitted.
Here's how perplexity works: 1) It takes your query, and given the complexity might expand it to several search queries using an LLM. ("rephrasing") 2) It runs queries against a web search index (I think it was using Bing or Brave at first, but they probably have their own by now), and uses an LLM to decide which are the best/most relevant documents. It starts writing a summary while it dives into sources (see next).…
What’s wrong with it downloading documents when the user asks it to? My browser also downloads whole documents and sometimes even prefetches documents I haven’t even clicked on yet. Toss in a adblocker or reader mode and my browser also strips all the ads. Why is it okay for me to ask my browser to do this but I can’t ask my LLM to do the same?
Perplexity is using stealth, undeclared crawlers to evade no-crawl directives
341–350 of 799 posts
Re: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives
#342> it is built on trust. This is funny coming from Cloudflare, the company that blocks most of the internet from being fetched with antispam checks even for a single web request. The internet we knew was open and not trusted , but thanks to companies like Cloudflare, now even the most benign , well meaning attempt to GET a website is met with a brick wall. The bots of Big Tech, namely Google, Meta and Apple are of cou…
> The internet we knew was open and not trusted , but thanks to companies like Cloudflare, now even the most benign , well meaning attempt to GET a website is met with a brick wall I don't think it's fair to blame Cloudflare for that. That's looking at a pool of blood and not what caused it: the bots/traffic which predate LLMs. And Cloudflare is working to fix it with the PrivacyPass standard (which Apple joined). Ea…
Re: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives
#343Earlier quoted context omitted.
> The bots of Big Tech, namely Google, Meta and Apple are of course exempt from this by pretty much every website and by cloudflare. But try being anyone other than them , no luck. Cloudflare is the biggest enabler of this monopolistic behavior Plenty of site/service owners explicitly want Google, Meta and Apple bots (because they believe they have a symbiotic relationship with it) and don't want your bot because the…
they didnt seem to mind when openai et al. took all their content to train LLMs when they were still parasites that didn't have a symbiotic relationship. This thinking is kind of too pro-monopolist for me
Re: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives
#344> it is built on trust. This is funny coming from Cloudflare, the company that blocks most of the internet from being fetched with antispam checks even for a single web request. The internet we knew was open and not trusted , but thanks to companies like Cloudflare, now even the most benign , well meaning attempt to GET a website is met with a brick wall. The bots of Big Tech, namely Google, Meta and Apple are of cou…
> The internet we knew was open and not trusted ... monopolistic behavior Monopolistic is the wrong word, because you have the problem backwards. Cloudflare isnt helping Apple/Google... It's helping its paying consumers and those are the only services those consumers want to let through. Do you know how I can predict that AI agents, the sort that end users use to accomplish real tasks, will never take off? Because th…
Re: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives
#345> it is built on trust. This is funny coming from Cloudflare, the company that blocks most of the internet from being fetched with antispam checks even for a single web request. The internet we knew was open and not trusted , but thanks to companies like Cloudflare, now even the most benign , well meaning attempt to GET a website is met with a brick wall. The bots of Big Tech, namely Google, Meta and Apple are of cou…
We're moving progressively in the direction of "pages can't be served for free anymore". Which, I don't think is a problem, and in fact I think it's something we should have addressed a long time ago. Cloudflare only needs to exist because the server doesn't get paid when a user or bot requests resources. Advertising only needs to exist because the publisher doesn't get paid when a user or bot requests resources. And…
Your idea of micro transacting web requests would play into it and probably end up with a system like Netflix where your ISP has access to a set of content creators to whom they grant ‘unlimited’ access as part of the service fee.
I’d imagine that accessing any content creators which are not part of their package will either be blocked via a paywall (buy an addon to access X creators outside our network each month) or charged at an insane price per MB as is the case with mobile data.
Obvious this is all super hypothetical but weirder stuff has happened in my lifetime
Re: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives
#346Earlier quoted context omitted.
Depending on the implementation (a big if) it would help smaller websites, because it would make hosting much cheaper. ISPs don’t choose what sites users visit, only what they pay. As long as the ISP isn’t giving significant discounts to visiting big sites (just charging a fixed rate per bytes downloads and uploaded) and charging something reasonable, visiting a small site would be so cheap (a few cents at most, but…
But users depend on major sites like google [insert service] still and will prioritize their usage accordingly like limited minutes and texts back in the day, right?
The average American allegedly* downloads 650-700GB/month, or >20GB/day. 10MB is more than enough for a webpage (honestly, 1MB is usually enough), so that means on average, ISPs serve over 2000 webpages worth of data per day. And the average internet plan is allegedly** $73/month, or That’s cheap enough, wrapped in a monthly bill, users won’t even pay attention to what sites they visit. The only people hurt by an ideal (granted, ideal) implementation are those who abuse fixed rates and download unreasonable amounts of data, like web crawlers who visit the same page seconds apart for many pages in parallel.
* https://www.astound.com/learn/internet/average-internet-data...
** https://www.nerdwallet.com/article/finance/how-much-is-inter...
Re: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives
#347Earlier quoted context omitted.
> The internet we knew was open and not trusted , but thanks to companies like Cloudflare, now even the most benign , well meaning attempt to GET a website is met with a brick wall I don't think it's fair to blame Cloudflare for that. That's looking at a pool of blood and not what caused it: the bots/traffic which predate LLMs. And Cloudflare is working to fix it with the PrivacyPass standard (which Apple joined). Ea…
do you think that every well-meaning GET request should be treated the same way as a distributed attack ? The latter is the reason why people use CF not the former.
Re: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives
#348robots.txt is not a blocking mechanism; it's a hint to indicate which parts of a site might be of interest to indexing.
People started using robots.txt to lie and declare things like no part of their site is interesting, and so of course that gets ignored.
Re: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives
#349> it is built on trust. This is funny coming from Cloudflare, the company that blocks most of the internet from being fetched with antispam checks even for a single web request. The internet we knew was open and not trusted , but thanks to companies like Cloudflare, now even the most benign , well meaning attempt to GET a website is met with a brick wall. The bots of Big Tech, namely Google, Meta and Apple are of cou…
I'm sorry, but that's some crazy take. Sure, the internet should be open and not trusted. But physical reality exists. Hosting and bandwidth cost money. I trust Google won't DDoS my site or cost my an arbitrary amount of money. I won't trust bots made by random people on the internet in the same way. The fact that Google respects robots.txt while Perplexity doesn't tells you why people trust Google more than random b…
Google already has access to any webpage because its own search Crawlers are allowed by most websites, and google crawls recursively. Thus Gemini has an advantage of this synergy with google search. Perplexity does not crawl recursively (i presume -- therefore it does not need to consult robots.txt), and it doesn't have synergies with a major search engine.
Re: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives
#350> it is built on trust. This is funny coming from Cloudflare, the company that blocks most of the internet from being fetched with antispam checks even for a single web request. The internet we knew was open and not trusted , but thanks to companies like Cloudflare, now even the most benign , well meaning attempt to GET a website is met with a brick wall. The bots of Big Tech, namely Google, Meta and Apple are of cou…
> The internet we knew was open and not trusted ... monopolistic behavior Monopolistic is the wrong word, because you have the problem backwards. Cloudflare isnt helping Apple/Google... It's helping its paying consumers and those are the only services those consumers want to let through. Do you know how I can predict that AI agents, the sort that end users use to accomplish real tasks, will never take off? Because th…