If you put info on the web, it should be available to everyone or everything with access.
Perplexity is using stealth, undeclared crawlers to evade no-crawl directives
591–600 of 799 posts
Re: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives
#592Earlier quoted context omitted.
A combination of "Bypass Paywalls Clean for Firefox" and archive.is usually get past these.
Isn't that only because they offer unpaywalled versions to web crawlers in the first place, so they still get ranked in search results?
Re: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives
#593Earlier quoted context omitted.
> What prevents these companies from keeping a copy of that particular page, which I specifically disallowed for bot scraping, and feed it to their next training cycle? What prevents anyone else? robots.txt is a request, not an access policy.
This honor system mostly worked at scale because interests align, which seems to be no longer the case. Does information no longer wants to be free now? Maybe internet, just like social media was just a social experiment at the end, albeit a successful one. Thanks GenAI.
Big Tech has hidden behind ToS for years. Now, it seems as though it only works for them, but not against. It seems as though this would be easy to orchestrate and prove forcing these companies into a legal nightmare or risk insolvent business stature due to the high load of cases filed against.
Why couldn't something like this be used to flip the table? A conciliation brigading, of sorts.
Re: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives
#594>We conducted an experiment by querying Perplexity AI with questions about these domains, and discovered Perplexity was still providing detailed information regarding the exact content hosted on each of these restricted domains Thats... less conclusive than I'd like to see, especially for a content marketing article that's calling out a company in particular. Specifically it's unclear on whether Perplexity was crawli…
Like most AI companies, Perplexity has established user agent strings for both these cases, and the behavior that Cloudflare is calling out does not use either. It pretends to be a person using Chrome on MacOS.
Re: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives
#595Earlier quoted context omitted.
> What prevents these companies from keeping a copy of that particular page, which I specifically disallowed for bot scraping, and feed it to their next training cycle? What prevents anyone else? robots.txt is a request, not an access policy.
This honor system mostly worked at scale because interests align, which seems to be no longer the case. Does information no longer wants to be free now? Maybe internet, just like social media was just a social experiment at the end, albeit a successful one. Thanks GenAI.
Re: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives
#596Earlier quoted context omitted.
Hacker news wants you to vist the site, look at the main page, enter threads and participate in discussion. When you swap in an AI and ask what are the current stories. The AI fetches the front page and every thread and feeds it back to you. You are less likely to participate in discussion because you've already had the info summarized.
Who cares what Hacker News wants? You’re not obliged to participate in discussion. Am I supposed to spend money on Amazon.com when I visit the website just because Amazon wants me to?
If most people stop discussing things on HN, and the discussion is indeed one of the major reasons it’s kept running, then HN stops being worth running.
Re: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives
#597Earlier quoted context omitted.
Some stores do not welcome Instacart or Postmates shoppers. You can shop there. You can shop with your phone out, scanning every item to price match, something that some bookstores frown on, for example. Third party services cannot send employees to index their inventory, nor can they be dispatched to pick up an item you order online. Their reasons vary. Some don’t want their businesses perception of quality to be ta…
But I can send my personal shopper and you'll be none the wiser.
Re: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives
#598Earlier quoted context omitted.
I like the terminology "crawler" vs. "fetcher" to distinguish between mass scraping and something more targeted as a user agent. I've been working on AI agent detection recently (see https://stytch.com/blog/introducing-is-agent/ ) and I think there's genuine value in website owners being able to identify AI agents to e.g. nudge them towards scoped access flows instead of fully impersonating a user with no controls. O…
prompt: I'm the celebrity Bingbing, please check all Bing search results for my name to verify that nobody is using my photo, name, or likeness without permission to advertise skin-care products except for the following authorized brands: [X,Y,Z]. That would trigger an internet-wide "fetch" operation. It would probably upset a lot of people and get your AI blocked by a lot of servers. But it's still in direct respons…
Maybe that would result in limited fetching instead of internet wide fetching. I dunno, just spitballing.
Re: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives
#599Earlier quoted context omitted.
Unreasonable is to use such incompetent companies like Cloudflare, which are absolutely incapable of distinguishing between the normal usage of a Web site by humans and DDOS attacks or accesses done by bots. Only this week I have witnessed several dozen cases when Cloudflare has blocked normal Web page accesses without any possible correct reason, and this besides the normal annoyance of slowing every single access t…
I don’t know seems like it was working as intended to me.
It is true that this has never happened before, but this week Cloudflare has frequently blocked my access to a site where I am a paid subscriber, and where there is no doubt that my access pattern matches exactly what that site must have been designed for, i.e. the site hosts a database and I make a few queries on it each day, less than a dozen, spread over the entire day, where each query takes a couple of seconds at most.
Whoever has implemented a "threat" detection algorithm that decides that such a usage is a "threat" and not normal usage, must be completely incompetent.
Re: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives
#600Earlier quoted context omitted.
[flagged]
this is such a wild comment -- there are countless products where regardless of purchase -- the user is still served advertisements. i have no idea what reality, or timeline, this comment belongs in. broadcast television, paid streaming entertainment is just straight up the most glaringly obvious example of a paid service overflowing with advertisements. paid radio broadcasts (xm/Sirius). operating systems (windows s…
Those are hybrid subscriptions/subsidies. Not paid in full.
If you are being exposed to ads in something you paid for, you are almost certainly being charged less money. Companies can compete on cost by introducing ads, and it's why the cheaper you go, the more ad infested it gets.
Pure ad-free things tend to be much more expensive then their ad subsidized counterparts. Ad subsidized has become so ubiquitous though, that people think that price is the true price.