Earlier quoted context omitted.
Some stores do not welcome Instacart or Postmates shoppers. You can shop there. You can shop with your phone out, scanning every item to price match, something that some bookstores frown on, for example. Third party services cannot send employees to index their inventory, nor can they be dispatched to pick up an item you order online. Their reasons vary. Some don’t want their businesses perception of quality to be ta…
But I can send my personal shopper and you'll be none the wiser.
Perplexity is using stealth, undeclared crawlers to evade no-crawl directives
131–140 of 799 posts
Re: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives
#132Earlier quoted context omitted.
Some stores do not welcome Instacart or Postmates shoppers. You can shop there. You can shop with your phone out, scanning every item to price match, something that some bookstores frown on, for example. Third party services cannot send employees to index their inventory, nor can they be dispatched to pick up an item you order online. Their reasons vary. Some don’t want their businesses perception of quality to be ta…
But I can send my personal shopper and you'll be none the wiser.
Re: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives
#133I find this problem quite difficult to solve: 1. If I as a human request a website, then I should be shown the content. Everyone agrees. 2. If I as the human request the software on my computer to modify the content before displaying it, for example by installing an ad-blocker into my user agent, then that's my choice and the website should not be notified about it. Most users agree, some websites try to nag you into…
And isn't the obvious solution to just make some sort of browsers add-on for the LLM summary so the request comes from your browser and then gets sent to the LLM? I think the main concern here is the huge amount of traffic from crawling just for content for pre-training.
Re: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives
#134> How can you protect yourself? Put your valuable content behind a paywall.
Re: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives
#135Earlier quoted context omitted.
I think it's an issue of scale. The next step in your progression here might be: If / when people have personal research bots that go and look for answers across a number of sites, requesting many pages much faster than humans do - what's the tipping point? Is personal web crawling ok? What if it gets a bit smarter and tried to anticipate what you'll ask and does a bunch of crawling to gather information regularly to…
Maybe we should just institutionalize and explicitly legalize the Internet Archive and Archive Team. Then, I can download a complete and halfway current crawl of domain X from the IA and that way, no additional costs are incurred for domain X. But of course, most website publishers would hate that. Because they don't want people to access their content, they want people to look at the ads that pay them. That's why to…
>Common Crawl maintains a free, open repository of web crawl data that can be used by anyone.
Re: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives
#136Earlier quoted context omitted.
If I put time and effort into a food recipe should I (get) compensation? the answer is apparently "no", and I don't really how recipe books have suffered as a result of less gatekeeping. "How will the internet work"? Probably better in some ways. There is plenty of valuable content on the internet given for free, it's being buried in low-value AI slop.
You understand that HN is ad supported too, right?
But what is your point? Is the value in HN primarily in its hosting, or the non-ad-supported community?
Re: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives
#137I do not really get why user-agent blocking measures are despised for browsers but celebrated for agents? It’s a different UI, sure, but there should be no discrimination towards it as there should be no discrimination towards, say, Links terminal browser, or some exotic Firefox derivative.
>I do not really get why user-agent blocking measures are despised for browsers but celebrated for agents? AI broke the brains of many people. The internet isn't a monolith, but prior to the AI boom you'd be hard pressed to find people who were pro-copyright (except maybe a few who wanted to use it to force companies to comply with copyleft obligations), pro user-agent restrictions, or anti-scraping. Now such positio…
People can believe that corporations are using the power asymmetry between them and individuals through copywrite law to stifle the individual to protect profits. People can also believe that corporations are using the power asymmetry between them and individuals through AI to steal intellectual labor done by individuals to protect their profits. People’s position just might be that the law should be used to protect the rights of parties when there is a large power asymmetry.
Re: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives
#138I am sorry, Cloudafre is the internet police now?
Re: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives
#139Earlier quoted context omitted.
Some stores do not welcome Instacart or Postmates shoppers. You can shop there. You can shop with your phone out, scanning every item to price match, something that some bookstores frown on, for example. Third party services cannot send employees to index their inventory, nor can they be dispatched to pick up an item you order online. Their reasons vary. Some don’t want their businesses perception of quality to be ta…
But I can send my personal shopper and you'll be none the wiser.
Re: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives
#140I don't really know anything about DRM except it is used to take down sites that violate it. Perhaps it is possible for cloudflare (or anyone else) to file a take down notice with Perplexity. That might at least confuse them.
Corporations use this to protect their content. I should be able to protect mine as well. What's good for the goose.