Earlier quoted context omitted.
Http doesnt have emotions or thought last time I checked.
It seems that a 403 makes you sad though.
Perplexity is using stealth, undeclared crawlers to evade no-crawl directives
541–550 of 799 posts
Re: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives
#542Earlier quoted context omitted.
[flagged]
Whoa, please don't post like this. We end up banning accounts that do. https://news.ycombinator.com/newsguidelines.html
Re: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives
#543Re: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives
#544Earlier quoted context omitted.
Some stores do not welcome Instacart or Postmates shoppers. You can shop there. You can shop with your phone out, scanning every item to price match, something that some bookstores frown on, for example. Third party services cannot send employees to index their inventory, nor can they be dispatched to pick up an item you order online. Their reasons vary. Some don’t want their businesses perception of quality to be ta…
These are more like a store putting up a billboard or catalog and asking people to turn off their meta AI glasses nearby because the store doesn't want AI translating it on your behalf as a tourist.
Re: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives
#545Now, it's a gazillion of AI crawlers and python crawlers, MCP servers that offer the same feature to anyone "building (personal workflow) automation" incl. bypass of various, standard protection mechanisms.
Re: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives
#546Re: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives
#547I find this problem quite difficult to solve: 1. If I as a human request a website, then I should be shown the content. Everyone agrees. 2. If I as the human request the software on my computer to modify the content before displaying it, for example by installing an ad-blocker into my user agent, then that's my choice and the website should not be notified about it. Most users agree, some websites try to nag you into…
is it just on your behalf? or is it on Perplexity's behalf? are they not archiving the pages to train on?
it's the difference between using Google Chrome vs. Chrome beaming full page snapshots to train Gemini on.
Re: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives
#548Earlier quoted context omitted.
Repeat after me - intentional discrimination of computer programs over humans is a good and praise worthy thing. We can and should make execution of computer programs harder and harder, even disproportionately so, if that makes lives of humans better and easier. LLM programs does not have human rights.
"if that makes lives of humans better" is doing a lot of heavy lifting, and remains to be explained. Computer programs don't take actions, people do. If I use a web browser, or scrape some site to make an LLM, that's me doing it, not the program. And I have human rights. If you think training LLMs should be illegal, just say that. If you think LLM companies are putting an undue strain on computer networks and they sh…
For example - humans can learn, programs can't. The "learning" cop out for LLM-corpos shouldn't be accepted by anyone, let alone by law. Humans have a fair use carve out of the copyright laws, not because it's something axiomatic, it's because some humans with empathy have forced others to allow all humans a leeway in legally using other's IP works. Just because such law exist for humans, doesn't mean that random computer programs should be applicable to it. Scraping web for LLMs should not be considered "fair use" because a) it is clearly not (commercialized later) and b) programs aren't humans and don't have equal rights.
And the list goes on. Now, I do get that train has long left the station and we are all collectively living in the anecdote about stealing a bicycle and asking god for forgiveness. But that doesn't mean I agree with this state. I'm just shouting my displeasure towards that passing train cause I'm weird like that. It's like with climate change - we are doing nothing that matters, no one discusses what really matters and I just accepted that nothing will really change. Doesn't mean I like the situation.
PS: tl;dr - LLMs clearly should be legal, it's just simple code is all. LLM corporations who steal IP content without compensation to the authors should be illegal, but of course they won't ever be.
PPS: there is a huge, gigantic gap between a single person scraping a few thousand pages for a personal use, maybe even some small local commercial use (though that's a grey area already) and a billion dollar megacorp, intent on destroying everything of value for humans in the internet for profit.
Re: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives
#549Earlier quoted context omitted.
Your comment and the above comment of course show different cases. An agent making a request on the explicit behalf of someone else is probably something most of us agree is reasonable. "What are the current stories on Hacker News?" -- the agent is just doing the same request to the same website that I would have done anyways. But the sort of non-explicit just-in-case crawling that Perplexity might do for a general q…
As a person who has a couple of sites out there, and witnesses AI crawlers coming and fetching pages from these sites, I have a question: What prevents these companies from keeping a copy of that particular page, which I specifically disallowed for bot scraping, and feed it to their next training cycle? Pinky promises? Ethics? Laws? Technical limitations? Leeroy Jenkins?
Re: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives
#550Earlier quoted context omitted.
Unless I am misunderstanding you, you are talking about something different than the article. The article is talking about web-crawling. You are talking about local / personal LLM usage. No one has any problems with local / personal LLM usage. It's when Perplexity uses web crawlers that an issue arises.
You probably need a computer that costs $250,000 or more to run the kind of LLM that Perplexity uses, but with batching it costs pennies to have the same LLM fetch a page for you, summarize the content, and tell you what is on it. And the power usage similarly, running the LLM for a single user will cost you a huge amount of money relative to the power it takes in a cloud environment with many users. Perplexity's "we…