Earlier quoted context omitted.
> "Stealth" crawlers are always going to win the game. no, because we'll end up with remote attestation needed to access any site of value
Yes, because there's always the option for a camera pointed at the screen and a robot arm moving the mouse. AI is hoping to solve much harder problems.
Perplexity is using stealth, undeclared crawlers to evade no-crawl directives
491–500 of 799 posts
Re: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives
#492AI companies continuing to have problems with the concept of "consent" is increasingly alarming god help us if they ever manage to build anything more than shitty chatbots
Do you ask for consent before you visit a website? If I told you, you personally, to stop visiting my blog, would you stop?
LLM programs does not have human rights.
Re: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives
#493Earlier quoted context omitted.
The way to prevent people from downloading your pages and using them is to take them off the public internet. There are laws to prevent people from violating your copyright or from preventing access to your service (by excessive traffic). But there is (thankfully) no magical right that stops people from reading your content and describing it.
Many site operators want people to access their content, but prevent AI companies from scraping their sites for training data. People who think like that made tools like Anubis, and it works. I also want to keep this distinction on the sites I own. I also use licenses to signal that this site is not good to use for AI training, because it's CC BY-NC-SA-2.0. So, I license my content appropriately (No derivative, Non-c…
Your license is probably not relevant. I can go to the cinema and watch a movie, then come on this website and describe the whole plot. That isn't copyright infringement. Even if I told it to the whole world, it wouldn't be copyright infringement. Probably the movie seller would prefer it if I didn't tell anyone. Why should I care?
I actually agree that AI companies are generally bad and should be stopped - because they use an exorbitant amount of bandwidth and harm the services for other users. At least they should be heavily taxed. I don't even begrudge people for using Anubis, at least in some cases. But it is wrong-headed (and actually wrong in fact) to try to say someone may or may not use my content for some purpose because it hurts my feelings or it messes with my ad revenue. We have laws against copyright infringement, and to prevent service disruption. We should not have laws that say, yes you can read my site but no you can't use it to train an LLM, or to build a search index. That would be unethical. Call for a windfall tax if they piss you off so much.
Re: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives
#494Earlier quoted context omitted.
How about I open a proxy, replace all ads with my ads, redirect the content to you and we share the ad revenue?
That's somewhat antisocial, but perfectly legal in the US. It's called PayPal Honey, for example, and has been running for 13 years now.
> PayPal Honey is a browser extension that automatically finds and applies coupon codes at checkout with a single click.
Re: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives
#495Earlier quoted context omitted.
But I can send my personal shopper and you'll be none the wiser.
Sure. There's lots of things you could do, but you don't do them because they are wrong. Might does not make right.
It's like saying a web browser that is customized in any way is wrong. If one configures their browser to eagerly load links so that their next click is instant, is that now wrong?
Re: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives
#496This is why Perplexity is my preferred deep search engine. The no-crawl directives don't really make sense when I'm doing research and want my tool of choice to be able to pull from any relevant source. If a site doesn't want particular users to access their content, put it behind a login. The only way I - and eventually many others - will see it in the first place anyway is when it pops up as a cited source in the L…
> The no-crawl directives don't really make sense when I'm doing research and want my tool of choice to be able to pull from any relevant source. If you are the source I think they could make plenty of sense. As an example, I run a website where I've spent a lot of time documenting the history of a somewhat niche activity. Much of this information isn't available online anywhere else. As it happens I'm happy to let b…
How do you square these two? Of course big companies profit from your work, this is why they send all these bots to crawl your site.
Re: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives
#497Question for those in this thread who are okay with this: If I have endpoints that are computationally expensive server-side, what mechanism do you propose I could use to avoid being overwhelmed? The web will be a much worse place if such services are all forced behind captchas or logins.
Re: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives
#498Earlier quoted context omitted.
No, they're supposed to rally together and fight for better laws and enforcement of those laws. Which is, arguably, exactly what they've done just in a way that you and I don't like.
What kind of laws and enforcement would stop a foreign actor from effectively DDoSing your site? What if the actor has (illegally) hacked tech-illiterate users so they have domestic residential IP addresses?
The kind of laws and enforcement that would block that entire country from the internet if it doesn't get its criminal act together.
Re: Perplexity is using stealth, undeclared crawlers to evade no-crawl directives
#499if I am willing to pay a penny a page, i and the people like me won't have to put up with clickwrap nonsense
free access doesn't have to be shut off (ok, it will be, but it doesn't have to be, and doesn't that tell you something?)
reddit could charge stiffer fees, but refund quality content to encourage better content. i've fantacized about ideas like "you pay upfront a deposit; you get banned, you lose your deposit; withdraw, have your deposit back", the goal being simplify the moderation task while encouraging quality.
because where the internet is headed is just more and more trash.
here's another idea, pay a penny per search at google/search engine of choice. if you don't like the results, you can take the penny back. google's ai can figure out how to please you. if the pennies don't keep coming in, they serve you ad-infested results; serve up ad-infested results, you can send your penny to a different search engine.