curious, will the robots.txt be really honored? maybe legal issue if not?
It will not be honored. The ones that honor it will lose to the ones that do.
Go ahead and block AI web crawlers
11–20 of 24 posts
Re: Go ahead and block AI web crawlers
#12curious, will the robots.txt be really honored? maybe legal issue if not?
robots.txt being an honor system has always been a bit weird and search engine treatment is odder still. Google will still index your disallowed pages but refuse to show a description since they are couldn't read the page. The premise being that Google could have discovered your URLs from other pages and they don't consider a robots.txt disallow to indicate that you don't want them in search, just that you don't want…
Re: Go ahead and block AI web crawlers
#13What faith can be placed in User-Agent strings. The contents of this header have been faked since the birth of the www in the early 90s.
How does anyone know what someone will do with the data they have crawled. There are no transparency requirements, there are no legally-enforceable agreements. There are only IP addresses and User-Agent headers.
No one can look at a User-Agent header or an IP address and conclude, "I know what someone will do with the data I let them access because this header or address confirms it". Those accessing the data could use it for any purpose or transfer it to someone else to use for any purpose. Unless the website operator has a binding agreement with the organisation doing the crawling, any assumptions about future behaviour made on basis of a header or IP address offer no control whatsoever.
Perhaps an IP address could be used to conclude the HTTP requests are being sent by Company X, but the address does not indicate what Company X will do with the data it collects. Company X can do whatever it wants. It does not need to tell anyone what it is doing; it could make up a story about what it is doing that conveniently conceals facts.
These so-called "tech" companies that are training "AI" are secretive and non-transparent. Further, they will lie when it suits their interests. They do not ask for permission, they only ask for forgiveness. They are strategic and unfortunately deceptive in what they tell the public. "Trust" at own risk.
Although it may be useless as a means of controlling how crawling data is used, it still makes sense to me to put something in robots.txt to indicate there is no consent given for crawling for the purpose of training "AI". Better would be to publish some explicit notice to the public that no consent is given to use data from the website to train "AI".
Put the restrictions in a license. Let the so-called "tech" companies assent to that license. Then, when the evidence of unauthorised data use becomes availalble, enforce the license.
Re: Go ahead and block AI web crawlers
#14Didn’t the FBI admit that it just bought U.S. citizen data instead of spying on them? Almost two years in a row?
This has been happening non-stop for over 20 years without any repercussions.
I don’t think robots.txt will make a difference. If your site is internet facing, nothing will prevent it from being crawled and scraped.
If anything, i can almost guarantee AI crawling will be tired to SEO if it’s not already.
Re: Go ahead and block AI web crawlers
#15curious, will the robots.txt be really honored? maybe legal issue if not?
It will not be honored. The ones that honor it will lose to the ones that do.
Today, people make money from web traffic, and therefore want their sites indexed by search indices. If the same happens with AI eventually, authors will probably follow. What people get from having their sites ingested by AI is still unknown today — if the model doesn’t send traffic to your site, you can’t make ad money and can’t support your site (if it’s big). Data licensing that allows AI platforms to pay authors for useful content may be a solution someday
Re: Go ahead and block AI web crawlers
#16Re: Go ahead and block AI web crawlers
#17How does one differentiate between "AI web crawlers" and "non-AI web crawlers". What faith can be placed in User-Agent strings. The contents of this header have been faked since the birth of the www in the early 90s. How does anyone know what someone will do with the data they have crawled. There are no transparency requirements, there are no legally-enforceable agreements. There are only IP addresses and User-Agent…
Re: Go ahead and block AI web crawlers
#18Earlier quoted context omitted.
It will not be honored. The ones that honor it will lose to the ones that do.
Given that Google has honored robots.txt for many years, not sure this is true. It mostly means that blocked sites might not show up on the biggest, most trustworthy platforms. If these AI platforms become a big way people consume content from the internet, authors will have an interest in being there. Today, people make money from web traffic, and therefore want their sites indexed by search indices. If the same hap…
Re: Go ahead and block AI web crawlers
#19Earlier quoted context omitted.
robots.txt being an honor system has always been a bit weird and search engine treatment is odder still. Google will still index your disallowed pages but refuse to show a description since they are couldn't read the page. The premise being that Google could have discovered your URLs from other pages and they don't consider a robots.txt disallow to indicate that you don't want them in search, just that you don't want…
OpenAI documents how their crawler honors robots.txt. https://platform.openai.com/docs/gptbot
Re: Go ahead and block AI web crawlers
#20Earlier quoted context omitted.
OpenAI documents how their crawler honors robots.txt. https://platform.openai.com/docs/gptbot
After they ingested everything