How is GPTBot allowed or disallowed?
medium.com
How is GPTBot allowed or disallowed?
1–6 of 6 posts
Re: How is GPTBot allowed or disallowed?
#2Re: How is GPTBot allowed or disallowed?
#3GPTBot — the official crawler of OpenAI, has been announced for nearly 2 months. GPTBot is for crawling web information to improve the models of OpenAI, e.g. GPT-4. We are wondering what the reactions from the Internet are. Is the bot being accepted or rejected?
Re: How is GPTBot allowed or disallowed?
#4GPTBot — the official crawler of OpenAI, has been announced for nearly 2 months. GPTBot is for crawling web information to improve the models of OpenAI, e.g. GPT-4. We are wondering what the reactions from the Internet are. Is the bot being accepted or rejected?
In my poor armchair quarterback opinion if people wish for something to not be crawled then they must make a best effort to ensure only humans are accessing it with strong authentication, legal agreements, best-effort bot detection and also have binding legal contracts that implement punitive actions for doing something with data it was not approved for and then actually follow through with legal action for breach of contract.
Re: How is GPTBot allowed or disallowed?
#5GPTBot — the official crawler of OpenAI, has been announced for nearly 2 months. GPTBot is for crawling web information to improve the models of OpenAI, e.g. GPT-4. We are wondering what the reactions from the Internet are. Is the bot being accepted or rejected?
I suspect not many website operators/developers are aware this exists. Usage of robots.txt is unenforceable and would only show intent to OpenAI. This would not be useful for other LLM's as Google, Bing and other search engines already have decades of ingested data to feed their LLM's. In my poor armchair quarterback opinion if people wish for something to not be crawled then they must make a best effort to ensure on…
Companies like OpenAI also have to do a lot of things to ensure the compliance to the regulation.
Re: How is GPTBot allowed or disallowed?
#6Earlier quoted context omitted.
I suspect not many website operators/developers are aware this exists. Usage of robots.txt is unenforceable and would only show intent to OpenAI. This would not be useful for other LLM's as Google, Bing and other search engines already have decades of ingested data to feed their LLM's. In my poor armchair quarterback opinion if people wish for something to not be crawled then they must make a best effort to ensure on…
The number of disallow we found in the robots.txt files actually surprises us. Companies like OpenAI also have to do a lot of things to ensure the compliance to the regulation.
Did legislation pass requiring people and their bots to obey robots.txt? If so I totally missed it. That would be big news if so.