What’s the incentive for people to allow the crawler at all? Unlike search engines, chatgpt doesn’t cite references at all (last I tried) or even if it does it often makes up nonexistent references. And because it rephrases the content, there’s often no way to prove they got the material from a particular source, so harder to litigate plagiarism too. How would contributing to the weights of this LLM help content crea…
Just one example: As a software vendor, you probably want OpenAI to index your documentation, so questions about your software can be answered by ChatGPT. Not everybody who creates content is a "content creator" (when did this word get the specific meaning "people who earn money or reputation from creating content").
GPTBot – OpenAI’s Web Crawler
151–160 of 327 posts
Re: GPTBot – OpenAI’s Web Crawler
#152Friendship ended with SEO. Now LEO [1] is my best friend. [1]: LLM Engine Optimization
If (generally) more data is better. As a site owner I might be happy to give free access to most of my public data and urls. You be at the mercy of my sites unique formatting and hiccups
But for some fee, I might be happy to provide api-like data access to some of my historic data that are more rich with information and promised some sort of format-encoding (AI-JSON ?)
The economics I think are still currently be discovered.
Re: GPTBot – OpenAI’s Web Crawler
#153Man what a time we live in :) It's like history is being written (ok compiled & backpropagated) right under our feet ! I can see a bots.txt entry in the near future that discern the site's data-usage for bots vs humans User-agent-class: AI Data-Policy-Allow: /news/* /articles/* Data-Policy-Deny: */comments
Although no human is going to read robots.txt or bots.txt
It'll end up as a small section in the EULA of the website which nobody reads:
> Before you click the 'reply' button please be aware we are allowing AI to crawl our comment section for training. Thank you for your consideration.
There's a little problem though:
1) Websites don't have an incentive to inform their users about this, and websites don't have an incentive to allow AI to crawl their content unless they get something back from it (e.g. payment). From this PoV, its time for OpenAI to start paying.
2) The competition (China, Russia) doesn't care about bots.txt or robots.txt and will just crawl whatever the hell they can.
Re: GPTBot – OpenAI’s Web Crawler
#154Yet another bot that completely ignores the "429 Too Many Requests" response header and happily continues hammering your tiny little side project [1] to death. Luckily, I already block the IP address they're using as it has been used for (other?) malicious bots before. [1] In my case, it relies on third-party APIs that are heavily rate limited. Any bot ignoring rate limitation measures will effectively (D)DOS my serv…
Re: GPTBot – OpenAI’s Web Crawler
#155If you scrape my hobby website about photography, scuba diving or let's say baking or gardening which improves your model by let's say a delta of 0.00000000001 than shouldn't I get some free credits to use that model or proportionate share in the revenue stream? EDIT: scuba diving NOT scooba diving
Counterpoint (not just to be annoying — I think you pose a very interesting unanswered question): If I read your hobby website about photography and use it to take 1% better pictures, do I owe you 1% of what my clients pay me? I think that probably most people would say no, assuming you could even determine that 1% in a way that both parties agreed was fair. I think generally, we have an understanding that some stuff…
Re: GPTBot – OpenAI’s Web Crawler
#156I wonder what kind of mischief facts people are going to start sneaking into OpenAI's newer models, by selectively feeding different responses to OpenAI when their crawler is identified.
Re: GPTBot – OpenAI’s Web Crawler
#157Nice of them to respect crawling after they've already trained their model. Presumably these headers don't affect any pages they've already crawled to train GPT(?)
It’s so now they can lobby for anti scraping regulation and hamper any possible catch-up.
Re: GPTBot – OpenAI’s Web Crawler
#158What’s the incentive for people to allow the crawler at all? Unlike search engines, chatgpt doesn’t cite references at all (last I tried) or even if it does it often makes up nonexistent references. And because it rephrases the content, there’s often no way to prove they got the material from a particular source, so harder to litigate plagiarism too. How would contributing to the weights of this LLM help content crea…
Re: GPTBot – OpenAI’s Web Crawler
#159What’s the incentive for people to allow the crawler at all? Unlike search engines, chatgpt doesn’t cite references at all (last I tried) or even if it does it often makes up nonexistent references. And because it rephrases the content, there’s often no way to prove they got the material from a particular source, so harder to litigate plagiarism too. How would contributing to the weights of this LLM help content crea…
Just one example: As a software vendor, you probably want OpenAI to index your documentation, so questions about your software can be answered by ChatGPT. Not everybody who creates content is a "content creator" (when did this word get the specific meaning "people who earn money or reputation from creating content").
But the book writer who wrote a detailed, expert book on how to deal with the software ("Photoshop for Dummies"?). OpenAI might be seen as a competitor.
A government would be easier to say all their data isn't allowed to be crawled, so they can sue later or just say no later on when they figure something classified was in there, or simply when they change their mind.
I believe the default response should be 'no, we'll look into it' for anyone, and then carefully let legal take a look at it (gonna be expensive). For the software vendor, too. Although their crown jewels are likely the source code to their product(s).
Re: GPTBot – OpenAI’s Web Crawler
#160I wonder what kind of mischief facts people are going to start sneaking into OpenAI's newer models, by selectively feeding different responses to OpenAI when their crawler is identified.