Live data from Hacker News

GPTBot – OpenAI’s Web Crawler

platform.openai.com

21–30 of 327 posts

Re: GPTBot – OpenAI’s Web Crawler

#21
post #5

If you scrape my hobby website about photography, scuba diving or let's say baking or gardening which improves your model by let's say a delta of 0.00000000001 than shouldn't I get some free credits to use that model or proportionate share in the revenue stream? EDIT: scuba diving NOT scooba diving

If I learn something from your StackOverflow answers, do you expect me to share a percentage of my future salary with you?

Re: GPTBot – OpenAI’s Web Crawler

#22
post #5

If you scrape my hobby website about photography, scuba diving or let's say baking or gardening which improves your model by let's say a delta of 0.00000000001 than shouldn't I get some free credits to use that model or proportionate share in the revenue stream? EDIT: scuba diving NOT scooba diving

You don't have to publish anything on the internet. And when you do, you may limit the allowed audience to just the group of your friends etc. Why publish anything if you worry that someone may consume it?

Re: GPTBot – OpenAI’s Web Crawler

#23
I wonder how much the regression of ChatGPT is due to it adding new content which has its origin from ChatGPT. The blog and SEO spam with ChatGPT fluff is going through the roof, eventually all of that will get crawled too and the model will just get positively reinforced on its own output. Or is that not a concern?

Re: GPTBot – OpenAI’s Web Crawler

#24
post #12

> Web pages crawled with the GPTBot user agent may potentially be used to improve future models > To disallow GPTBot to access your site you can add the GPTBot to your site’s robots.txt Too late - they already grabbed content from my personal website.

If they implemented this properly, they should be retroactively filtering all their content that is no longer allowed in the robots.txt, or carries the #NoAI tag.

Re: GPTBot – OpenAI’s Web Crawler

#25
post #19

Just a random thought: While they are already earning money by scraping data, would it not be nice if they pay the site owners a certain amount of the money they earn?

If a human read a website and profited from the knowledge obtained, I don’t think we’d expect them to pay royalties to the site owner.

Re: GPTBot – OpenAI’s Web Crawler

#26
post #4

Nice of them to respect crawling after they've already trained their model. Presumably these headers don't affect any pages they've already crawled to train GPT(?)

Hoping this is what they’ll use to train future models and deprecate the older ones before the legal cases proceed any further.

Re: GPTBot – OpenAI’s Web Crawler

#27
post #21
post #5

If you scrape my hobby website about photography, scuba diving or let's say baking or gardening which improves your model by let's say a delta of 0.00000000001 than shouldn't I get some free credits to use that model or proportionate share in the revenue stream? EDIT: scuba diving NOT scooba diving

If I learn something from your StackOverflow answers, do you expect me to share a percentage of my future salary with you?

Can you share your knowledge with millions at once?

If so, then pay.

Re: GPTBot – OpenAI’s Web Crawler

#29
post #4

Nice of them to respect crawling after they've already trained their model. Presumably these headers don't affect any pages they've already crawled to train GPT(?)

GPT-4 finished training in August 2022, before the release of ChatGPT.

If they had announced this sooner hardly anyone on the internet would have noticed. Props to them for adding it now.

Re: GPTBot – OpenAI’s Web Crawler

#30
post #4

Nice of them to respect crawling after they've already trained their model. Presumably these headers don't affect any pages they've already crawled to train GPT(?)

Hoping this is what they’ll use to train future models and deprecate the older ones before the legal cases proceed any further.

The legal cases don't mean anything. The rule of law has all but disappeared from the corporate world. The idea that courts or regulators will be able to control AI is laughable. They are too corrupt, and they are way too slow.
Post reply on HN