Live data from Hacker News

GPTBot – OpenAI’s Web Crawler

platform.openai.com

11–20 of 327 posts

Re: GPTBot – OpenAI’s Web Crawler

#11
post #5

If you scrape my hobby website about photography, scuba diving or let's say baking or gardening which improves your model by let's say a delta of 0.00000000001 than shouldn't I get some free credits to use that model or proportionate share in the revenue stream? EDIT: scuba diving NOT scooba diving

Depends on what the courts say. We'll have to see.

Re: GPTBot – OpenAI’s Web Crawler

#12
> Web pages crawled with the GPTBot user agent may potentially be used to improve future models

> To disallow GPTBot to access your site you can add the GPTBot to your site’s robots.txt

Too late - they already grabbed content from my personal website.

Re: GPTBot – OpenAI’s Web Crawler

#14
post #5

If you scrape my hobby website about photography, scuba diving or let's say baking or gardening which improves your model by let's say a delta of 0.00000000001 than shouldn't I get some free credits to use that model or proportionate share in the revenue stream? EDIT: scuba diving NOT scooba diving

Why? I wouldn't pay you for marginally improving my baking skills either.

It is an interesting question. I would have no qualms paying for a textbook or university course for curated learning (worth noting OpenAI has paid datasets too), but paying for (or being paid for) relatively diffuse and low quality content through hobby blogs seems at odds with my expectations as an individual, and as a society we were never (en masse) concerned about things like Google's search excerpt answers...

Re: GPTBot – OpenAI’s Web Crawler

#15
post #5

If you scrape my hobby website about photography, scuba diving or let's say baking or gardening which improves your model by let's say a delta of 0.00000000001 than shouldn't I get some free credits to use that model or proportionate share in the revenue stream? EDIT: scuba diving NOT scooba diving

Depends on what the courts say. We'll have to see.

OpenAI would love that kinda regulation, it would basically kill free models.

Re: GPTBot – OpenAI’s Web Crawler

#16
post #5

If you scrape my hobby website about photography, scuba diving or let's say baking or gardening which improves your model by let's say a delta of 0.00000000001 than shouldn't I get some free credits to use that model or proportionate share in the revenue stream? EDIT: scuba diving NOT scooba diving

Counterpoint (not just to be annoying — I think you pose a very interesting unanswered question):

If I read your hobby website about photography and use it to take 1% better pictures, do I owe you 1% of what my clients pay me?

I think that probably most people would say no, assuming you could even determine that 1% in a way that both parties agreed was fair. I think generally, we have an understanding that some stuff is put out into the world for other humans to learn from and use to make themselves better, and that they don’t owe the original authors anything other than the price of admission.

I guess it comes down to this: do we think that training a model is:

- like storing and later reproducing a version of some collected data, or

- like learning from collected data, and synthesizing new info?

Is there even a meaningful distinction, for a computer?

(Is there even a meaningful distinction for a human…?)

Re: GPTBot – OpenAI’s Web Crawler

#17
post #5

If you scrape my hobby website about photography, scuba diving or let's say baking or gardening which improves your model by let's say a delta of 0.00000000001 than shouldn't I get some free credits to use that model or proportionate share in the revenue stream? EDIT: scuba diving NOT scooba diving

This doesn't appear consistent with other visitors to your website. If a cafe owner uses info on your site to improve their baking, should they also be required to share their revenue with you?

Re: GPTBot – OpenAI’s Web Crawler

#18
post #5

If you scrape my hobby website about photography, scuba diving or let's say baking or gardening which improves your model by let's say a delta of 0.00000000001 than shouldn't I get some free credits to use that model or proportionate share in the revenue stream? EDIT: scuba diving NOT scooba diving

> shouldn't I get some free credits to use that model or proportionate share in the revenue stream

But you do get paid in kind - you "gave" information for the AI to train on, the aI gives you information back, contextualised to your needs. Sometimes those 1000 tokens are worth much more than $0.06

You still need to be able to pay for inference costs, it's crowded and expensive on GPUs nowadays.

Re: GPTBot – OpenAI’s Web Crawler

#20
post #5

If you scrape my hobby website about photography, scuba diving or let's say baking or gardening which improves your model by let's say a delta of 0.00000000001 than shouldn't I get some free credits to use that model or proportionate share in the revenue stream? EDIT: scuba diving NOT scooba diving

Why? I wouldn't pay you for marginally improving my baking skills either. It is an interesting question. I would have no qualms paying for a textbook or university course for curated learning (worth noting OpenAI has paid datasets too), but paying for (or being paid for) relatively diffuse and low quality content through hobby blogs seems at odds with my expectations as an individual, and as a society we were never (…

Because perfect information transfer isn’t usually possible by a human reading a book or website, whereas computer systems can usually do that.

If humans could perfectly remember information, I’m sure copyright would be very different.

Post reply on HN