Live data from Hacker News

GPTBot – OpenAI’s Web Crawler

platform.openai.com

51–60 of 327 posts

Re: GPTBot – OpenAI’s Web Crawler

#52

I'd like to know how the bot handles copyrighted information, as well as things like music and images licensed in different ways. Also, where is this data going? The existing ChatGPT says it has nothing past 2021.

> Also, where is this data going? The existing ChatGPT says it has nothing past 2021.

To the next ChatGPT

Re: GPTBot – OpenAI’s Web Crawler

#53
What’s the incentive for people to allow the crawler at all?

Unlike search engines, chatgpt doesn’t cite references at all (last I tried) or even if it does it often makes up nonexistent references. And because it rephrases the content, there’s often no way to prove they got the material from a particular source, so harder to litigate plagiarism too.

How would contributing to the weights of this LLM help content creators?

Re: GPTBot – OpenAI’s Web Crawler

#54
post #50

Earlier quoted context omitted.

If your website appears on the search results of Google and they show ads next to it, aren't you entitled to that revenue too?

Google allows me to limitlessly search their index that allows me to find other pages too and in turn, they sell my attention so it is somewhat fair proposition in contrast to a wall gardened AI model being charged by per token such as GPT 4 that includes my content as well.

Can you not use ChatGPT as well?

I think you'll find if you do try to push Google Search too far, its not quite "limitless" either.

Re: GPTBot – OpenAI’s Web Crawler

#55
post #16

Earlier quoted context omitted.

Counterpoint (not just to be annoying — I think you pose a very interesting unanswered question): If I read your hobby website about photography and use it to take 1% better pictures, do I owe you 1% of what my clients pay me? I think that probably most people would say no, assuming you could even determine that 1% in a way that both parties agreed was fair. I think generally, we have an understanding that some stuff…

Interesting point though I'd go with another analogy. You can go to a library to borrow a book, but you can't go to the library and copy all the books for your own use.

If the library owned an effectively infinite copies of each book why wouldn’t they let you borrow one copy of each book?

Re: GPTBot – OpenAI’s Web Crawler

#56
post #16

Earlier quoted context omitted.

Counterpoint (not just to be annoying — I think you pose a very interesting unanswered question): If I read your hobby website about photography and use it to take 1% better pictures, do I owe you 1% of what my clients pay me? I think that probably most people would say no, assuming you could even determine that 1% in a way that both parties agreed was fair. I think generally, we have an understanding that some stuff…

Interesting point though I'd go with another analogy. You can go to a library to borrow a book, but you can't go to the library and copy all the books for your own use.

I used to go to the library, find books with the relevant chapters related to what I wanted to learn, and the librarian would photo copy all the pages I wanted to take home. So I guess technically you could copy all the books for your own use.

It's just impractical to photocopy every page of every book in a library.

Re: GPTBot – OpenAI’s Web Crawler

#57
post #16
post #5

If you scrape my hobby website about photography, scuba diving or let's say baking or gardening which improves your model by let's say a delta of 0.00000000001 than shouldn't I get some free credits to use that model or proportionate share in the revenue stream? EDIT: scuba diving NOT scooba diving

Counterpoint (not just to be annoying — I think you pose a very interesting unanswered question): If I read your hobby website about photography and use it to take 1% better pictures, do I owe you 1% of what my clients pay me? I think that probably most people would say no, assuming you could even determine that 1% in a way that both parties agreed was fair. I think generally, we have an understanding that some stuff…

This is a very thought provoking point and it throughly stimulated me to think deeper and through. Purpose of my website is threefold, document my own knowledge, maybe some vanity and the urge to give back something to "someone" make a better living or similar.

Things get interesting at corporate scale. There are fat VC funds, executives, board of directors and what not - making far more and far more comfortable than an individual trying to get better at their craft to put food on the table. And on top of that, you don't give me access to the product that was refined on my input.

It is like someone learning photography from my website but later taking a really masterpiece shot but asking me for money each time I want to view the photo in their studio.

There are no easy answers, I concur.

Thanks for your comment though, really. :)

Re: GPTBot – OpenAI’s Web Crawler

#58
post #25

Earlier quoted context omitted.

If a human read a website and profited from the knowledge obtained, I don’t think we’d expect them to pay royalties to the site owner.

A human reads a website, watches/clicks on ad, buys merch, subscribes to website, sets a bookmark to a website, shares an article, invites others etc. What will be the point of sharing knowledge or content if it will no longer be associated to an individual or organization?

That sounds like an argument against any bot visiting a monetized website.

Some people publish content freely on the Internet as a form of note taking, publicity, public discourse or for the betterment of like minded individuals, akin to why we’re here commenting on HN.

I guess I assume public content defaults into this “for the benefit of the world” category, where it’s up to the publisher to gate content as desired.

Re: GPTBot – OpenAI’s Web Crawler

#59
post #31
post #24

Earlier quoted context omitted.

If they implemented this properly, they should be retroactively filtering all their content that is no longer allowed in the robots.txt, or carries the #NoAI tag.

My understanding is that it's not easy to untrain a model of data already fed to it. Regarding noai tags - is this respected or just wishful?

Like every time you put content on the internet: you depend on their good will to respect these tags, or robots.txt. OpenAI can decide to ignore it. It's wishful thinking.

Re: GPTBot – OpenAI’s Web Crawler

#60
post #16

Earlier quoted context omitted.

Counterpoint (not just to be annoying — I think you pose a very interesting unanswered question): If I read your hobby website about photography and use it to take 1% better pictures, do I owe you 1% of what my clients pay me? I think that probably most people would say no, assuming you could even determine that 1% in a way that both parties agreed was fair. I think generally, we have an understanding that some stuff…

Interesting point though I'd go with another analogy. You can go to a library to borrow a book, but you can't go to the library and copy all the books for your own use.

Not really sure that this analogy applies, because I could definitely photocopy as many books from the library as I physically can. No one is going to stop me.
Post reply on HN