Live data from Hacker News

GPTBot – OpenAI’s Web Crawler

platform.openai.com

181–190 of 327 posts

Re: GPTBot – OpenAI’s Web Crawler

#181

What’s the incentive for people to allow the crawler at all? Unlike search engines, chatgpt doesn’t cite references at all (last I tried) or even if it does it often makes up nonexistent references. And because it rephrases the content, there’s often no way to prove they got the material from a particular source, so harder to litigate plagiarism too. How would contributing to the weights of this LLM help content crea…

> What’s the incentive for people to allow the crawler at all?

So that LLMs can learn from it? Profit is not the only thing that motivates people. I’ve spent years contributing to Stack Overflow to help people solve their problems, with the understanding that they had an open data policy and anybody could access the data dump easily to build things with it. It pisses me off that they are now trying to lock that information away where LLMs can’t access it. The whole reason to contribute is to help people. Locking that information away instead of exploiting this new channel to help people more effectively is antithetical to the reason I contributed in the first place.

Re: GPTBot – OpenAI’s Web Crawler

#182
post #16

Earlier quoted context omitted.

Counterpoint (not just to be annoying — I think you pose a very interesting unanswered question): If I read your hobby website about photography and use it to take 1% better pictures, do I owe you 1% of what my clients pay me? I think that probably most people would say no, assuming you could even determine that 1% in a way that both parties agreed was fair. I think generally, we have an understanding that some stuff…

Interesting point though I'd go with another analogy. You can go to a library to borrow a book, but you can't go to the library and copy all the books for your own use.

> You can go to a library to borrow a book, but you can't go to the library and copy all the books for your own use.

Why? What's stopping me from doing that? The only limitation is time.

Re: GPTBot – OpenAI’s Web Crawler

#183
post #80

Earlier quoted context omitted.

They are business, not some random visitor. Google scrapes websites all the time, provide Adsense, a way to earn money.

Adsense pays you money for showing ads to human visitors, you don't get paid for allowing their crawler.

You are letting them to crawl your site and earn money, they came up with a genius idea to pay you. Simple

Re: GPTBot – OpenAI’s Web Crawler

#184
post #180

Earlier quoted context omitted.

One option is to completely ban openai’s crawler ip addresses. They steal content without credit anyway - as most ai companies do - so there’s no benefit in allowing them access.

>so there’s no benefit in allowing them access. Well, you're helping improving the model.

Which benefits the company.

Re: GPTBot – OpenAI’s Web Crawler

#185
post #161

If I grab a copy of Adobe Photoshop (yeah, I know it runs in on a remote computer nowadays called 'the cloud') and I use it not to create the creative content its meant to be used for (manipulating cat pics, obviously) but to run it through IDA or Ghidra, or to study and use it to create a competitor (GIMP or make GIMP more like Photoshop) then even though I don't use it for its primary purpose; it is still copyright…

>it is still copyright infringement.

I doubt it's copyright infringement in this case, at most it's just against the ToS.

>Its making a derivative work

A derivative work includes major copyrightable elements of a first, previously created original work, and that's how it's treated in court. Most AIs will not generate derivative works (unless you ask them to).

Re: GPTBot – OpenAI’s Web Crawler

#186
post #180

Earlier quoted context omitted.

>so there’s no benefit in allowing them access. Well, you're helping improving the model.

Which benefits the company.

Yeah but I don't think it's an inherently bad thing. +100M people use ChatGPT without paying anything, in this it benefits much more than the company.

Re: GPTBot – OpenAI’s Web Crawler

#187
post #30

Earlier quoted context omitted.

Hoping this is what they’ll use to train future models and deprecate the older ones before the legal cases proceed any further.

The legal cases don't mean anything. The rule of law has all but disappeared from the corporate world. The idea that courts or regulators will be able to control AI is laughable. They are too corrupt, and they are way too slow.

(fortunately)

Re: GPTBot – OpenAI’s Web Crawler

#188
post #180

Earlier quoted context omitted.

>so there’s no benefit in allowing them access. Well, you're helping improving the model.

Which benefits the company.

If you're making a library/package/rubygem/crate, allowing ChatGPT to understand your API and being able to generate code using it can help the adoption.

Re: GPTBot – OpenAI’s Web Crawler

#189
post #119

Earlier quoted context omitted.

It’s so now they can lobby for anti scraping regulation and hamper any possible catch-up.

That would be a hilariously bad idea for them. Their business is based on fair use. The only way to enforce restrictions against scraping is through copyright law because obviously you can run the spidering code from any jurisdiction you want, so any law that says “thou shall not scrape” is toothless unless it acts through copyright. Any workable restrictions against using scraped data would also make ChatGPT illegal…

Nonsense. Regulation rarely works retroactively. Their model is trained and they have the money to license incremental data going forward, potentially exclusively.

Re: GPTBot – OpenAI’s Web Crawler

#190
post #180

Earlier quoted context omitted.

One option is to completely ban openai’s crawler ip addresses. They steal content without credit anyway - as most ai companies do - so there’s no benefit in allowing them access.

>so there’s no benefit in allowing them access. Well, you're helping improving the model.

That's of no benefit to me. Quite the contrary.
Post reply on HN