Live data from Hacker News

GPTBot – OpenAI’s Web Crawler

platform.openai.com

221–230 of 327 posts

Re: GPTBot – OpenAI’s Web Crawler

#221
post #189

Earlier quoted context omitted.

Nonsense. Regulation rarely works retroactively. Their model is trained and they have the money to license incremental data going forward, potentially exclusively.

Copyright laws do in fact (or have in fact) acted retroactively.

When? Not doubting, just curious about scope and type of scenarios where it's happened.

Re: GPTBot – OpenAI’s Web Crawler

#222
post #31
post #24

Earlier quoted context omitted.

If they implemented this properly, they should be retroactively filtering all their content that is no longer allowed in the robots.txt, or carries the #NoAI tag.

My understanding is that it's not easy to untrain a model of data already fed to it. Regarding noai tags - is this respected or just wishful?

It's easy for them to delete the model and start from scratch though.

Re: GPTBot – OpenAI’s Web Crawler

#223
post #209

Earlier quoted context omitted.

That's of no benefit to me. Quite the contrary.

Why? If it helps people it should be good. Why bother posting something on the public web if not to help people. Sure a large org is receiving some ancillary benefit, but do you feel the same hostility for people working at [large corp] using what you worked on to help them at work? I honestly don't understand the hostility towards llms using public data

There is no such thing as 'public data'. There is public domain, but data always belongs to someone if not expressed otherwise.

Re: GPTBot – OpenAI’s Web Crawler

#225

What’s the incentive for people to allow the crawler at all? Unlike search engines, chatgpt doesn’t cite references at all (last I tried) or even if it does it often makes up nonexistent references. And because it rephrases the content, there’s often no way to prove they got the material from a particular source, so harder to litigate plagiarism too. How would contributing to the weights of this LLM help content crea…

> What’s the incentive for people to allow the crawler at all? So that LLMs can learn from it? Profit is not the only thing that motivates people. I’ve spent years contributing to Stack Overflow to help people solve their problems, with the understanding that they had an open data policy and anybody could access the data dump easily to build things with it. It pisses me off that they are now trying to lock that infor…

By the same token (no pun intended), locking up such data in a closed (in many senses) LLM wouldn’t be a desirable outcome?

Re: GPTBot – OpenAI’s Web Crawler

#227
post #193

Earlier quoted context omitted.

How's that gonna work when they need to update their model? Also, how would they compete with companies like FB that have an insane amount of conversational data, or Google, a company that literally indexes the internet?

Spend money on licensing deals, lock out the competition. The value of the LLM isn’t up to date data, it’s the concepts of extracts. There’s very limited value in a large amount of crap if chinchilla is to be believed. I don’t think stack overflow is all that valuable once your model has access to github due to their good friends at MS. The money in proprietary AI is on the top end now, open source / edge is destroyi…

> The value of the LLM isn’t up to date data

As a heavy ChatGPT user I disagree. Lack of up to date data is one of the biggest issues I face every day - technology changes fast, libraries change APIs, new tech comes out, etc.

Re: GPTBot – OpenAI’s Web Crawler

#228
post #189

Earlier quoted context omitted.

Nonsense. Regulation rarely works retroactively. Their model is trained and they have the money to license incremental data going forward, potentially exclusively.

Copyright laws do in fact (or have in fact) acted retroactively.

It’s a red herring - There’s many ways to regulate scraping that don’t involve changing copyright.

Meta has been lobbying hard around that for years.

Re: GPTBot – OpenAI’s Web Crawler

#229
post #193

Earlier quoted context omitted.

Spend money on licensing deals, lock out the competition. The value of the LLM isn’t up to date data, it’s the concepts of extracts. There’s very limited value in a large amount of crap if chinchilla is to be believed. I don’t think stack overflow is all that valuable once your model has access to github due to their good friends at MS. The money in proprietary AI is on the top end now, open source / edge is destroyi…

> The value of the LLM isn’t up to date data As a heavy ChatGPT user I disagree. Lack of up to date data is one of the biggest issues I face every day - technology changes fast, libraries change APIs, new tech comes out, etc.

I’m saying if they feed the source into ChatGPT (from their friends at Github) they have everything they need already.

Re: GPTBot – OpenAI’s Web Crawler

#230
post #229

Earlier quoted context omitted.

> The value of the LLM isn’t up to date data As a heavy ChatGPT user I disagree. Lack of up to date data is one of the biggest issues I face every day - technology changes fast, libraries change APIs, new tech comes out, etc.

I’m saying if they feed the source into ChatGPT (from their friends at Github) they have everything they need already.

Oh. Hm, yeah, that sounds possible. We'll see. There are a lot of places besides Github where people talk about code.
Post reply on HN