Earlier quoted context omitted.
Nonsense. Regulation rarely works retroactively. Their model is trained and they have the money to license incremental data going forward, potentially exclusively.
Copyright laws do in fact (or have in fact) acted retroactively.
GPTBot – OpenAI’s Web Crawler
221–230 of 327 posts
Re: GPTBot – OpenAI’s Web Crawler
#222Earlier quoted context omitted.
If they implemented this properly, they should be retroactively filtering all their content that is no longer allowed in the robots.txt, or carries the #NoAI tag.
My understanding is that it's not easy to untrain a model of data already fed to it. Regarding noai tags - is this respected or just wishful?
Re: GPTBot – OpenAI’s Web Crawler
#223Earlier quoted context omitted.
That's of no benefit to me. Quite the contrary.
Why? If it helps people it should be good. Why bother posting something on the public web if not to help people. Sure a large org is receiving some ancillary benefit, but do you feel the same hostility for people working at [large corp] using what you worked on to help them at work? I honestly don't understand the hostility towards llms using public data
Re: GPTBot – OpenAI’s Web Crawler
#224Re: GPTBot – OpenAI’s Web Crawler
#225What’s the incentive for people to allow the crawler at all? Unlike search engines, chatgpt doesn’t cite references at all (last I tried) or even if it does it often makes up nonexistent references. And because it rephrases the content, there’s often no way to prove they got the material from a particular source, so harder to litigate plagiarism too. How would contributing to the weights of this LLM help content crea…
> What’s the incentive for people to allow the crawler at all? So that LLMs can learn from it? Profit is not the only thing that motivates people. I’ve spent years contributing to Stack Overflow to help people solve their problems, with the understanding that they had an open data policy and anybody could access the data dump easily to build things with it. It pisses me off that they are now trying to lock that infor…
Re: GPTBot – OpenAI’s Web Crawler
#226Re: GPTBot – OpenAI’s Web Crawler
#227Earlier quoted context omitted.
How's that gonna work when they need to update their model? Also, how would they compete with companies like FB that have an insane amount of conversational data, or Google, a company that literally indexes the internet?
Spend money on licensing deals, lock out the competition. The value of the LLM isn’t up to date data, it’s the concepts of extracts. There’s very limited value in a large amount of crap if chinchilla is to be believed. I don’t think stack overflow is all that valuable once your model has access to github due to their good friends at MS. The money in proprietary AI is on the top end now, open source / edge is destroyi…
As a heavy ChatGPT user I disagree. Lack of up to date data is one of the biggest issues I face every day - technology changes fast, libraries change APIs, new tech comes out, etc.
Re: GPTBot – OpenAI’s Web Crawler
#228Earlier quoted context omitted.
Nonsense. Regulation rarely works retroactively. Their model is trained and they have the money to license incremental data going forward, potentially exclusively.
Copyright laws do in fact (or have in fact) acted retroactively.
Meta has been lobbying hard around that for years.
Re: GPTBot – OpenAI’s Web Crawler
#229Earlier quoted context omitted.
Spend money on licensing deals, lock out the competition. The value of the LLM isn’t up to date data, it’s the concepts of extracts. There’s very limited value in a large amount of crap if chinchilla is to be believed. I don’t think stack overflow is all that valuable once your model has access to github due to their good friends at MS. The money in proprietary AI is on the top end now, open source / edge is destroyi…
> The value of the LLM isn’t up to date data As a heavy ChatGPT user I disagree. Lack of up to date data is one of the biggest issues I face every day - technology changes fast, libraries change APIs, new tech comes out, etc.
Re: GPTBot – OpenAI’s Web Crawler
#230Earlier quoted context omitted.
> The value of the LLM isn’t up to date data As a heavy ChatGPT user I disagree. Lack of up to date data is one of the biggest issues I face every day - technology changes fast, libraries change APIs, new tech comes out, etc.
I’m saying if they feed the source into ChatGPT (from their friends at Github) they have everything they need already.