If you scrape my hobby website about photography, scuba diving or let's say baking or gardening which improves your model by let's say a delta of 0.00000000001 than shouldn't I get some free credits to use that model or proportionate share in the revenue stream? EDIT: scuba diving NOT scooba diving
Counterpoint (not just to be annoying — I think you pose a very interesting unanswered question): If I read your hobby website about photography and use it to take 1% better pictures, do I owe you 1% of what my clients pay me? I think that probably most people would say no, assuming you could even determine that 1% in a way that both parties agreed was fair. I think generally, we have an understanding that some stuff…
GPTBot – OpenAI’s Web Crawler
121–130 of 327 posts
Re: GPTBot – OpenAI’s Web Crawler
#122What’s the incentive for people to allow the crawler at all? Unlike search engines, chatgpt doesn’t cite references at all (last I tried) or even if it does it often makes up nonexistent references. And because it rephrases the content, there’s often no way to prove they got the material from a particular source, so harder to litigate plagiarism too. How would contributing to the weights of this LLM help content crea…
Chatgpt 4 provides pretty good citations on request.
Re: GPTBot – OpenAI’s Web Crawler
#123What’s the incentive for people to allow the crawler at all? Unlike search engines, chatgpt doesn’t cite references at all (last I tried) or even if it does it often makes up nonexistent references. And because it rephrases the content, there’s often no way to prove they got the material from a particular source, so harder to litigate plagiarism too. How would contributing to the weights of this LLM help content crea…
Just one example: As a software vendor, you probably want OpenAI to index your documentation, so questions about your software can be answered by ChatGPT. Not everybody who creates content is a "content creator" (when did this word get the specific meaning "people who earn money or reputation from creating content").
On a second thought, I guess a lot of marketing content would also love to be crawled by anything that crawls…
Re: GPTBot – OpenAI’s Web Crawler
#124Earlier quoted context omitted.
Interesting point though I'd go with another analogy. You can go to a library to borrow a book, but you can't go to the library and copy all the books for your own use.
I used to go to the library, find books with the relevant chapters related to what I wanted to learn, and the librarian would photo copy all the pages I wanted to take home. So I guess technically you could copy all the books for your own use. It's just impractical to photocopy every page of every book in a library.
Re: GPTBot – OpenAI’s Web Crawler
#125I wonder how much the regression of ChatGPT is due to it adding new content which has its origin from ChatGPT. The blog and SEO spam with ChatGPT fluff is going through the roof, eventually all of that will get crawled too and the model will just get positively reinforced on its own output. Or is that not a concern?
My reasons are:
- I don't recall seeing any evidence that OpenAI has included new data in pretraining beyond the previous limit (Sept. 2021?) for GPT-3.5 or GPT-4
- Maybe they did finetuning or RLHF on new data but this is likely to be highly curated data
- AI generated content should be absolutely tiny in comparison to the data they are already working with.
Re: GPTBot – OpenAI’s Web Crawler
#126Earlier quoted context omitted.
Counterpoint (not just to be annoying — I think you pose a very interesting unanswered question): If I read your hobby website about photography and use it to take 1% better pictures, do I owe you 1% of what my clients pay me? I think that probably most people would say no, assuming you could even determine that 1% in a way that both parties agreed was fair. I think generally, we have an understanding that some stuff…
This is a very thought provoking point and it throughly stimulated me to think deeper and through. Purpose of my website is threefold, document my own knowledge, maybe some vanity and the urge to give back something to "someone" make a better living or similar. Things get interesting at corporate scale. There are fat VC funds, executives, board of directors and what not - making far more and far more comfortable than…
Re: GPTBot – OpenAI’s Web Crawler
#127Earlier quoted context omitted.
A human reads a website, watches/clicks on ad, buys merch, subscribes to website, sets a bookmark to a website, shares an article, invites others etc. What will be the point of sharing knowledge or content if it will no longer be associated to an individual or organization?
Some would argue that it allows for cutting edge research that could potentially massively benefit humanity. Whether or not that pans out is to be determined, but that is the stated goal. The point then, as advocates for this technology believe is for the betterment of civilization, that by training a neural network on this knowledge will make that knowledge more accessible to the rest of us. One could of course deba…
I repeat, what will be the motivation of an individual to share or provide valuable information if you decrease or eliminate any control of where and how that information appears?
Re: GPTBot – OpenAI’s Web Crawler
#128Earlier quoted context omitted.
My understanding is that it's not easy to untrain a model of data already fed to it. Regarding noai tags - is this respected or just wishful?
Like every time you put content on the internet: you depend on their good will to respect these tags, or robots.txt. OpenAI can decide to ignore it. It's wishful thinking.
However, it's trivial to know whether the bot crawled your site or stopped at robots.txt.
Re: GPTBot – OpenAI’s Web Crawler
#129Re: GPTBot – OpenAI’s Web Crawler
#130Earlier quoted context omitted.
The legal cases don't mean anything. The rule of law has all but disappeared from the corporate world. The idea that courts or regulators will be able to control AI is laughable. They are too corrupt, and they are way too slow.
I think a key idea is that with the amount of jurisdictions and number of courts the odds that a clean and sympathetic judge can be found approach one. I would argue that European jurisdictions are inherently less likely to be in pockets of American corporate interest and they are more likely to hear cases where fundamental human freedoms are at stake because both of these are existential threats to European independ…
Let's say the French government decides that OpenAI must change something about their business practices if they want to continue operating in France. OpenAI says "nope", and blocks access to French users.
Suddenly French companies aren't able to use GPT-X anymore – while their competitors in other countries can. How long do you think it will take before a storm of corporate outrage forces the government to relent?
Any individual government (except, perhaps, the combined US and EU governments) is powerless against today's technology megacorporations, because they can take much more away from a country than that country can take from them. If push ever comes to shove, it will become obvious where the true power lies. So far, the corporations have barely even tried to throw their weight around.