Earlier quoted context omitted.
Hardly unethical, considering OpenAI is doing exactly this.
You can quibble about the ethics of web scraping for ML in general but I think you're conflating issues. OpenAI and Google both scour the web for human-generated content. What Google cares about here is the learnings from OpenAI's proprietary RLHF dataset, for which they had to contract a large sum of human labelers. Finding a roundabout way to extract the value of a direct competitor's purpose-built, costly data fee…
Google denies training Bard on ChatGPT chats from ShareGPT
201–210 of 342 posts
Re: Google denies training Bard on ChatGPT chats from ShareGPT
#202Earlier quoted context omitted.
The clauses always have a trap door: "[outside of] our provision of the Service" means they can do anything as long as it's a service they provide. Under definitions: The “Service” refers to the applications, software, products, and services provided by GitHub, including any Beta Previews.
I think there's a misunderstanding over what the word "relinquish" means. The terms make clear that uploading code to GitHub gives GitHub the right to "store, archive, parse, and display Your Content, and make incidental copies, as necessary to provide the Service, including improving the Service over time" while the code is hosted on GitHub. However, that's not the same thing as relinquishing (giving up) licensing r…
GitHub would argue that it is, and they'd likely argue that charging for access to copilot is akin to charging for access to private repositories.
Others would say that copilot is somehow separate from the services Github provides, so using their code for CoPilot wouldn't be covered by the ToS.
Re: Google denies training Bard on ChatGPT chats from ShareGPT
#203Re: Google denies training Bard on ChatGPT chats from ShareGPT
#204Earlier quoted context omitted.
Google has no contract with OpenAI though. They used a third party site to scrape conversations. If the outputs themselves are not copyrighted, and they never agreed to the terms of service, it should be fine, right? Albeit unethical and embarrassing.
> Albeit unethical and embarrassing. I really don’t understand this angle. In fact, I am fairly positive that the training set for GPT-4 contains many thousands of conversations with AI agents not developed by OpenAI. Do AI companies need to manually sift through the corpus and scrub webpages that contain competitor LLM output? (“Yes” is an acceptable answer to this, but then it applies to OpenAI’s currently existing…
Re: Google denies training Bard on ChatGPT chats from ShareGPT
#205Earlier quoted context omitted.
Hardly unethical, considering OpenAI is doing exactly this.
You can quibble about the ethics of web scraping for ML in general but I think you're conflating issues. OpenAI and Google both scour the web for human-generated content. What Google cares about here is the learnings from OpenAI's proprietary RLHF dataset, for which they had to contract a large sum of human labelers. Finding a roundabout way to extract the value of a direct competitor's purpose-built, costly data fee…
Yes, this latest instance with OpenAI outputs is shady, but I think it's in the same spirit as scraping news organizations for content which journalists were paid to write, and then showing portions of it directly in response to queries so people don't go directly to the news organization's pages, and it's in the same spirit as showing answers to query-questions that are excerpts from scraped pages which another organization paid to produce.
Re: Google denies training Bard on ChatGPT chats from ShareGPT
#206Earlier quoted context omitted.
> It's definitely a derived work as far as copyright is concerned - the output would simply not exist without the copyrighted training data. Can you point to a legal case that confirms this? Because it’s not at all clear that this is true from a legal standpoint. “X would not exist without Y” is not a sufficient test for derivative works - it’s far more nuanced.
United States copyright law in quite clear on the matter: >A "derivative work" is a work based upon one or more preexisting works, such as a translation, musical arrangement, dramatization, fictionalization, motion picture version, sound recording, art reproduction, abridgment, condensation, or any other form in which a work may be recast, transformed, or adapted . The emphasis part clearly applies: not only the AI m…
If I wrote a novel that contained once sentence from 1,000 people's novels, it would probably be fair use since I hardly took anything from any individual person and because my novel is probably not harming those other writers.
If I wrote a bot that did the same thing, same result, because my bot uses only a little from everyone's novel and doesn't harm the original novelist, so it's likely fair use.
Now I think a J.K. Rowling A.I. probably takes at least a little from her when it produces output, but it's not clear to me how much is actually based on J.K. Rowling and how much is a dataset of how words tend to be associated with other words. You could design a J.K. Rowling A.I. that uses nothing from J.K. Rowling, just data that is said to be J.K. Rowling-esque.
Re: Google denies training Bard on ChatGPT chats from ShareGPT
#207Earlier quoted context omitted.
I think there's a misunderstanding over what the word "relinquish" means. The terms make clear that uploading code to GitHub gives GitHub the right to "store, archive, parse, and display Your Content, and make incidental copies, as necessary to provide the Service, including improving the Service over time" while the code is hosted on GitHub. However, that's not the same thing as relinquishing (giving up) licensing r…
The question turns on whether you consider copilot part of the "GitHub service." GitHub would argue that it is, and they'd likely argue that charging for access to copilot is akin to charging for access to private repositories. Others would say that copilot is somehow separate from the services Github provides, so using their code for CoPilot wouldn't be covered by the ToS.
I'll repeat the definition of service: The “Service” refers to the applications, software, products, and services provided by GitHub, including any Beta Previews.
Re: Google denies training Bard on ChatGPT chats from ShareGPT
#208There is also a rumor that there has been a falling out between Google and Deepmind so I’m wondering what the story is there.
Re: Google denies training Bard on ChatGPT chats from ShareGPT
#209Apart from the open questions of the quality of such once-removed-from-human-generated training data... I can't speak to the legality of the situation, but the morality of using, without their consent, data generated by someone's AI engine... ... that was, itself, trained on other people's data without their consent... ... should be, at the very least, equivalently evil to the original AI's training.
Re: Google denies training Bard on ChatGPT chats from ShareGPT
#210Earlier quoted context omitted.
The clauses always have a trap door: "[outside of] our provision of the Service" means they can do anything as long as it's a service they provide. Under definitions: The “Service” refers to the applications, software, products, and services provided by GitHub, including any Beta Previews.
I think there's a misunderstanding over what the word "relinquish" means. The terms make clear that uploading code to GitHub gives GitHub the right to "store, archive, parse, and display Your Content, and make incidental copies, as necessary to provide the Service, including improving the Service over time" while the code is hosted on GitHub. However, that's not the same thing as relinquishing (giving up) licensing r…