Earlier quoted context omitted.
How could that be ever really be enforceable? If I use an AI tool to design my Superhero, can't I just submit it without disclosing the help I received from an AI. I get that it would be very nice to prevent AI SPAM copyrighting of every possible superhero, but if I use the AI to come up with a concept, then quickly redraw it myself with pen and paper, I feel like it would never be provable that it came from an AI.
you would be committing fraud. what happens if a criminal commits fraud?
Google denies training Bard on ChatGPT chats from ShareGPT
291–300 of 342 posts
Re: Google denies training Bard on ChatGPT chats from ShareGPT
#292Earlier quoted context omitted.
you would be committing fraud. what happens if a criminal commits fraud?
Redrawing something by hand creates a new copyrightable work, so it certainly isn't fraud to claim you own the copyright in a work of art you drew based on an AI output.
Re: Google denies training Bard on ChatGPT chats from ShareGPT
#293even if true which it does not seem to be the case, the whole thing sounds pretty marginal, in order to train a model that is most likely significantly bigger than 100b parameters, one also needs orders of magnitude more training data than the small 120k chat that were shared on the ShareGPT website
Such logs would not be used for training the base model, but rather for fine-tuning the model for instruction following. Instruction tuning requires far less data than is needed for pre-training the foundation model. Stanford Alpaca showed surprisingly strong results from fine-tuning Meta's LLaMA model on just 52k ChatGPT-esque interactions ( https://crfm.stanford.edu/2023/03/13/alpaca.html ).
"The cat is finally out of the bag – Google relied heavily on @ShareGPT 's data when training Bard.
This was also why we took down ShareGPT's Explore page – which has over 112K shared conversations – last week.
Insanity."
Fine-tunning is not exactly the same as "relying heavily", I bet they got way more fine-tunning data from simply asking their 100k employees to pre-beta test for a couple of months
Re: Google denies training Bard on ChatGPT chats from ShareGPT
#294Earlier quoted context omitted.
Your one sentence from one thousand works is likely seen as transformative. https://www.copyright.gov/fair-use/ > Additionally, “transformative” uses are more likely to be considered fair. Transformative uses are those that add something new, with a further purpose or different character, and do not substitute for the original use of the work. Creating a model from copyrighted works is likely sufficiently transformat…
"Creating a model from copyrighted works is likely sufficiently transformative to be non-infringing even if it is found to be a derivative work." Maybe, but one of the factors of fair use is whether it deprives the copyright owner of income or undermines a new or potential market for the copyrighted work. If ChatGPD gets so good at writing J.K. Rowling novels that it hurts the sales of the next J.K. Rowling book, tha…
GPT isn't spitting out novels in the style of J.K. Rowling and sending them to publishers - a human is.
GPT being instructed to tell a Harry Potter story itself is no more infringing than a child asking a parent for a made up Harry Potter bed time story. They equally infringe and undermine new or potential markets for copyrighted work.
The question is "what do you do with the material?" If a human took the output of GPT writing as J.K. Rowling or a parent took their collected Harry Potter bedtime stories - those are equally problematic.
If I was to take a portrait of Marilyn Monroe and send it through a plugin called Warholize in Photoshop ( https://www.adobe.com/creativecloud/photography/hub/guides/c... ) , it's not the plugin or photoshop that is infringing - it would be me, the human who created an infringing work. If I print it out and hang it on my wall, that didn't particularly impact on the income for the Warhol estate nor deprive them of new markets. If I print out copies of it and sell them - then that is a different matter.
The question is what you - the human with agency - do with the infringing work after you create it. You can't blame photoshop for creating a Warhol infringing work nor can you blame GPT for writing in the style of J.K. Rowling if you instruct it to do so.
Re: Google denies training Bard on ChatGPT chats from ShareGPT
#295I don't care at all about this from a copyright or data ownership perspective, but I am a little skeptical that it's a good idea to be this incestuous with training data in the long run. It's one thing to do fine tuning or knowledge distillation for specialized domains or shrinking models. But if you're trying to train your own foundation model, is relying on output from other foundation models going to make them lea…
Re: Google denies training Bard on ChatGPT chats from ShareGPT
#296Earlier quoted context omitted.
And how are you going to distinguish those interactions from chatbots trying to sell you something?
A network of trust, backed by a social graph, which can be used to filter untrusted content.
Re: Google denies training Bard on ChatGPT chats from ShareGPT
#297Earlier quoted context omitted.
Right, but training an LLM on the output of another LLM can certainly exacerbate these issues
Maybe, but we are fast approaching the point (or more likely have crossed it already) where distinguishing between human and AI generated data isn't really possible. If Google indexes a blog, how does it know whether it was written with AI assistance and therefore should not be used for training? Heck, how does OpenAI itself prevent such a feedback loop from its own output (or that of other LLMs)?
Re: Google denies training Bard on ChatGPT chats from ShareGPT
#298Earlier quoted context omitted.
Redrawing something by hand creates a new copyrightable work, so it certainly isn't fraud to claim you own the copyright in a work of art you drew based on an AI output.
It depends if your redrawing is substantially different enough from the original image to earn copyright on its own. Your changes to an image from ChatGPT do not affect the copyrightability of the original content. If you've simply redrawn what the computer designed it may not be substantial enough to earn copyright. If you've made changes, it may only be copyrightable for those changes.
It would be pretty much impossible for a hand drawn work of art to not be sufficiently original. Hand drawn art doesn't look the same as what a computer produces. Originality has a very low threshold, simply pointing my camera at something and hitting click is almost always enough to show originality.
At any rate it isn't fraud to take the legal position that you are an original enough artist to have copyright in the work. If taking a legal position was "fraud" any attorney who lost a court motion would be whisked away to jail.
Edited to add: the copyright registration form asks if you are the "author" not if you are "original."
Re: Google denies training Bard on ChatGPT chats from ShareGPT
#2991. Google denies doing it, so at the very least the title should have an "allegedly". 2. Even if they did – so what? The output from ChatGPT is not copyrightable by OpenAI. In fact it is OpenAI that is training its models on copyrighted data, pictures, code from all over the internet.
My stronger opinion is that the people who can do this stuff via having a crawled corpus of the Internet need to keep in mind that it's all our "user-generated content" that they've freely appropriated to build their models, and so whatever the technical copyright rules are (or become): you don't ethically own something that's closely imitating stuff we all wrote over the years.