Hard to believe that is true, or else Bard would probably not perform so bad.
The logs on ShareGPT are merely a drop in the bucket.
51–60 of 342 posts
Hard to believe that is true, or else Bard would probably not perform so bad.
The logs on ShareGPT are merely a drop in the bucket.
Thankfully archive.org exists, otherwise it would not be possible to get good training data in a few years when the internet is flooded with AI content.
Isn't most of the internet available through common crawl? I don't know what percentage of training data is just that data set but i assume it's enough for anyone with enough compute and ingenuity to create a reasonable LLM
How the turn tables. Remember when Google called out Microsoft in 2011 for using Google results? https://googleblog.blogspot.com/2011/02/microsofts-bing-uses... >We look forward to competing with genuinely new search algorithms out there—algorithms built on core innovation, and not on recycled search results from a competitor.
First off, the whole argument behind these models has been from day one that training on copyrighted material is fair use. At most this would be a TOS violation. Second off, AI output is not subject to copyright, so it has even less protection than the original works it was trained on.
Copyright maximalism for me, but not for thee. It's just so silly for someone working at OpenAI to complain about this.
Earlier quoted context omitted.
Our writing, our code, our artwork... Furthermore, the U.S. Copyright Office (USCO) concluded that AI-generated works on their own cannot be copyright, so these ChatGPT logs are free game. It would be hypocritical to think that Google is wrong and OpenAI is not.
its not even that on their own those works cant be copywritten. its that even when you make changes to those works, your changes might qualify for copyright but they do not affect the copyright status of the ai generated portions of the work. if you used ai to design a new superhero and then added pink shoes, yellow hair, and a beard, only those three elements would possibly be able to be protected by copywrite. your…
Wouldn't that depend heavily on the prompt used (among other factors such as image to image and ControlNet)? You could be specifying lots of detail about the design in your prompt, and the AI could only be generating concept artwork with little variation from what you already provided.
If I'm already providing the pose, the face, and the outfit for a character (say via ControlNet and Textual Inversion), generating should be no different from generating , that is to say, the copyright already exists thanks to my work and the AI is just a tool, the output of which should have no bearing on who owns that copyright (DC is going to be perfectly able to challenge my commercial use of AI generated superman artwork).
Hard to believe that is true, or else Bard would probably not perform so bad.
Google only has a fraction of the training data. OpenAI had a huge head start and has been collecting training data for years now. ChatGPT is also wildly popular which has given them tons more training data. It's estimated that ChatGPT gained over 100 million users in the first two months alone, and may have over 13 million active users daily. The logs on ShareGPT are merely a drop in the bucket.
Uh, what? The same Google that has been crawling, indexing, and letting people search the entire Internet for the last 25 years? They have owned DeepMind for nearly twice as long as OpenAI has been in existence!
If anything this is proof that no one at Google can get anything done anymore, and lack of training data ain't the problem.
Good luck to them. AI models are automated plagiarism, top to bottom. None of us gave OpenAI permission to derive their model from our writing, surely billions of dollars worth, but they took it anyway. Copyright hasn't caught up so all that stolen value rests securely with OpenAI. If we're not getting that back, I don't see why AI competitors should have any qualms about borrowing each others' work.
What companies like OpenAI want is a system where everything they build is protected, and nothing that anyone else builds is protected. It's wildly hypocritical, what's good for the goose is good for the gander.
That some AI proponents are now freaking out about how model output can be legally used shows that on some level those people weren't really honestly engaging with artists who were freaking out about their work being appropriated to copy them. It's all just "learning from the art" until it affects somebody's competitive moat, and then suddenly people do understand how LLM weights could be seen as a derivative work of their inputs.
According to the article, the story goes this way: This engineer Jacob Devlin raised his concerns on training Bard with ShareGPT data. Then he directly joined OpenAI. He also claims that Google were about to do it, and then they stopped after his warnings. And presumably removed every trace of openai's responses. A couple of things: 1. So, Bard could have been trained on ShareGPT but it's not - according to the same…