Live data from Hacker News

Google denies training Bard on ChatGPT chats from ShareGPT

twitter.com

331–340 of 342 posts

Re: Google denies training Bard on ChatGPT chats from ShareGPT

#331
post #57

Earlier quoted context omitted.

> if you used ai to design a new superhero and then added pink shoes, yellow hair, and a beard Wouldn't that depend heavily on the prompt used (among other factors such as image to image and ControlNet)? You could be specifying lots of detail about the design in your prompt, and the AI could only be generating concept artwork with little variation from what you already provided. If I'm already providing the pose, the…

According to the copyright board a promot is not anymore than any person commissioning a work from an artist, which does not provide copyright, and the lack of human authorship for the design decisions still stops it from being protected by copyright.

Textual inversion involves providing self-created images, which should confer copyright in the same way AI images of DC's superman are considered to fall under the copyright of DC. In other words, commissioning fanart still allows the original owner of the IP to exert copyright -- shouldn't that be the case here?

Re: Google denies training Bard on ChatGPT chats from ShareGPT

#332

Earlier quoted context omitted.

The example was redrawing something by hand that was computer generated originally. It would be pretty much impossible for a hand drawn work of art to not be sufficiently original. Hand drawn art doesn't look the same as what a computer produces. Originality has a very low threshold, simply pointing my camera at something and hitting click is almost always enough to show originality. At any rate it isn't fraud to tak…

If you think you own a design because you hand drew a version of it someone else invented, you're gonna have a bad time. Please redraw a superman picture someone else made and then go to have it copyrighted, and tell me how that goes for you.

[deleted]

Re: Google denies training Bard on ChatGPT chats from ShareGPT

#333

Earlier quoted context omitted.

The example was redrawing something by hand that was computer generated originally. It would be pretty much impossible for a hand drawn work of art to not be sufficiently original. Hand drawn art doesn't look the same as what a computer produces. Originality has a very low threshold, simply pointing my camera at something and hitting click is almost always enough to show originality. At any rate it isn't fraud to tak…

If you think you own a design because you hand drew a version of it someone else invented, you're gonna have a bad time. Please redraw a superman picture someone else made and then go to have it copyrighted, and tell me how that goes for you.

It would go fine, since I see the form has a question "is this a derivative work". I put yes, and this means my claim is only for what was original to me when I drew the drawing based on another drawing of Superman.

But I see we've moved away from the orignal point that it would be difficult for anybody to know an AI helped someone make the drawing if they redrew and didn't disclose it was a redrawing.

Re: Google denies training Bard on ChatGPT chats from ShareGPT

#334
post #255
post #88

1. Google denies doing it, so at the very least the title should have an "allegedly". 2. Even if they did – so what? The output from ChatGPT is not copyrightable by OpenAI. In fact it is OpenAI that is training its models on copyrighted data, pictures, code from all over the internet.

Ok, I've added that information to the title—thanks. There's also https://www.theverge.com/2023/3/29/23662621/google-bard-chat... . Unfortunately the original report ( https://www.theinformation.com/articles/alphabets-google-and... ) is hardwalled.

The Verge doesn’t have an Information sub?

Re: Google denies training Bard on ChatGPT chats from ShareGPT

#335
post #309

Earlier quoted context omitted.

That's okay if they can leapfrog. Google Maps wasn't the first map product. We all used mapquest way before. But Google Maps was technologically advanced. Ajax made maps usable for the first time. Gmail wasn't the first webmail. Hotmail had millions of customers already. But Google gave people unlimited space to store old email, whereas email in the old days filled up your inboxes and needed to be deleted. Question i…

>But Google gave people unlimited space to store old email This isn't correct. Gmail launched with 1GB per user, which was way higher than other services, and they did keep doubling the storage space year-after-year, but it was never unlimited until Google Apps offered unlimited storage for businesses and schools.

That is true. You're technically correct.

In 2004 1GB of email was effectively unlimited though. Keep in mind hotmail offered 2MB and only increased that 125x when Gmail launched.

https://www.cnet.com/tech/services-and-software/hotmail-to-o...

Re: Google denies training Bard on ChatGPT chats from ShareGPT

#336

Earlier quoted context omitted.

>Even if they did – so what? Amplification of biases, propagation of errors, echolalia and over-optimization, lack of diverse data, overfitting

Not to mention it's embarrassing. Google playing second banana to OpenAI.

Ascribing actual human emotion to a giant corporation like a google is probably not a good idea. Their motivations aren't going be heavily dictated by feelings of shame at being late out of the gate.

Re: Google denies training Bard on ChatGPT chats from ShareGPT

#337

Earlier quoted context omitted.

Worst case scenario they just start only training on pre-2020 data and then finetuning on a dataset which they somehow know to be 'clean'. In practice though I doubt that AI contamination is actually a problem. Otherwise how would e.g. AlphaZero work so well (which is effectively only trained on its own data).

The parallels with AlphaZero are not so easy. The problem is you need some sort of arbiter of who has "won" a conversation but if the arbiter is just another transformer emitting a score, the models will compete to match the incomplete picture of reasoning given by the arbiter.

We have that though, just train on the buzzfeed articles that get the most attention or the tweets that get the most likes, etc.

Re: Google denies training Bard on ChatGPT chats from ShareGPT

#338
post #294

Earlier quoted context omitted.

"Creating a model from copyrighted works is likely sufficiently transformative to be non-infringing even if it is found to be a derivative work." Maybe, but one of the factors of fair use is whether it deprives the copyright owner of income or undermines a new or potential market for the copyrighted work. If ChatGPD gets so good at writing J.K. Rowling novels that it hurts the sales of the next J.K. Rowling book, tha…

But the model doesn't have any agency. GPT isn't spitting out novels in the style of J.K. Rowling and sending them to publishers - a human is. GPT being instructed to tell a Harry Potter story itself is no more infringing than a child asking a parent for a made up Harry Potter bed time story. They equally infringe and undermine new or potential markets for copyrighted work. The question is "what do you do with the ma…

"But the model doesn't have any agency"

This is a red herring. The issue before the court will be whether creation and release of the model effects J.K. Rowling.

Think of it this way- suppose I make a bunch of super mario brothers video games and try to sell them without Nintendo's permision.

If Nintendo sued me, I can't say "This cartridge has no agency. This will only effect Nintendo if humans use this video game instead of playing super mario brothers."

Re: Google denies training Bard on ChatGPT chats from ShareGPT

#339

People complained that new AI is "stealing" from artists. But stealing from other AI turns out to often be easier. And this is where things get fun, because companies like OpenAI want to be able to train on all the data without any explicit permissions from the creators, but the moment people do the same to them they likely (we will see) be very much against it. So it will be interesting if they will be able to both…

I can see the GitHub Copilot controversy being resolved in this way. If Microsoft, GitHub, and OpenAI successfully use the fair use defense for Copilot's appropriation of proprietary and incompatibly licensed code, then a free and open source alternative to Copilot can be trained on Copilot's outputs. After all, the GitHub Copilot Product Specific Terms say: > 2. Ownership of Suggestions and Your Code > GitHub does n…

Why would it need to be trained on Copilot’s output? Its training data is publicly available code on GitHub, so just use that directly. ChatGPT is different because they specifically trained it as an assistant with a private dataset

Re: Google denies training Bard on ChatGPT chats from ShareGPT

#340
I don't do this stuff at the training level, I just [ask AIs to] make pictures and stories where horrible things happen to people I do not like.

That said, given that everything that came out of ChatGPT is processed inputs from the real world, wouldn't feeding that output into training another AI basically be some weird new combination of coprophagy and inbreeding (in a digital sense)?

Post reply on HN