Live data from Hacker News

Google denies training Bard on ChatGPT chats from ShareGPT

twitter.com

281–290 of 342 posts

Re: Google denies training Bard on ChatGPT chats from ShareGPT

#281

So? First off, the whole argument behind these models has been from day one that training on copyrighted material is fair use. At most this would be a TOS violation. Second off, AI output is not subject to copyright, so it has even less protection than the original works it was trained on. Copyright maximalism for me, but not for thee. It's just so silly for someone working at OpenAI to complain about this.

> AI output is not subject to copyright The chats include human output too, which is presumably copyrighted, and is presumably necessary for training purposes.

OpenAI doesn't own the copyright on the human aspects of the chats, so it still doesn't really have a claim to make around them. And even if it did own that copyright, we loop right back around to "wait, training an AI on copyrighted material isn't fair use now?"

There's no way that ChatGPT's conversations are going to be subject to more intellectual property protection than the human chats it was trained on.

Re: Google denies training Bard on ChatGPT chats from ShareGPT

#282
post #214

So? First off, the whole argument behind these models has been from day one that training on copyrighted material is fair use. At most this would be a TOS violation. Second off, AI output is not subject to copyright, so it has even less protection than the original works it was trained on. Copyright maximalism for me, but not for thee. It's just so silly for someone working at OpenAI to complain about this.

> It's just so silly for someone working at OpenAI to complain about this. Who from OpenAI is complaining?

My understanding is that the Twitter thread author works at OpenAI. Maybe I'm wrong about that.

Re: Google denies training Bard on ChatGPT chats from ShareGPT

#283

Earlier quoted context omitted.

Take what action? Pretty sure that’s not illegal, especially since the training data is ai generated and therefore can’t be copyrighted.

Not illegal, but that won't stop people from finding it amusing that the company considered the world's beacon of innovation is copying someone else's homework. It's hard being the favorite horse.

tech companies steal ideas all the time. snapchat invented stories and now whatsapp, facebook, instagram, tiktok, youtube have them

Re: Google denies training Bard on ChatGPT chats from ShareGPT

#284
post #279

Earlier quoted context omitted.

That's an interesting analysis. The issue isn't really whether the A.I. has creative ability, though, if we're talking about whether it infringes copyright. I think comparing the A.I. to a really simple bot is informative. If I wrote a novel that contained once sentence from 1,000 people's novels, it would probably be fair use since I hardly took anything from any individual person and because my novel is probably no…

Your one sentence from one thousand works is likely seen as transformative. https://www.copyright.gov/fair-use/ > Additionally, “transformative” uses are more likely to be considered fair. Transformative uses are those that add something new, with a further purpose or different character, and do not substitute for the original use of the work. Creating a model from copyrighted works is likely sufficiently transformat…

"Creating a model from copyrighted works is likely sufficiently transformative to be non-infringing even if it is found to be a derivative work."

Maybe, but one of the factors of fair use is whether it deprives the copyright owner of income or undermines a new or potential market for the copyrighted work.

If ChatGPD gets so good at writing J.K. Rowling novels that it hurts the sales of the next J.K. Rowling book, that's a strong argument against the use being fair, even if it is transformative.

If J.K. Rowling signs an exclusive agreement with Google to train on J.K. Rowling novels, that's another factors that would suggest OpenAI's use is not fair, because she's shown OpenAI is hurting a potential market for J.K. Rowling selling the use of her novels to train A.I.

Re: Google denies training Bard on ChatGPT chats from ShareGPT

#285
post #88

1. Google denies doing it, so at the very least the title should have an "allegedly". 2. Even if they did – so what? The output from ChatGPT is not copyrightable by OpenAI. In fact it is OpenAI that is training its models on copyrighted data, pictures, code from all over the internet.

[deleted]

Re: Google denies training Bard on ChatGPT chats from ShareGPT

#286

Earlier quoted context omitted.

Which is a baseless hyperbole. We get it, blog spam is annoying. That doesn’t change the fact that humans generate a ton of data just interacting with one another online.

As a forum moderator, I have transitioned to relying heavily on AI-generated responses to users. These responses can range from short and concise ("Friendly reminder: please ensure that all content posted adheres to our rules regarding hate speech. Let's work together to maintain a safe and inclusive community for everyone") to lengthy explanations of underlying issues. By using AI-generated content, a small moderati…

> Original comment

The original comment is much better, please stop rewriting your comments using OpenAI.

> In fact, I did it for this comment.

Yes, it was obvious from the second sentence. The way ChatGPT structures text by default is very different from how most humans writes. Always the same "By using", "These X can range from" etc.

Padding your text with more words doesn't make it better, more words makes it worse, this isn't school.

Re: Google denies training Bard on ChatGPT chats from ShareGPT

#287
post #191

Earlier quoted context omitted.

Maybe, but we are fast approaching the point (or more likely have crossed it already) where distinguishing between human and AI generated data isn't really possible. If Google indexes a blog, how does it know whether it was written with AI assistance and therefore should not be used for training? Heck, how does OpenAI itself prevent such a feedback loop from its own output (or that of other LLMs)?

I'm only half joking.... I think we likely will end up with flags for human generated/curated content (and it will have to be that way round, as I can't imagine spammers bothering to put flags on AI-generated stuff), and we probably already should have an equivalent of robots.txt protocol that allows users to specify which parts of their website they would and wouldn't like used in the training of LLMs.

I think something like this will definitely happen, and your suggestion is the cleanest implementation idea I've seen for it. I imagine there will be a service provided by Google and OpenAI where they verify your identity as a human and then grant you a token to put into your meta tags (wait a second... this sounds like sama's worldcoin idea...).

It will need to be based somewhat on the honor system (just because someone's proved they're a human doesn't mean they won't put their attestation on auto-generated text), but it definitely sounds better than nothing.

They'll still need to incentivize it somehow, though. Why do I as a human want to add that meta tag? If the answer is "better search ranking" then it renders the whole scheme mostly pointless because obviously spammers will want to acquire the attestation and attach it to their auto-generated content.

Re: Google denies training Bard on ChatGPT chats from ShareGPT

#288

Earlier quoted context omitted.

>Even if they did – so what? Amplification of biases, propagation of errors, echolalia and over-optimization, lack of diverse data, overfitting

Not to mention it's embarrassing. Google playing second banana to OpenAI.

Many of Google's products were second, or later, to market. Google does not care.

Re: Google denies training Bard on ChatGPT chats from ShareGPT

#289

Earlier quoted context omitted.

Take what action? Pretty sure that’s not illegal, especially since the training data is ai generated and therefore can’t be copyrighted.

Not illegal, but that won't stop people from finding it amusing that the company considered the world's beacon of innovation is copying someone else's homework. It's hard being the favorite horse.

I don't think google has really been considered a beacon of innovation for a number of years.

Re: Google denies training Bard on ChatGPT chats from ShareGPT

#290
post #57

Earlier quoted context omitted.

> if you used ai to design a new superhero and then added pink shoes, yellow hair, and a beard Wouldn't that depend heavily on the prompt used (among other factors such as image to image and ControlNet)? You could be specifying lots of detail about the design in your prompt, and the AI could only be generating concept artwork with little variation from what you already provided. If I'm already providing the pose, the…

According to the copyright board a promot is not anymore than any person commissioning a work from an artist, which does not provide copyright, and the lack of human authorship for the design decisions still stops it from being protected by copyright.

Copyright.gov says the copyright office has started an initiative to examine AI copyright issues.

That makes me think even the copyright board is not confident of their answer.

Post reply on HN