Live data from Hacker News

Google denies training Bard on ChatGPT chats from ShareGPT

twitter.com

271–280 of 342 posts

Re: Google denies training Bard on ChatGPT chats from ShareGPT

#271

So? First off, the whole argument behind these models has been from day one that training on copyrighted material is fair use. At most this would be a TOS violation. Second off, AI output is not subject to copyright, so it has even less protection than the original works it was trained on. Copyright maximalism for me, but not for thee. It's just so silly for someone working at OpenAI to complain about this.

> AI output is not subject to copyright

The chats include human output too, which is presumably copyrighted, and is presumably necessary for training purposes.

Re: Google denies training Bard on ChatGPT chats from ShareGPT

#272
post #191

Earlier quoted context omitted.

Maybe, but we are fast approaching the point (or more likely have crossed it already) where distinguishing between human and AI generated data isn't really possible. If Google indexes a blog, how does it know whether it was written with AI assistance and therefore should not be used for training? Heck, how does OpenAI itself prevent such a feedback loop from its own output (or that of other LLMs)?

I'm only half joking.... I think we likely will end up with flags for human generated/curated content (and it will have to be that way round, as I can't imagine spammers bothering to put flags on AI-generated stuff), and we probably already should have an equivalent of robots.txt protocol that allows users to specify which parts of their website they would and wouldn't like used in the training of LLMs.

Reminds me of the old “evil bit” RFC[1]

[1] https://www.ietf.org/rfc/rfc3514.txt

Re: Google denies training Bard on ChatGPT chats from ShareGPT

#273

Earlier quoted context omitted.

Not to mention it's embarrassing. Google playing second banana to OpenAI.

That assumes that training on the output of another language model somehow gives you the ability to improve your model and to catch up somehow

Well, it does, that's how we got Alpaca from LLaMA.

Re: Google denies training Bard on ChatGPT chats from ShareGPT

#274
post #191

Earlier quoted context omitted.

Maybe, but we are fast approaching the point (or more likely have crossed it already) where distinguishing between human and AI generated data isn't really possible. If Google indexes a blog, how does it know whether it was written with AI assistance and therefore should not be used for training? Heck, how does OpenAI itself prevent such a feedback loop from its own output (or that of other LLMs)?

I'm only half joking.... I think we likely will end up with flags for human generated/curated content (and it will have to be that way round, as I can't imagine spammers bothering to put flags on AI-generated stuff), and we probably already should have an equivalent of robots.txt protocol that allows users to specify which parts of their website they would and wouldn't like used in the training of LLMs.

If content with a "human-generated" flag is rated more highly in some way -- e.g. search results -- then of course spammers will automatically add that flag to their AI-generated garbage. How do you propose to prevent them?

Re: Google denies training Bard on ChatGPT chats from ShareGPT

#275

Earlier quoted context omitted.

>Even if they did – so what? Amplification of biases, propagation of errors, echolalia and over-optimization, lack of diverse data, overfitting

Not to mention it's embarrassing. Google playing second banana to OpenAI.

Google’s been second banana to openai for a few years now, right?

Re: Google denies training Bard on ChatGPT chats from ShareGPT

#276
post #88

1. Google denies doing it, so at the very least the title should have an "allegedly". 2. Even if they did – so what? The output from ChatGPT is not copyrightable by OpenAI. In fact it is OpenAI that is training its models on copyrighted data, pictures, code from all over the internet.

> Google denies doing it Read their statement carefully and it's actually not a denial of the allegation. > But Google is firmly and clearly denying the data was used: “Bard is not trained on any data from ShareGPT or ChatGPT,” spokesperson Chris Pappas tells The Verge * Allegation: Google used ShareGPT to train Bard. * Rebuttal: The current production version of Bard is not trained on ShareGPT data Both things can b…

Trained would mean the current model wasn't trained at all from ShareGPT data, not that was trained on it previously, and isn't being trained anymore.

This association makes no sense.

Re: Google denies training Bard on ChatGPT chats from ShareGPT

#277
post #274

Earlier quoted context omitted.

I'm only half joking.... I think we likely will end up with flags for human generated/curated content (and it will have to be that way round, as I can't imagine spammers bothering to put flags on AI-generated stuff), and we probably already should have an equivalent of robots.txt protocol that allows users to specify which parts of their website they would and wouldn't like used in the training of LLMs.

If content with a "human-generated" flag is rated more highly in some way -- e.g. search results -- then of course spammers will automatically add that flag to their AI-generated garbage. How do you propose to prevent them?

I assume, like the actual meta generator tags, it wouldn't actually be a massive boon for regular search results

Re: Google denies training Bard on ChatGPT chats from ShareGPT

#279

Earlier quoted context omitted.

United States copyright law in quite clear on the matter: >A "derivative work" is a work based upon one or more preexisting works, such as a translation, musical arrangement, dramatization, fictionalization, motion picture version, sound recording, art reproduction, abridgment, condensation, or any other form in which a work may be recast, transformed, or adapted . The emphasis part clearly applies: not only the AI m…

That's an interesting analysis. The issue isn't really whether the A.I. has creative ability, though, if we're talking about whether it infringes copyright. I think comparing the A.I. to a really simple bot is informative. If I wrote a novel that contained once sentence from 1,000 people's novels, it would probably be fair use since I hardly took anything from any individual person and because my novel is probably no…

Your one sentence from one thousand works is likely seen as transformative.

https://www.copyright.gov/fair-use/

> Additionally, “transformative” uses are more likely to be considered fair. Transformative uses are those that add something new, with a further purpose or different character, and do not substitute for the original use of the work.

Creating a model from copyrighted works is likely sufficiently transformative to be non-infringing even if it is found to be a derivative work.

Post reply on HN