Live data from Hacker News

Google denies training Bard on ChatGPT chats from ShareGPT

twitter.com

231–240 of 342 posts

Re: Google denies training Bard on ChatGPT chats from ShareGPT

#231
post #88

1. Google denies doing it, so at the very least the title should have an "allegedly". 2. Even if they did – so what? The output from ChatGPT is not copyrightable by OpenAI. In fact it is OpenAI that is training its models on copyrighted data, pictures, code from all over the internet.

if ChatGPT trained using Bard data, this site would be LIT UP because of OpenAI's association with Microsoft.

but it's google so no big deal right?

Re: Google denies training Bard on ChatGPT chats from ShareGPT

#232

Earlier quoted context omitted.

Which is a baseless hyperbole. We get it, blog spam is annoying. That doesn’t change the fact that humans generate a ton of data just interacting with one another online.

And how are you going to distinguish those interactions from chatbots trying to sell you something?

OpenAI at least can track the hashes of all content it's ever output, and filter that content out of future training data. Of course they won't be able to do this for the output of other LLMs, but maybe we'll see something like a federated bloom index or something.

Agreed there is no perfect solution though, and it will definitely be a problem finding high quality training data in the future.

Re: Google denies training Bard on ChatGPT chats from ShareGPT

#233

Earlier quoted context omitted.

It's definitely a derived work as far as copyright is concerned: the output would simply not exist without the copyrighted training data. > It's finding patterns same as anyone studying the code base would do. No, it's quite unlike anyone studying data, because it's not a person with legal rights, such as fair use, but an automated algorithm. There is absolutely no legal debate that copyright applies only to human au…

The output of human copyrighted work wouldn't exist if it weren't for humans training on the output of other humans. Humans constantly use cliches in their writing and speech, and most of what they produce is a repackaged version of what someone else has written or said, yet no one's up in arms against this mass of unoriginality as long as it's human-generated. This is anti-AI bias, pure and simple.

There is a difference between a computer and a human and we tried them already differently in copyright law. For example copying a program from disk into memory is typically already considered a copy on a computer (hence many licences grant you the licence to do this copy), no such licence is required for a human.

Re: Google denies training Bard on ChatGPT chats from ShareGPT

#234

Earlier quoted context omitted.

Are OpenAI saying they have adhered to the terms of service of all the content they have used?

Content is not subject to terms of service . Services are subject to terms of service. (If content is received through a service, the terms of service may govern use of it, but that’s not a feature of the content, but the acquisition route.)

Terms of Service, Terms and Conditions, and Terms of Use are all the same thing. There is no legal difference between them.

> that’s not a feature of the content, but the acquisition route.

It's neither. It's a feature of contract law.

Re: Google denies training Bard on ChatGPT chats from ShareGPT

#235
post #88

1. Google denies doing it, so at the very least the title should have an "allegedly". 2. Even if they did – so what? The output from ChatGPT is not copyrightable by OpenAI. In fact it is OpenAI that is training its models on copyrighted data, pictures, code from all over the internet.

> Google denies doing it

Read their statement carefully and it's actually not a denial of the allegation.

> But Google is firmly and clearly denying the data was used: “Bard is not trained on any data from ShareGPT or ChatGPT,” spokesperson Chris Pappas tells The Verge

* Allegation: Google used ShareGPT to train Bard.

* Rebuttal: The current production version of Bard is not trained on ShareGPT data

Both things can be true:

* Google did use ShareGPT to train Bard

* Bard is not currently trained on any data from ShareGPT or ChatGPT.

It depends on what the meaning of is is ;)

Re: Google denies training Bard on ChatGPT chats from ShareGPT

#236

Apart from the open questions of the quality of such once-removed-from-human-generated training data... I can't speak to the legality of the situation, but the morality of using, without their consent, data generated by someone's AI engine... ... that was, itself, trained on other people's data without their consent... ... should be, at the very least, equivalently evil to the original AI's training.

No, it shouldn't. Maybe you should be, at the very least, considered a questionable person. I do not in any way or form consider anything to be wrong with what they're doing, but I question the senses of someone thinking this is immoral or even evil. Keep your subjective nonsense out of this.

Every opinion is subjective.

Re: Google denies training Bard on ChatGPT chats from ShareGPT

#237

Earlier quoted context omitted.

> Albeit unethical and embarrassing. I really don’t understand this angle. In fact, I am fairly positive that the training set for GPT-4 contains many thousands of conversations with AI agents not developed by OpenAI. Do AI companies need to manually sift through the corpus and scrub webpages that contain competitor LLM output? (“Yes” is an acceptable answer to this, but then it applies to OpenAI’s currently existing…

How did you come about being "fairly positive" that GPT-4 is trained on other AI conversations?

Many AI conversations have been floating around internet forums since the original GPT was released. As OpenAI hasn't shared anything about its training set, to err on the side of caution I would assume that they didn't filter these conversations out. If they aren't even marked as such, it may not even be possible to do. I think it would be very hard to prove that no AI conversations are included in the training set, even if it wasn't secret.

Re: Google denies training Bard on ChatGPT chats from ShareGPT

#238
post #131
post #125

I love that OpenAI uses a ton of other peoples work to train their model, yet when someone uses OpenAI to train their model, they get all up in arms. As far as I'm concerned, OpenAI has decided terms of use don't exist anymore.

OpenAI is training on data that is against their terms of use? That reads like a serious allegation. What is this all about?

OpenAI is training on copyrighted data without a licence. I would argue copyright law has much stronger legal standing than some ToS.

Now OpenAI is arguing their training is fair use, but that has certainly not been legally established so far and could just as much be used as a defence against ToS violation.

So in short yes OpenAI is pretty much doing the same thing.

Re: Google denies training Bard on ChatGPT chats from ShareGPT

#239

"What's sauce for the goose is sauce for the gander" as the legal cliche goes. OpenAI cannot on the one hand claim that google did something wrong if they used their outputs as part of the bard training while simultaneously on the other hand claiming they themselves are free to use everyone on the internets content to train their model. Either they believe that training should respect copyright (in which case they co…

[flagged]

Re: Google denies training Bard on ChatGPT chats from ShareGPT

#240
post #88

1. Google denies doing it, so at the very least the title should have an "allegedly". 2. Even if they did – so what? The output from ChatGPT is not copyrightable by OpenAI. In fact it is OpenAI that is training its models on copyrighted data, pictures, code from all over the internet.

>Even if they did – so what? Amplification of biases, propagation of errors, echolalia and over-optimization, lack of diverse data, overfitting

Not to mention it's embarrassing. Google playing second banana to OpenAI.
Post reply on HN