Live data from Hacker News

Google denies training Bard on ChatGPT chats from ShareGPT

twitter.com

91–100 of 342 posts

Re: Google denies training Bard on ChatGPT chats from ShareGPT

#91
I hope they trained it on the insane ChatGPT conversations. Maybe it could be the very start of generated data ruining the ability to train these models on massive amounts of genuine human-created data. Hopefully the models will stagnate or regress because they're just training on older models' output.

Re: Google denies training Bard on ChatGPT chats from ShareGPT

#92
post #88

1. Google denies doing it, so at the very least the title should have an "allegedly". 2. Even if they did – so what? The output from ChatGPT is not copyrightable by OpenAI. In fact it is OpenAI that is training its models on copyrighted data, pictures, code from all over the internet.

This is an argument in bad faith but at this point I have zero trust in corporations and feel like you can generally count on them to do shitty things if they can benefit from it so I can be easily swayed by little proof at this point.

Re: Google denies training Bard on ChatGPT chats from ShareGPT

#93
post #90

Earlier quoted context omitted.

Take what action? Pretty sure that’s not illegal, especially since the training data is ai generated and therefore can’t be copyrighted.

OpenAI could have blocked Google's accounts, for example. Nothing really to do with legality.

No one is alleging that Google directly used OpenAI's API to get training data (which would be unambiguously against TOS). The claim is that they downloaded examples from ShareGPT.

Re: Google denies training Bard on ChatGPT chats from ShareGPT

#94

Earlier quoted context omitted.

Missed the point - they are saying that, in the future, there will be no human generated content left on the Internet.

Which is a baseless hyperbole. We get it, blog spam is annoying. That doesn’t change the fact that humans generate a ton of data just interacting with one another online.

And how are you going to distinguish those interactions from chatbots trying to sell you something?

Re: Google denies training Bard on ChatGPT chats from ShareGPT

#95
post #56

even if true which it does not seem to be the case, the whole thing sounds pretty marginal, in order to train a model that is most likely significantly bigger than 100b parameters, one also needs orders of magnitude more training data than the small 120k chat that were shared on the ShareGPT website

Such logs would not be used for training the base model, but rather for fine-tuning the model for instruction following. Instruction tuning requires far less data than is needed for pre-training the foundation model. Stanford Alpaca showed surprisingly strong results from fine-tuning Meta's LLaMA model on just 52k ChatGPT-esque interactions (https://crfm.stanford.edu/2023/03/13/alpaca.html).

Re: Google denies training Bard on ChatGPT chats from ShareGPT

#96

Earlier quoted context omitted.

Missed the point - they are saying that, in the future, there will be no human generated content left on the Internet.

Which is a baseless hyperbole. We get it, blog spam is annoying. That doesn’t change the fact that humans generate a ton of data just interacting with one another online.

I think their comment was meant to be taken as humour rather than a literal prediction.

Re: Google denies training Bard on ChatGPT chats from ShareGPT

#97
post #62

Earlier quoted context omitted.

Students in school also will not never learn to read without being exposed to text. Does this mean that teachers who write exercise sheets and school text book publishers now own the copyright of everything students do?

AI is not a human being or a student in school. It’s a software tool, stop comparing the two.

Being in school is also just a tool to knowing stuff, being able to read, and being around similar aged peers, etc.

Whether the knowledge is directly in your brain or in a device you operate (directly or through an API) shouldn't really matter.

If it's forbidden for a human to move a stone with manual labour, then it's also forbidden to move that stone with an excavator. This has nothing to do with the person being a human and the other person being an excavator controlled by a human: it's not authorized.

I think that we should allow humans to move stones up the hill with excavators too. There is no stealing of excavator fuel from human food sources going on (let's assume it's not biofuel operated :p).

Re: Google denies training Bard on ChatGPT chats from ShareGPT

#98
post #88

1. Google denies doing it, so at the very least the title should have an "allegedly". 2. Even if they did – so what? The output from ChatGPT is not copyrightable by OpenAI. In fact it is OpenAI that is training its models on copyrighted data, pictures, code from all over the internet.

This is an argument in bad faith but at this point I have zero trust in corporations and feel like you can generally count on them to do shitty things if they can benefit from it so I can be easily swayed by little proof at this point.

What's the argument? What's been done by anyone that's shitty? I don't even understand the point of this post. As far as I know, the current wave of text-based AIs is trained on all text accessible on the internet. Would it be a scandal to learn that ChatGPT is trained on wikipedia? Reddit? What is even the argument here, good faith or otherwise?

Re: Google denies training Bard on ChatGPT chats from ShareGPT

#99
post #83

Earlier quoted context omitted.

Building things and maintaining it as a trade secret can be protected as a trade secret. Trade secrets don't need to be copyrightable (e.g. list of customer numbers is a trade secret but not copyrightable). https://copyrightalliance.org/faqs/difference-copyright-pate... > Trade secret protection protects secrets from unauthorized disclosure and use by others. A trade secret is information that has an economic benefit…

I am not a lawyer, but I don’t believe a trade secret would prevent someone from reverse engineering your model’s knowledge from it’s output though, in the same way that it doesn’t prevent someone from reverse engineering your hot sauce from buying a bunch and experimenting with the ingredients until it tastes similar.

Yep, that's correct.

My point was more of there are protections for things that aren't copyrightable. If the model is protected as a trade secret, then it is a trade secret.

The example of the hot sauce recipe is quite apt - the recipe isn't copyrightable, but you can be certain that the secret formula for how to make Coca-Cola syrup is protected as a trade secret.

https://www.coca-colacompany.com/company/history/coca-cola-f...

Re: Google denies training Bard on ChatGPT chats from ShareGPT

#100
post #17

Hard to believe that is true, or else Bard would probably not perform so bad.

Bard is only a week old and has a large "experimental" sticker on it. Besides its UI is better and the answers are succinct which I prefer.

They literally copied the Chatgpt UI, lol, only it looks like a dated Google UI. How do you prefer answers with less data?... that's crazy.
Post reply on HN