Live data from Hacker News

Google denies training Bard on ChatGPT chats from ShareGPT

twitter.com

11–20 of 342 posts

Re: Google denies training Bard on ChatGPT chats from ShareGPT

#11
post #8
post #5

Earlier quoted context omitted.

Google has already denied this. https://www.theverge.com/2023/3/29/23662621/google-bard-chat... (For whatever that's worth.)

The engineer's testimony and the scandal might be enough for OpenAI to try to get an injunction against Google to block their AI development. If that happens, it's game over for Google in the AI race. Disclaimer IANAL and all that, this is not legal advice.

Maybe we should all get one against OpenAI considering they've basically used everyone's material in one way or another and profited from it?

Re: Google denies training Bard on ChatGPT chats from ShareGPT

#12

Did Google ever agree to these terms of service? Why should they care? From a legal point of view this doesn't matter and from a moral point of view it's hilarious.

If a Google employee working on this thing ever agreed to OpenAI's terms of service, they might be screwed.

From OpenAI's terms:

(c) Restrictions. You may not (i) use the Services in a way that infringes, misappropriates or violates any person’s rights; (ii) reverse assemble, reverse compile, decompile, translate or otherwise attempt to discover the source code or underlying components of models, algorithms, and systems of the Services (except to the extent such restrictions are contrary to applicable law); (iii) use output from the Services to develop models that compete with OpenAI;

(j) Equitable Remedies. You acknowledge that if you violate or breach these Terms, it may cause irreparable harm to OpenAI and its affiliates, and OpenAI shall have the right to seek injunctive relief against you in addition to any other legal remedies.

Those two very clearly establish that if you use the output of their service to develop your own models, then you are in breach of the terms and they can seek injunctive relief against you (stop you from working until the case is resolved).

Re: Google denies training Bard on ChatGPT chats from ShareGPT

#13
post #6

Good luck to them. AI models are automated plagiarism, top to bottom. None of us gave OpenAI permission to derive their model from our writing, surely billions of dollars worth, but they took it anyway. Copyright hasn't caught up so all that stolen value rests securely with OpenAI. If we're not getting that back, I don't see why AI competitors should have any qualms about borrowing each others' work.

Our writing, our code, our artwork... Furthermore, the U.S. Copyright Office (USCO) concluded that AI-generated works on their own cannot be copyright, so these ChatGPT logs are free game. It would be hypocritical to think that Google is wrong and OpenAI is not.

Re: Google denies training Bard on ChatGPT chats from ShareGPT

#14
Google accused Microsoft Bing of using them for page rankings a few years ago. Setup a sting to show that when you searched for something unique on Google using MS Explorer, shortly afterwards the same search result would start showing up on Bing.

This was seen as deeply embarrassing for Microsoft at the time.

Re: Google denies training Bard on ChatGPT chats from ShareGPT

#15
post #8
post #5

Earlier quoted context omitted.

Google has already denied this. https://www.theverge.com/2023/3/29/23662621/google-bard-chat... (For whatever that's worth.)

The engineer's testimony and the scandal might be enough for OpenAI to try to get an injunction against Google to block their AI development. If that happens, it's game over for Google in the AI race. Disclaimer IANAL and all that, this is not legal advice.

Injunction on which grounds? Even if OpenAI had copyright over ChatGPT output (which is not at all clear), Google isn't distributing those, they just trained a model on them. So from a copyright perspective there's nothing to complain about. Unless OpenAI would want to argue that you need rights to your training data, but something tells me that that's not in their best interest.

Re: Google denies training Bard on ChatGPT chats from ShareGPT

#16
post #6

Good luck to them. AI models are automated plagiarism, top to bottom. None of us gave OpenAI permission to derive their model from our writing, surely billions of dollars worth, but they took it anyway. Copyright hasn't caught up so all that stolen value rests securely with OpenAI. If we're not getting that back, I don't see why AI competitors should have any qualms about borrowing each others' work.

I don't get this sentiment.

For some cases sure, if it repurposes your code that ignores the license fine. But it's rarely wholesale copying. It's finding patterns same as anyone studying the code base would do.

As for the majority of content written on the internet through reddit or some social media, what's the harm in ingesting that? It's an incredibly useful tool that will add huge value to everyone. It's relatively open, cheap and highly available. It's worth to it's owners is only a fraction of the value it will add to society. It has the chance to have as big of an impact on progress as something like the microprocessor.

I agree it's free game for other llms to use gpt output as training data and that's positive. Although it signals desperation and panic that the largest "ai first" company with more data than any org in history is caught so flat footed and has to rely on it.

Do you really think it would be a better world in which a large LLM would never be able to be developed?

Re: Google denies training Bard on ChatGPT chats from ShareGPT

#18

Thankfully archive.org exists, otherwise it would not be possible to get good training data in a few years when the internet is flooded with AI content.

Isn't most of the internet available through common crawl? I don't know what percentage of training data is just that data set but i assume it's enough for anyone with enough compute and ingenuity to create a reasonable LLM
Post reply on HN