Live data from Hacker News

Google denies training Bard on ChatGPT chats from ShareGPT

twitter.com

131–140 of 342 posts

Re: Google denies training Bard on ChatGPT chats from ShareGPT

#131
post #125

I love that OpenAI uses a ton of other peoples work to train their model, yet when someone uses OpenAI to train their model, they get all up in arms. As far as I'm concerned, OpenAI has decided terms of use don't exist anymore.

OpenAI is training on data that is against their terms of use? That reads like a serious allegation. What is this all about?

Re: Google denies training Bard on ChatGPT chats from ShareGPT

#132

Earlier quoted context omitted.

The argument is these companies are using our ideas created by us humans in this thing called the interenet for free and without attribution and it's problematic.

You can't own ideas, they got their own life-cycle.

Right, but I do think you can "own" (by which I mean our societally-mediated legal definition of ownership in the anglosphere) specific sequences of text or at least the right to copy them?

Re: Google denies training Bard on ChatGPT chats from ShareGPT

#133
post #88

1. Google denies doing it, so at the very least the title should have an "allegedly". 2. Even if they did – so what? The output from ChatGPT is not copyrightable by OpenAI. In fact it is OpenAI that is training its models on copyrighted data, pictures, code from all over the internet.

But remember many years back when it was news that Bing used Google search results to improve its results.

It's not quite the same thing, because Bing was getting the data from a browser toolbar and watching the search terms used and where the user went afterwards.

A closer equivalent would be if someone had made a ShareSERP site and people posted their favorite search terms and the results Google gave and Bing crawled that and incorporated the search terms to links connections into their search graph.

The actual actions had maybe gone too far (personally I thought it was more funny than "copying"), the hypothetical would be pretty much what you'd expect to happen. Even google would probably crawl ShareSERP and inadvertently reinforce their own results (the same way OpenAI presumably gets more than a bit of their own results back at them in any new crawls of reddit, hn, etc even if they avoid sites like ShareGPT deliberately).

Re: Google denies training Bard on ChatGPT chats from ShareGPT

#134
post #88

1. Google denies doing it, so at the very least the title should have an "allegedly". 2. Even if they did – so what? The output from ChatGPT is not copyrightable by OpenAI. In fact it is OpenAI that is training its models on copyrighted data, pictures, code from all over the internet.

OpenAI Terms of service forbid training competitor models via their ML outputs (LoRa alpaca laundering is probably not allowed for commercial use).

Re: Google denies training Bard on ChatGPT chats from ShareGPT

#135

People complained that new AI is "stealing" from artists. But stealing from other AI turns out to often be easier. And this is where things get fun, because companies like OpenAI want to be able to train on all the data without any explicit permissions from the creators, but the moment people do the same to them they likely (we will see) be very much against it. So it will be interesting if they will be able to both…

I can see the GitHub Copilot controversy being resolved in this way. If Microsoft, GitHub, and OpenAI successfully use the fair use defense for Copilot's appropriation of proprietary and incompatibly licensed code, then a free and open source alternative to Copilot can be trained on Copilot's outputs.

After all, the GitHub Copilot Product Specific Terms say:

> 2. Ownership of Suggestions and Your Code

> GitHub does not claim any ownership rights in Suggestions. You retain ownership of Your Code.

https://github.com/customer-terms/github-copilot-product-spe...

Re: Google denies training Bard on ChatGPT chats from ShareGPT

#136
post #25

Earlier quoted context omitted.

Injunction on which grounds? Even if OpenAI had copyright over ChatGPT output (which is not at all clear), Google isn't distributing those, they just trained a model on them. So from a copyright perspective there's nothing to complain about. Unless OpenAI would want to argue that you need rights to your training data, but something tells me that that's not in their best interest.

Again, IANAL. But it could be extremely damaging to OpenAI for their biggest openly declared competition (Google), to have used OpenAI's tech to improve their own. So it could seem reasonable to a judge to grant temporary/preliminary injunction relief to OpenAI against Google until discovery can happen or an audience can be held.

Google could respond by seeding Bard output across the public internet, then if they can prove that GPT-5 is trained on this output, then they can sue back and AI development can stop altogether. Win for everybody!

Re: Google denies training Bard on ChatGPT chats from ShareGPT

#137
This is also bad because the risk of AI "inbreeding" is real. I have seen invisible artifact amplification happen in a single generation training ESRGAN on itself.

Maybe it wont happen in a single LLM generation, but perhaps gen 3 or 5 will start having really weird speech patterns or hallucinations because of this.

Re: Google denies training Bard on ChatGPT chats from ShareGPT

#138
post #88

1. Google denies doing it, so at the very least the title should have an "allegedly". 2. Even if they did – so what? The output from ChatGPT is not copyrightable by OpenAI. In fact it is OpenAI that is training its models on copyrighted data, pictures, code from all over the internet.

OpenAI Terms of service forbid training competitor models via their ML outputs (LoRa alpaca laundering is probably not allowed for commercial use).

Where exactly does it do that? I looked a bit and could t find it, but likely I was just wrong

Re: Google denies training Bard on ChatGPT chats from ShareGPT

#139
post #88

1. Google denies doing it, so at the very least the title should have an "allegedly". 2. Even if they did – so what? The output from ChatGPT is not copyrightable by OpenAI. In fact it is OpenAI that is training its models on copyrighted data, pictures, code from all over the internet.

OpenAI Terms of service forbid training competitor models via their ML outputs (LoRa alpaca laundering is probably not allowed for commercial use).

I love it how they don't want others to use their model output but they have no qualms about training their model on the copyrighted works of others? Isn't this a stunning level of hypocrisy?

Re: Google denies training Bard on ChatGPT chats from ShareGPT

#140
post #88

1. Google denies doing it, so at the very least the title should have an "allegedly". 2. Even if they did – so what? The output from ChatGPT is not copyrightable by OpenAI. In fact it is OpenAI that is training its models on copyrighted data, pictures, code from all over the internet.

OpenAI Terms of service forbid training competitor models via their ML outputs (LoRa alpaca laundering is probably not allowed for commercial use).

So, to verify, are you claiming that if someone added a similar clause to their source code and then GitHub went ahead and trained Copilot against it, that would be an issue?
Post reply on HN