Live data from Hacker News

Google denies training Bard on ChatGPT chats from ShareGPT

twitter.com

21–30 of 342 posts

Re: Google denies training Bard on ChatGPT chats from ShareGPT

#21
post #6

Good luck to them. AI models are automated plagiarism, top to bottom. None of us gave OpenAI permission to derive their model from our writing, surely billions of dollars worth, but they took it anyway. Copyright hasn't caught up so all that stolen value rests securely with OpenAI. If we're not getting that back, I don't see why AI competitors should have any qualms about borrowing each others' work.

Our writing, our code, our artwork... Furthermore, the U.S. Copyright Office (USCO) concluded that AI-generated works on their own cannot be copyright, so these ChatGPT logs are free game. It would be hypocritical to think that Google is wrong and OpenAI is not.

But clearly everything generated by an AI isn’t automatically in the public domain. That would be a trivial way of copyright laundering.

"Sorry, while this looks like a bit for bit copy of a popular Hollywood movie, it was actually entirely dreamt up by our new, sophisticated, definitely AI-using identity function."

Re: Google denies training Bard on ChatGPT chats from ShareGPT

#22
post #12

Did Google ever agree to these terms of service? Why should they care? From a legal point of view this doesn't matter and from a moral point of view it's hilarious.

If a Google employee working on this thing ever agreed to OpenAI's terms of service, they might be screwed. From OpenAI's terms: (c) Restrictions. You may not (i) use the Services in a way that infringes, misappropriates or violates any person’s rights; (ii) reverse assemble, reverse compile, decompile, translate or otherwise attempt to discover the source code or underlying components of models, algorithms, and syst…

I hereby set a terms of service for everything I post on the internet from now on. OpenAI may not train future GPT models on my words or my code without my express written permission.

Somehow, I don’t think they’ll care.

Re: Google denies training Bard on ChatGPT chats from ShareGPT

#23
post #16
post #6

Good luck to them. AI models are automated plagiarism, top to bottom. None of us gave OpenAI permission to derive their model from our writing, surely billions of dollars worth, but they took it anyway. Copyright hasn't caught up so all that stolen value rests securely with OpenAI. If we're not getting that back, I don't see why AI competitors should have any qualms about borrowing each others' work.

I don't get this sentiment. For some cases sure, if it repurposes your code that ignores the license fine. But it's rarely wholesale copying. It's finding patterns same as anyone studying the code base would do. As for the majority of content written on the internet through reddit or some social media, what's the harm in ingesting that? It's an incredibly useful tool that will add huge value to everyone. It's relativ…

No, but I believe a large language model is a work that is 99.9% derivative of its inputs, with all that implies for authorship and copyright. Right now it's just a heist.

Re: Google denies training Bard on ChatGPT chats from ShareGPT

#24
post #16
post #6

Good luck to them. AI models are automated plagiarism, top to bottom. None of us gave OpenAI permission to derive their model from our writing, surely billions of dollars worth, but they took it anyway. Copyright hasn't caught up so all that stolen value rests securely with OpenAI. If we're not getting that back, I don't see why AI competitors should have any qualms about borrowing each others' work.

I don't get this sentiment. For some cases sure, if it repurposes your code that ignores the license fine. But it's rarely wholesale copying. It's finding patterns same as anyone studying the code base would do. As for the majority of content written on the internet through reddit or some social media, what's the harm in ingesting that? It's an incredibly useful tool that will add huge value to everyone. It's relativ…

It's definitely a derived work as far as copyright is concerned: the output would simply not exist without the copyrighted training data.

> It's finding patterns same as anyone studying the code base would do.

No, it's quite unlike anyone studying data, because it's not a person with legal rights, such as fair use, but an automated algorithm. There is absolutely no legal debate that copyright applies only to human authors, or only to the human created part of a mixed work, there is vast jurisprudence on this; by extension, any fair use rights too, exist only for human users of the works. Derivation by automated means - for the express economic purpose of out-competing the creator in the market place, no less - is completely outside the spirit of copyright.

Re: Google denies training Bard on ChatGPT chats from ShareGPT

#25
post #8

Earlier quoted context omitted.

The engineer's testimony and the scandal might be enough for OpenAI to try to get an injunction against Google to block their AI development. If that happens, it's game over for Google in the AI race. Disclaimer IANAL and all that, this is not legal advice.

Injunction on which grounds? Even if OpenAI had copyright over ChatGPT output (which is not at all clear), Google isn't distributing those, they just trained a model on them. So from a copyright perspective there's nothing to complain about. Unless OpenAI would want to argue that you need rights to your training data, but something tells me that that's not in their best interest.

Again, IANAL. But it could be extremely damaging to OpenAI for their biggest openly declared competition (Google), to have used OpenAI's tech to improve their own.

So it could seem reasonable to a judge to grant temporary/preliminary injunction relief to OpenAI against Google until discovery can happen or an audience can be held.

Re: Google denies training Bard on ChatGPT chats from ShareGPT

#26
post #5

Paywalled upstream source: https://www.theinformation.com/articles/alphabets-google-and...

Google has already denied this. https://www.theverge.com/2023/3/29/23662621/google-bard-chat... (For whatever that's worth.)

They are a public company so they cannot lie so openly right? Usually you see categorial denies. Here the statement is in no way categorical at all.

> But Google is firmly and clearly denying the data was used: “Bard is not trained on any data from ShareGPT or ChatGPT,” spokesperson Chris Pappas tells The Verge

Re: Google denies training Bard on ChatGPT chats from ShareGPT

#27
post #6

Good luck to them. AI models are automated plagiarism, top to bottom. None of us gave OpenAI permission to derive their model from our writing, surely billions of dollars worth, but they took it anyway. Copyright hasn't caught up so all that stolen value rests securely with OpenAI. If we're not getting that back, I don't see why AI competitors should have any qualms about borrowing each others' work.

Our writing, our code, our artwork... Furthermore, the U.S. Copyright Office (USCO) concluded that AI-generated works on their own cannot be copyright, so these ChatGPT logs are free game. It would be hypocritical to think that Google is wrong and OpenAI is not.

its not even that on their own those works cant be copywritten. its that even when you make changes to those works, your changes might qualify for copyright but they do not affect the copyright status of the ai generated portions of the work.

if you used ai to design a new superhero and then added pink shoes, yellow hair, and a beard, only those three elements would possibly be able to be protected by copywrite. your additions do not change the status of the underlying ai work which cannot be protected and is available for anyone to use.

Re: Google denies training Bard on ChatGPT chats from ShareGPT

#28
How the turn tables. Remember when Google called out Microsoft in 2011 for using Google results?

https://googleblog.blogspot.com/2011/02/microsofts-bing-uses...

>We look forward to competing with genuinely new search algorithms out there—algorithms built on core innovation, and not on recycled search results from a competitor.

Re: Google denies training Bard on ChatGPT chats from ShareGPT

#29
post #6

Good luck to them. AI models are automated plagiarism, top to bottom. None of us gave OpenAI permission to derive their model from our writing, surely billions of dollars worth, but they took it anyway. Copyright hasn't caught up so all that stolen value rests securely with OpenAI. If we're not getting that back, I don't see why AI competitors should have any qualms about borrowing each others' work.

I am strongly in favor of eliminating copyright completely everywhere, soooo I am pretty fine with that. The other direction should be more enforce-able: stuff derived from open data must also be made open again, like the GPL but for data (and therefore ML stuff).

Re: Google denies training Bard on ChatGPT chats from ShareGPT

#30
post #21

Earlier quoted context omitted.

Our writing, our code, our artwork... Furthermore, the U.S. Copyright Office (USCO) concluded that AI-generated works on their own cannot be copyright, so these ChatGPT logs are free game. It would be hypocritical to think that Google is wrong and OpenAI is not.

But clearly everything generated by an AI isn’t automatically in the public domain. That would be a trivial way of copyright laundering. "Sorry, while this looks like a bit for bit copy of a popular Hollywood movie, it was actually entirely dreamt up by our new, sophisticated, definitely AI-using identity function."

The person using something similar to something else may be infringing but the ai work cannot be protected by copyright as it lacks human authorship. Those are two separate issues.
Post reply on HN