Live data from Hacker News

Google denies training Bard on ChatGPT chats from ShareGPT

twitter.com

81–90 of 342 posts

Re: Google denies training Bard on ChatGPT chats from ShareGPT

#81
post #62

Earlier quoted context omitted.

It's definitely a derived work as far as copyright is concerned: the output would simply not exist without the copyrighted training data. > It's finding patterns same as anyone studying the code base would do. No, it's quite unlike anyone studying data, because it's not a person with legal rights, such as fair use, but an automated algorithm. There is absolutely no legal debate that copyright applies only to human au…

Students in school also will not never learn to read without being exposed to text. Does this mean that teachers who write exercise sheets and school text book publishers now own the copyright of everything students do?

AI is not a human being or a student in school. It’s a software tool, stop comparing the two.

Re: Google denies training Bard on ChatGPT chats from ShareGPT

#82
post #6

Good luck to them. AI models are automated plagiarism, top to bottom. None of us gave OpenAI permission to derive their model from our writing, surely billions of dollars worth, but they took it anyway. Copyright hasn't caught up so all that stolen value rests securely with OpenAI. If we're not getting that back, I don't see why AI competitors should have any qualms about borrowing each others' work.

I'm not a copyright maximalist, and I kind of agree that training should be fair use. Maybe I'm right about that, maybe I'm wrong. BUT importantly, that has to go hand in hand with an acknowledgement that AI material is not copyrightable and that training on other model output is fine. What companies like OpenAI want is a system where everything they build is protected, and nothing that anyone else builds is protecte…

That shouldn't be hard. Are Google's results copyrightable?

Re: Google denies training Bard on ChatGPT chats from ShareGPT

#83
post #6

Good luck to them. AI models are automated plagiarism, top to bottom. None of us gave OpenAI permission to derive their model from our writing, surely billions of dollars worth, but they took it anyway. Copyright hasn't caught up so all that stolen value rests securely with OpenAI. If we're not getting that back, I don't see why AI competitors should have any qualms about borrowing each others' work.

I'm not a copyright maximalist, and I kind of agree that training should be fair use. Maybe I'm right about that, maybe I'm wrong. BUT importantly, that has to go hand in hand with an acknowledgement that AI material is not copyrightable and that training on other model output is fine. What companies like OpenAI want is a system where everything they build is protected, and nothing that anyone else builds is protecte…

Building things and maintaining it as a trade secret can be protected as a trade secret.

Trade secrets don't need to be copyrightable (e.g. list of customer numbers is a trade secret but not copyrightable).

https://copyrightalliance.org/faqs/difference-copyright-pate...

> Trade secret protection protects secrets from unauthorized disclosure and use by others. A trade secret is information that has an economic benefit due to its secret nature, has value to others who cannot legitimately obtain it, and is subject to reasonable efforts to maintain its secrecy. The protections afforded by trade secret law are very different from others forms of IP.

Re: Google denies training Bard on ChatGPT chats from ShareGPT

#84
post #18

Earlier quoted context omitted.

Isn't most of the internet available through common crawl? I don't know what percentage of training data is just that data set but i assume it's enough for anyone with enough compute and ingenuity to create a reasonable LLM

Missed the point - they are saying that, in the future, there will be no human generated content left on the Internet.

Which is a baseless hyperbole. We get it, blog spam is annoying. That doesn’t change the fact that humans generate a ton of data just interacting with one another online.

Re: Google denies training Bard on ChatGPT chats from ShareGPT

#85

Earlier quoted context omitted.

It's definitely a derived work as far as copyright is concerned: the output would simply not exist without the copyrighted training data. > It's finding patterns same as anyone studying the code base would do. No, it's quite unlike anyone studying data, because it's not a person with legal rights, such as fair use, but an automated algorithm. There is absolutely no legal debate that copyright applies only to human au…

The output of human copyrighted work wouldn't exist if it weren't for humans training on the output of other humans. Humans constantly use cliches in their writing and speech, and most of what they produce is a repackaged version of what someone else has written or said, yet no one's up in arms against this mass of unoriginality as long as it's human-generated. This is anti-AI bias, pure and simple.

Most definitely. Good luck telling the difference between traditional and AI-empowered art in the near future.

It's just a new tool for artists, and this anti-AI sentiment towards copyright is only going to hurt individual artists, while doing nothing for large corporations with enough money to play the game.

Re: Google denies training Bard on ChatGPT chats from ShareGPT

#86

According to the article, the story goes this way: This engineer Jacob Devlin raised his concerns on training Bard with ShareGPT data. Then he directly joined OpenAI. He also claims that Google were about to do it, and then they stopped after his warnings. And presumably removed every trace of openai's responses. A couple of things: 1. So, Bard could have been trained on ShareGPT but it's not - according to the same…

Take what action? Pretty sure that’s not illegal, especially since the training data is ai generated and therefore can’t be copyrighted.

I think the oomph behind the story is due to it being embarrassing, rather than illegal.

Re: Google denies training Bard on ChatGPT chats from ShareGPT

#87

Google accused Microsoft Bing of using them for page rankings a few years ago. Setup a sting to show that when you searched for something unique on Google using MS Explorer, shortly afterwards the same search result would start showing up on Bing. This was seen as deeply embarrassing for Microsoft at the time.

Embarrassing, maybe, but imitation is the sincerest form of flattery.

Re: Google denies training Bard on ChatGPT chats from ShareGPT

#88
1. Google denies doing it, so at the very least the title should have an "allegedly".

2. Even if they did – so what? The output from ChatGPT is not copyrightable by OpenAI. In fact it is OpenAI that is training its models on copyrighted data, pictures, code from all over the internet.

Re: Google denies training Bard on ChatGPT chats from ShareGPT

#89
post #83

Earlier quoted context omitted.

I'm not a copyright maximalist, and I kind of agree that training should be fair use. Maybe I'm right about that, maybe I'm wrong. BUT importantly, that has to go hand in hand with an acknowledgement that AI material is not copyrightable and that training on other model output is fine. What companies like OpenAI want is a system where everything they build is protected, and nothing that anyone else builds is protecte…

Building things and maintaining it as a trade secret can be protected as a trade secret. Trade secrets don't need to be copyrightable (e.g. list of customer numbers is a trade secret but not copyrightable). https://copyrightalliance.org/faqs/difference-copyright-pate... > Trade secret protection protects secrets from unauthorized disclosure and use by others. A trade secret is information that has an economic benefit…

I am not a lawyer, but I don’t believe a trade secret would prevent someone from reverse engineering your model’s knowledge from it’s output though, in the same way that it doesn’t prevent someone from reverse engineering your hot sauce from buying a bunch and experimenting with the ingredients until it tastes similar.

Re: Google denies training Bard on ChatGPT chats from ShareGPT

#90

According to the article, the story goes this way: This engineer Jacob Devlin raised his concerns on training Bard with ShareGPT data. Then he directly joined OpenAI. He also claims that Google were about to do it, and then they stopped after his warnings. And presumably removed every trace of openai's responses. A couple of things: 1. So, Bard could have been trained on ShareGPT but it's not - according to the same…

Take what action? Pretty sure that’s not illegal, especially since the training data is ai generated and therefore can’t be copyrighted.

OpenAI could have blocked Google's accounts, for example. Nothing really to do with legality.
Post reply on HN