Live data from Hacker News

Google denies training Bard on ChatGPT chats from ShareGPT

twitter.com

101–110 of 342 posts

Re: Google denies training Bard on ChatGPT chats from ShareGPT

#101

Earlier quoted context omitted.

It's definitely a derived work as far as copyright is concerned: the output would simply not exist without the copyrighted training data. > It's finding patterns same as anyone studying the code base would do. No, it's quite unlike anyone studying data, because it's not a person with legal rights, such as fair use, but an automated algorithm. There is absolutely no legal debate that copyright applies only to human au…

> It's definitely a derived work as far as copyright is concerned - the output would simply not exist without the copyrighted training data. Can you point to a legal case that confirms this? Because it’s not at all clear that this is true from a legal standpoint. “X would not exist without Y” is not a sufficient test for derivative works - it’s far more nuanced.

United States copyright law in quite clear on the matter:

>A "derivative work" is a work based upon one or more preexisting works, such as a translation, musical arrangement, dramatization, fictionalization, motion picture version, sound recording, art reproduction, abridgment, condensation, or any other form in which a work may be recast, transformed, or adapted.

The emphasis part clearly applies: not only the AI model needs to be trained on massive amounts of copyrighted works *); but without these input works, it displays no intrinsic creative ability, it has no capacity to produce a single intelligible word or sketch. All creative features of its productions are a transformation of (and only of) the creative features of the inputs, the AI algorithm has no "intelligence" in the common meaning of the word and no ability to create original works.

*) by that, I mean a specific instance of the model with certain desirable features, for example the ability to imitate the style of J.K Rowling

Re: Google denies training Bard on ChatGPT chats from ShareGPT

#102
People complained that new AI is "stealing" from artists.

But stealing from other AI turns out to often be easier.

And this is where things get fun, because companies like OpenAI want to be able to train on all the data without any explicit permissions from the creators, but the moment people do the same to them they likely (we will see) be very much against it.

So it will be interesting if they will be able to both have and eat the cake (e.g. by using Microsoft lobby to push absurd law) or will they fall apart due to cannibalization making it non profitable to create better AI.

EDIT: This comment isn't specific to Google/Bert, so it doesn't matter weather Google actually did so or weather not.

Re: Google denies training Bard on ChatGPT chats from ShareGPT

#103

Earlier quoted context omitted.

This is an argument in bad faith but at this point I have zero trust in corporations and feel like you can generally count on them to do shitty things if they can benefit from it so I can be easily swayed by little proof at this point.

What's the argument? What's been done by anyone that's shitty? I don't even understand the point of this post. As far as I know, the current wave of text-based AIs is trained on all text accessible on the internet. Would it be a scandal to learn that ChatGPT is trained on wikipedia? Reddit? What is even the argument here, good faith or otherwise?

The argument is these companies are using our ideas created by us humans in this thing called the interenet for free and without attribution and it's problematic.

Re: Google denies training Bard on ChatGPT chats from ShareGPT

#104
"What's sauce for the goose is sauce for the gander" as the legal cliche goes. OpenAI cannot on the one hand claim that google did something wrong if they used their outputs as part of the bard training while simultaneously on the other hand claiming they themselves are free to use everyone on the internets content to train their model.

Either they believe that training should respect copyright (in which case they could not do what they do) or they believe that training is fair use (in which case they cannot possibly object to Google doing the same as them).

Re: Google denies training Bard on ChatGPT chats from ShareGPT

#106
post #97

Earlier quoted context omitted.

AI is not a human being or a student in school. It’s a software tool, stop comparing the two.

Being in school is also just a tool to knowing stuff, being able to read, and being around similar aged peers, etc. Whether the knowledge is directly in your brain or in a device you operate (directly or through an API) shouldn't really matter. If it's forbidden for a human to move a stone with manual labour, then it's also forbidden to move that stone with an excavator. This has nothing to do with the person being a…

> If it's forbidden for a human to move a stone with manual labour, then it's also forbidden to move that stone with an excavator.

Sure, but the reverse is false: I can walk on my own feet through Hyde Park, but I can't ride my excavator there.

Laws are made by humans for the benefit of humans, it's a political struggle. Now, large corporation try to exploit loopholes in the existing copyright framework in order to expropriate creators of their works. It's standard uberisation: disrupt existing economic models, insert yourself as a unavoidable middle man and pauperize the workforce the provides the actual service.

Re: Google denies training Bard on ChatGPT chats from ShareGPT

#107
Regardless of whether this happened or not, would training Bard on ChatGPT output be good or bad for Bard's product quality? I imagine there's a risk of AIs recursively reinforcing bad data in their models. This problem seems unavoidable as more web content becomes AI-generated content and spam.

Re: Google denies training Bard on ChatGPT chats from ShareGPT

#108

Earlier quoted context omitted.

Bard is only a week old and has a large "experimental" sticker on it. Besides its UI is better and the answers are succinct which I prefer.

They literally copied the Chatgpt UI, lol, only it looks like a dated Google UI. How do you prefer answers with less data?... that's crazy.

doing a visual diff will show you it's not a literal copy

Re: Google denies training Bard on ChatGPT chats from ShareGPT

#109

Earlier quoted context omitted.

It's a bit more nuanced than that, what I mean is that the slow speed at which humans learn it's a foundation block of our society, if suddenly some new race of humans emerged that could read an entire book in a couple of minutes and achieve lifelong superhuman retention and assimilation of all that knowledge then we would have the exact same type of concerns than what we have today about AI, including how easily the…

Startup technologists have been acting like speed of actions doesn't matter for decades. If a person can do it, why shouldn't a computer do it 1000x faster? What could go wrong? It's always been a poor argument at best and a bad faith one at worst.

Well said. The mindless automation away of everything has only one logical conclusion in which the creators of such automations are automated themselves, and even if the optimists are right and we never get there it doesn't matter, the chaos it can make just by getting closer at faster rates than society can adapt is unprecedented, specially given that the population count is at all times high and there are many other simultaneous threats that need our attention (e.g. climate change)

Re: Google denies training Bard on ChatGPT chats from ShareGPT

#110
I don't care at all about this from a copyright or data ownership perspective, but I am a little skeptical that it's a good idea to be this incestuous with training data in the long run. It's one thing to do fine tuning or knowledge distillation for specialized domains or shrinking models. But if you're trying to train your own foundation model, is relying on output from other foundation models going to make them learn to imitate their own errors?
Post reply on HN