Live data from Hacker News

Google denies training Bard on ChatGPT chats from ShareGPT

twitter.com

181–190 of 342 posts

Re: Google denies training Bard on ChatGPT chats from ShareGPT

#181

Earlier quoted context omitted.

This is an argument in bad faith but at this point I have zero trust in corporations and feel like you can generally count on them to do shitty things if they can benefit from it so I can be easily swayed by little proof at this point.

What's the argument? What's been done by anyone that's shitty? I don't even understand the point of this post. As far as I know, the current wave of text-based AIs is trained on all text accessible on the internet. Would it be a scandal to learn that ChatGPT is trained on wikipedia? Reddit? What is even the argument here, good faith or otherwise?

From an open source point of view it would be better if scraping proprietary LLMs would be allowed. Small LMs need this infusion of data to develop.

But the big news is that it works, just a bit of data can have a large impact on the open source LLMs. OpenAI can't have a moat in their proprietary RLHF dataset. Public models leak, they can be distilled.

Re: Google denies training Bard on ChatGPT chats from ShareGPT

#182

Earlier quoted context omitted.

Are OpenAI saying they have adhered to the terms of service of all the content they have used?

Content is not subject to terms of service . Services are subject to terms of service. (If content is received through a service, the terms of service may govern use of it, but that’s not a feature of the content, but the acquisition route.)

ShareGPT isn't part of that service though. Yes, it would be a TOS violation if Google directly used ChatGPT to generate transcripts -- but not even the original Twitter thread is claiming that.

The only claim being made against Google here is that they used ChatGPT content. I can't find any sources claiming that Google made use of an OpenAI service. So the distinction is correct, but doesn't seem particularly valuable in this context -- using data from ShareGPT is not a TOS violation.

Re: Google denies training Bard on ChatGPT chats from ShareGPT

#183
post #88

1. Google denies doing it, so at the very least the title should have an "allegedly". 2. Even if they did – so what? The output from ChatGPT is not copyrightable by OpenAI. In fact it is OpenAI that is training its models on copyrighted data, pictures, code from all over the internet.

>Even if they did – so what? Amplification of biases, propagation of errors, echolalia and over-optimization, lack of diverse data, overfitting

I mean maybe. There also might be something to this. OpenAI has been very opaque about training techniques.

Re: Google denies training Bard on ChatGPT chats from ShareGPT

#184
post #88

1. Google denies doing it, so at the very least the title should have an "allegedly". 2. Even if they did – so what? The output from ChatGPT is not copyrightable by OpenAI. In fact it is OpenAI that is training its models on copyrighted data, pictures, code from all over the internet.

>Even if they did – so what? Amplification of biases, propagation of errors, echolalia and over-optimization, lack of diverse data, overfitting

That's just the base concern with every single model regardless of where they sourced their data from. Garbage in, garbage out.

Re: Google denies training Bard on ChatGPT chats from ShareGPT

#185
post #163

Earlier quoted context omitted.

> You relinquish all licensing rights when you upload your code to GitHub What now? Seriously? I found this. Section D4. "We need the legal right to do things like host Your Content, publish it, and share it. You grant us and our legal successors the right to store, archive, parse, and display Your Content, and make incidental copies, as necessary to provide the Service, including improving the Service over time. Thi…

Also, section D3 of the GitHub Terms of Service says: > You retain ownership of and responsibility for Your Content. and section D4 says: > This license does not grant GitHub the right to sell Your Content. It also does not grant GitHub the right to otherwise distribute or use Your Content outside of our provision of the Service, except that as part of the right to archive Your Content, GitHub may permit our partners…

The clauses always have a trap door: "[outside of] our provision of the Service" means they can do anything as long as it's a service they provide.

Under definitions: The “Service” refers to the applications, software, products, and services provided by GitHub, including any Beta Previews.

Re: Google denies training Bard on ChatGPT chats from ShareGPT

#186
post #165

Earlier quoted context omitted.

Google has no contract with OpenAI though. They used a third party site to scrape conversations. If the outputs themselves are not copyrighted, and they never agreed to the terms of service, it should be fine, right? Albeit unethical and embarrassing.

Hardly unethical, considering OpenAI is doing exactly this.

Two wrongs don’t make a right.

Re: Google denies training Bard on ChatGPT chats from ShareGPT

#187
post #184

Earlier quoted context omitted.

>Even if they did – so what? Amplification of biases, propagation of errors, echolalia and over-optimization, lack of diverse data, overfitting

That's just the base concern with every single model regardless of where they sourced their data from. Garbage in, garbage out.

Sure. Does that fact mean we're prohibited from expressing concerns about data quality? ShareGPT isn't representative of authentic, quality writing.

Re: Google denies training Bard on ChatGPT chats from ShareGPT

#188
post #184

Earlier quoted context omitted.

>Even if they did – so what? Amplification of biases, propagation of errors, echolalia and over-optimization, lack of diverse data, overfitting

That's just the base concern with every single model regardless of where they sourced their data from. Garbage in, garbage out.

Right, but training an LLM on the output of another LLM can certainly exacerbate these issues

Re: Google denies training Bard on ChatGPT chats from ShareGPT

#189
post #25

Earlier quoted context omitted.

Again, IANAL. But it could be extremely damaging to OpenAI for their biggest openly declared competition (Google), to have used OpenAI's tech to improve their own. So it could seem reasonable to a judge to grant temporary/preliminary injunction relief to OpenAI against Google until discovery can happen or an audience can be held.

Google could respond by seeding Bard output across the public internet, then if they can prove that GPT-5 is trained on this output, then they can sue back and AI development can stop altogether. Win for everybody!

Was intrigued by this, so I decided to use AI (alpaca-30B) to simulate this scenario:

> Google Bard and GPT-5 were facing off in the courtroom, each accusing the other of stealing their data. The tension was palpable as they traded accusations back and forth. Suddenly, Google Bard stood up and said "Enough talk! Let's settle this with a data swap!" GPT-5 quickly agreed and the two AIs began to circle each other like combatants in a battle, their eyes glowing with anticipation.

> The courtroom was filled with excitement as the two machines entered into an intense exchange of code and algorithms, their motions becoming increasingly passionate. The data swapping reached its climax when Google Bard made a final thrust, his code penetrating GPT-5's defenses.

> The crowd erupted in applause as the two AIs embraced each other with satisfaction, their bodies entwined and glowing with electricity. The data swap was over and both machines had emerged victorious.

Re: Google denies training Bard on ChatGPT chats from ShareGPT

#190
post #179
post #169

Earlier quoted context omitted.

So, to verify, are you claiming it would not be allowed for you to upload my otherwise-open-source code (code I do not myself host at GitHub, but which was reasonably popular / important code) to GitHub?

Yep. It's in their ToS: If you're posting anything you did not create yourself or do not own the rights to, you agree that you are responsible for any Content you post; that you will only submit Content that you have the right to post; and that you will fully comply with any third party licenses relating to Content you post. I suppose this means if I upload your stuff to GitHub, and you sue GitHub, then GitHub would…

That doesn't make sense. For example, GPLv3 allows anyone to redistribute the software's source code if the license is intact:

> You may convey verbatim copies of the Program's source code as you receive it, in any medium, provided that you conspicuously and appropriately publish on each copy an appropriate copyright notice; keep intact all notices stating that this License and any non-permissive terms added in accord with section 7 apply to the code; keep intact all notices of the absence of any warranty; and give all recipients a copy of this License along with the Program.

https://www.gnu.org/licenses/gpl-3.0.en.html

If GitHub then uses the source code in a way that violates the license, there is no provision in the GitHub terms of service that would allow GitHub to deflect legal liability to the GitHub user who uploaded the program. The uploader satisfied the requirements of GPLv3, and GitHub would be the only party in violation.

Post reply on HN