Live data from Hacker News

Google denies training Bard on ChatGPT chats from ShareGPT

twitter.com

241–250 of 342 posts

Re: Google denies training Bard on ChatGPT chats from ShareGPT

#242
post #239

"What's sauce for the goose is sauce for the gander" as the legal cliche goes. OpenAI cannot on the one hand claim that google did something wrong if they used their outputs as part of the bard training while simultaneously on the other hand claiming they themselves are free to use everyone on the internets content to train their model. Either they believe that training should respect copyright (in which case they co…

[flagged]

That's nonsensical. An AI is either transformative or it's not, it's an intrinsic quality that has nothing to do with the training data or the "product" type. If OpenAI is sufficiently transformative to claim fair use (which I don't believe for a second, alas), then any other AI built on similar fundamentals has the same claims and can crunch any data their creators see fit, including the output of other AIs.

Re: Google denies training Bard on ChatGPT chats from ShareGPT

#243
post #165

Earlier quoted context omitted.

Hardly unethical, considering OpenAI is doing exactly this.

You can quibble about the ethics of web scraping for ML in general but I think you're conflating issues. OpenAI and Google both scour the web for human-generated content. What Google cares about here is the learnings from OpenAI's proprietary RLHF dataset, for which they had to contract a large sum of human labelers. Finding a roundabout way to extract the value of a direct competitor's purpose-built, costly data fee…

> labelers. Finding a roundabout way to extract the value of a direct competitor's purpose-built, costly data feels meaningfully different from scraping the web in general as an input to a transformative use

There we go again, its, one law for the unwashed plebs and the other for us.

Why do you think that I, after spending my time and effort to write my blog, own my content to a lesser extent that OpenAI does their? Such hypocracy.

Re: Google denies training Bard on ChatGPT chats from ShareGPT

#244

Earlier quoted context omitted.

Missed the point - they are saying that, in the future, there will be no human generated content left on the Internet.

Which is a baseless hyperbole. We get it, blog spam is annoying. That doesn’t change the fact that humans generate a ton of data just interacting with one another online.

As a forum moderator, I have transitioned to relying heavily on AI-generated responses to users.

These responses can range from short and concise ("Friendly reminder: please ensure that all content posted adheres to our rules regarding hate speech. Let's work together to maintain a safe and inclusive community for everyone") to lengthy explanations of underlying issues.

By using AI-generated content, a small moderation team can efficiently manage a large group of users in a timely manner.

This approach is becoming increasingly common, as evidenced by the rise in AI-generated comments on popular sites such as HN, Reddit, Twitter, and Facebook.

Many users are also using AI tools to fix grammar issues and add extra content to their comments, which can be tempting but may result in unintentional changes to the original message.

In fact, I myself have used this technique to edit this very comment to provide an example.

---- Original comment:

As an online forum mod, I switched to mainly using AI to generate replies to users. Some are very short ("Hey! Remember the rules.") and some are long paragraphs explaining underlying issues. Someone training on my replies would pretty much train on AI generated content without knowing. It allows a small moderation team to moderate a large group quickly. I know that I am not alone in this.

There is also a raise in AI generated comments on sites like HN, Reddit, Twitter and Facebook. It's tempting to copy-paste a comment in AI for it to fix grammar issues, which often results in extra content being added to text. In fact, I did it for this comment.

Re: Google denies training Bard on ChatGPT chats from ShareGPT

#245

Earlier quoted context omitted.

>Even if they did – so what? Amplification of biases, propagation of errors, echolalia and over-optimization, lack of diverse data, overfitting

Not to mention it's embarrassing. Google playing second banana to OpenAI.

I think Amazon was first in the (free) banana business

Re: Google denies training Bard on ChatGPT chats from ShareGPT

#246
post #179

Earlier quoted context omitted.

Yep. It's in their ToS: If you're posting anything you did not create yourself or do not own the rights to, you agree that you are responsible for any Content you post; that you will only submit Content that you have the right to post; and that you will fully comply with any third party licenses relating to Content you post. I suppose this means if I upload your stuff to GitHub, and you sue GitHub, then GitHub would…

That doesn't make sense. For example, GPLv3 allows anyone to redistribute the software's source code if the license is intact: > You may convey verbatim copies of the Program's source code as you receive it, in any medium, provided that you conspicuously and appropriately publish on each copy an appropriate copyright notice; keep intact all notices stating that this License and any non-permissive terms added in accor…

Uploading is granting GitHub a license separate from the gpl license.

If you can't actually grant that separate license, you're misrepresenting your ownership and license to that code

Re: Google denies training Bard on ChatGPT chats from ShareGPT

#247

Earlier quoted context omitted.

OpenAI Terms of service forbid training competitor models via their ML outputs (LoRa alpaca laundering is probably not allowed for commercial use).

Google has no contract with OpenAI though. They used a third party site to scrape conversations. If the outputs themselves are not copyrighted, and they never agreed to the terms of service, it should be fine, right? Albeit unethical and embarrassing.

[deleted]

Re: Google denies training Bard on ChatGPT chats from ShareGPT

#248
post #88

1. Google denies doing it, so at the very least the title should have an "allegedly". 2. Even if they did – so what? The output from ChatGPT is not copyrightable by OpenAI. In fact it is OpenAI that is training its models on copyrighted data, pictures, code from all over the internet.

> Google denies doing it Read their statement carefully and it's actually not a denial of the allegation. > But Google is firmly and clearly denying the data was used: “Bard is not trained on any data from ShareGPT or ChatGPT,” spokesperson Chris Pappas tells The Verge * Allegation: Google used ShareGPT to train Bard. * Rebuttal: The current production version of Bard is not trained on ShareGPT data Both things can b…

Intent matters I guess.

Did they accidentally train on that public piece of info they scraped anyway because they are scraping the whole web?

Or did they intentionally scrape chatgpt output to see if that would help?

Re: Google denies training Bard on ChatGPT chats from ShareGPT

#249

Earlier quoted context omitted.

>Even if they did – so what? Amplification of biases, propagation of errors, echolalia and over-optimization, lack of diverse data, overfitting

Not to mention it's embarrassing. Google playing second banana to OpenAI.

That assumes that training on the output of another language model somehow gives you the ability to improve your model and to catch up somehow

Re: Google denies training Bard on ChatGPT chats from ShareGPT

#250
post #88

1. Google denies doing it, so at the very least the title should have an "allegedly". 2. Even if they did – so what? The output from ChatGPT is not copyrightable by OpenAI. In fact it is OpenAI that is training its models on copyrighted data, pictures, code from all over the internet.

>Even if they did – so what? Amplification of biases, propagation of errors, echolalia and over-optimization, lack of diverse data, overfitting

you talk like chatgpt was some bastion of curated perfectly correct content. get a grip. web scraping is web scraping.
Post reply on HN