Live data from Hacker News

Google denies training Bard on ChatGPT chats from ShareGPT

twitter.com

211–220 of 342 posts

Re: Google denies training Bard on ChatGPT chats from ShareGPT

#212

Apart from the open questions of the quality of such once-removed-from-human-generated training data... I can't speak to the legality of the situation, but the morality of using, without their consent, data generated by someone's AI engine... ... that was, itself, trained on other people's data without their consent... ... should be, at the very least, equivalently evil to the original AI's training.

No, it shouldn't. Maybe you should be, at the very least, considered a questionable person. I do not in any way or form consider anything to be wrong with what they're doing, but I question the senses of someone thinking this is immoral or even evil.

Keep your subjective nonsense out of this.

Re: Google denies training Bard on ChatGPT chats from ShareGPT

#213
post #165

Earlier quoted context omitted.

Hardly unethical, considering OpenAI is doing exactly this.

You can quibble about the ethics of web scraping for ML in general but I think you're conflating issues. OpenAI and Google both scour the web for human-generated content. What Google cares about here is the learnings from OpenAI's proprietary RLHF dataset, for which they had to contract a large sum of human labelers. Finding a roundabout way to extract the value of a direct competitor's purpose-built, costly data fee…

I see no difference. Any web scraping is a means to deflect revenue-generating traffic to yourself, and away from other websites. Fewer people will go to Stack Overflow because of Codex and Copilot. The point that the content was paid for vs volunteered becomes moot once it's posted publicly online for free, on ShareGPT.

Re: Google denies training Bard on ChatGPT chats from ShareGPT

#214

So? First off, the whole argument behind these models has been from day one that training on copyrighted material is fair use. At most this would be a TOS violation. Second off, AI output is not subject to copyright, so it has even less protection than the original works it was trained on. Copyright maximalism for me, but not for thee. It's just so silly for someone working at OpenAI to complain about this.

> It's just so silly for someone working at OpenAI to complain about this.

Who from OpenAI is complaining?

Re: Google denies training Bard on ChatGPT chats from ShareGPT

#215

Apart from the open questions of the quality of such once-removed-from-human-generated training data... I can't speak to the legality of the situation, but the morality of using, without their consent, data generated by someone's AI engine... ... that was, itself, trained on other people's data without their consent... ... should be, at the very least, equivalently evil to the original AI's training.

No, it shouldn't. Maybe you should be, at the very least, considered a questionable person. I do not in any way or form consider anything to be wrong with what they're doing, but I question the senses of someone thinking this is immoral or even evil. Keep your subjective nonsense out of this.

So were it to be the case that we should consider building an AI by scraping people's publicly-available work without their consent to be immoral (as many whose art was scraped to build e.g. stable diffusion would argue it should be)...

Do you not agree (in that context) we should consider scraping the output of an AI generated via such an immoral process to create yet another AI also immoral? At the very least, I'd think we would consider it further laundering of other people's labor with just extra steps.

Re: Google denies training Bard on ChatGPT chats from ShareGPT

#216

According to the article, the story goes this way: This engineer Jacob Devlin raised his concerns on training Bard with ShareGPT data. Then he directly joined OpenAI. He also claims that Google were about to do it, and then they stopped after his warnings. And presumably removed every trace of openai's responses. A couple of things: 1. So, Bard could have been trained on ShareGPT but it's not - according to the same…

Your comment doesn't make sense to me. > Bard team was heavily relying on ShareGPT. > He also claims that Google were about to do it, and then they stopped after his warnings. So were they heavily relying or were they about to and then stopped? It's unclear from your comment. Could you link where you're getting this info from? The Information article is walled, unfortunately.

[1] gives a jist as well.

What I meant to say was that: Acc to The Information article the engineer raised concerns because it appeared to him (article wording) Bard team was using (and heavily reliant on) ShareGPT for Bard training. The engineer wasnt working on Bard and presumably someone told him or somehow he got the impression that Bard team was reliant on ShareGPT. At the time he was at Google.

Then, when he raised concerns to Sundar Pichai, Bard team stopped doing it and also scrapped any traces of ShareGPT data. So, the headline is false and Bard (again presumably) is not trained on any of ShareGPT data.

[1]: https://www.theverge.com/2023/3/29/23662621/google-bard-chat...

Re: Google denies training Bard on ChatGPT chats from ShareGPT

#217
post #88

1. Google denies doing it, so at the very least the title should have an "allegedly". 2. Even if they did – so what? The output from ChatGPT is not copyrightable by OpenAI. In fact it is OpenAI that is training its models on copyrighted data, pictures, code from all over the internet.

Seems to me like it makes Google look kind of pathetic. That's worse than any legal issue here. (Caveat: assuming I understand the situation correctly)

Re: Google denies training Bard on ChatGPT chats from ShareGPT

#219
post #207
post #202

Earlier quoted context omitted.

The question turns on whether you consider copilot part of the "GitHub service." GitHub would argue that it is, and they'd likely argue that charging for access to copilot is akin to charging for access to private repositories. Others would say that copilot is somehow separate from the services Github provides, so using their code for CoPilot wouldn't be covered by the ToS.

It is certainly a service that's being provided. If not by GitHub, then by whom? I'll repeat the definition of service: The “Service” refers to the applications, software, products, and services provided by GitHub, including any Beta Previews.

So do you believe if you hosted a closed source project on GitHub, and GitHub decided they want to integrate this into their service they would simply be allowed to take the code?

Fortunately HN commenters are not judges. And I would wager any bet that MS lawyers would not try to argue based on their ToS either, that would be a recipe for loosing any court case.

Re: Google denies training Bard on ChatGPT chats from ShareGPT

#220

Earlier quoted context omitted.

Your comment doesn't make sense to me. > Bard team was heavily relying on ShareGPT. > He also claims that Google were about to do it, and then they stopped after his warnings. So were they heavily relying or were they about to and then stopped? It's unclear from your comment. Could you link where you're getting this info from? The Information article is walled, unfortunately.

[1] gives a jist as well. What I meant to say was that: Acc to The Information article the engineer raised concerns because it appeared to him (article wording) Bard team was using (and heavily reliant on) ShareGPT for Bard training. The engineer wasnt working on Bard and presumably someone told him or somehow he got the impression that Bard team was reliant on ShareGPT. At the time he was at Google. Then, when he ra…

I think I might be confused by your usage of “about to do it” in your original comment to mean “actively doing it.”

You claim that the very engineer accusing Google of training Bard on ShareGPT acknowledges that the final product was not. As far as I can tell, Devlin did no such thing.

Not sure why you would presume they restarted their expensive training process.

It just doesn’t seem like a good faith characterization to me.

Post reply on HN