Live data from Hacker News

Google denies training Bard on ChatGPT chats from ShareGPT

twitter.com

141–150 of 342 posts

Re: Google denies training Bard on ChatGPT chats from ShareGPT

#141

According to the article, the story goes this way: This engineer Jacob Devlin raised his concerns on training Bard with ShareGPT data. Then he directly joined OpenAI. He also claims that Google were about to do it, and then they stopped after his warnings. And presumably removed every trace of openai's responses. A couple of things: 1. So, Bard could have been trained on ShareGPT but it's not - according to the same…

Your comment doesn't make sense to me.

> Bard team was heavily relying on ShareGPT.

> He also claims that Google were about to do it, and then they stopped after his warnings.

So were they heavily relying or were they about to and then stopped? It's unclear from your comment. Could you link where you're getting this info from? The Information article is walled, unfortunately.

Re: Google denies training Bard on ChatGPT chats from ShareGPT

#142

Heh, imagine the day most of online content will be AI generated, good luck guaranteeing that AI X,Y,Z, ... etc. won't feed each other, possibly even circularly.

Circular reporting will be the only reporting!

https://en.wikipedia.org/wiki/Circular_reporting

Re: Google denies training Bard on ChatGPT chats from ShareGPT

#143
post #88

1. Google denies doing it, so at the very least the title should have an "allegedly". 2. Even if they did – so what? The output from ChatGPT is not copyrightable by OpenAI. In fact it is OpenAI that is training its models on copyrighted data, pictures, code from all over the internet.

> The output from ChatGPT is not copyrightable by OpenAI.

I think the argument here is over the OpenAI Terms of Service, not copyright.

Re: Google denies training Bard on ChatGPT chats from ShareGPT

#144

Earlier quoted context omitted.

They are a public company so they cannot lie so openly right? Usually you see categorial denies. Here the statement is in no way categorical at all. > But Google is firmly and clearly denying the data was used: “Bard is not trained on any data from ShareGPT or ChatGPT,” spokesperson Chris Pappas tells The Verge

Normally I would suspect this could be due to a misunderstanding from the ShareGPT author who could have misinterpreted a bunch of traffic from Googlebot as Google scraping it for Bard training data. But there is a Google engineer who says he resigned because of it.

And then went to work for OpenAI. I'm not saying he's lying but he is not an unbiased observer.

Re: Google denies training Bard on ChatGPT chats from ShareGPT

#145

So? First off, the whole argument behind these models has been from day one that training on copyrighted material is fair use. At most this would be a TOS violation. Second off, AI output is not subject to copyright, so it has even less protection than the original works it was trained on. Copyright maximalism for me, but not for thee. It's just so silly for someone working at OpenAI to complain about this.

> At most this would be a TOS violation

And would it be a ShareGPT TOS violation (assuming it had any)?

If OpenAI says "you can share these online but don't use them for AI training", people share them on another site, and then someone else comes along to scrape that site for AI training data, there's no relationship between OpenAI and the scraper for the TOS to apply to.

Normally I think you'd rely on copyright in that kind of case, but that doesn't apply to ChatGPT's output, so...

Re: Google denies training Bard on ChatGPT chats from ShareGPT

#147

Earlier quoted context omitted.

They literally copied the Chatgpt UI, lol, only it looks like a dated Google UI. How do you prefer answers with less data?... that's crazy.

I just don't want to be hit with a wall of text every single time, it gets the point across with minimal padding (high signal to noise ratio), ChatGPT feels like it gets paid by the word and they do actually charge by token if you use the API. As for the UI it's a take on the tried and true chat UI same as ChatGPT's, it spits the whole answer at once instead of feeding it to you one word at a time, it has an alternat…

You can combat that in the prompt, I use "just code, no words" which will also remove code comments from output. Bard doesn't respect the same request. You can be more succinct with chatgpt. Half the things I ask for in Bard give me this:

"I'm still learning coding skills, so at the moment I can't help with this. I'm trained to do things like help you write lists about different topics, compare things, or build travel itineraries. Do you want to try any of those now?"

Re: Google denies training Bard on ChatGPT chats from ShareGPT

#148
post #88

1. Google denies doing it, so at the very least the title should have an "allegedly". 2. Even if they did – so what? The output from ChatGPT is not copyrightable by OpenAI. In fact it is OpenAI that is training its models on copyrighted data, pictures, code from all over the internet.

OpenAI Terms of service forbid training competitor models via their ML outputs (LoRa alpaca laundering is probably not allowed for commercial use).

Google has no contract with OpenAI though. They used a third party site to scrape conversations. If the outputs themselves are not copyrighted, and they never agreed to the terms of service, it should be fine, right? Albeit unethical and embarrassing.

Re: Google denies training Bard on ChatGPT chats from ShareGPT

#149

"What's sauce for the goose is sauce for the gander" as the legal cliche goes. OpenAI cannot on the one hand claim that google did something wrong if they used their outputs as part of the bard training while simultaneously on the other hand claiming they themselves are free to use everyone on the internets content to train their model. Either they believe that training should respect copyright (in which case they co…

No one is alleging copyright violations. The claim is that they violated OpenAI's terms of service. We don't know whether Google ever even agreed to those terms of service in the first place.

Re: Google denies training Bard on ChatGPT chats from ShareGPT

#150
post #88

1. Google denies doing it, so at the very least the title should have an "allegedly". 2. Even if they did – so what? The output from ChatGPT is not copyrightable by OpenAI. In fact it is OpenAI that is training its models on copyrighted data, pictures, code from all over the internet.

OpenAI Terms of service forbid training competitor models via their ML outputs (LoRa alpaca laundering is probably not allowed for commercial use).

breaking terms of service is not punishable in any way. Facebook tried and lost in court
Post reply on HN