Live data from Hacker News

Google denies training Bard on ChatGPT chats from ShareGPT

twitter.com

61–70 of 342 posts

Re: Google denies training Bard on ChatGPT chats from ShareGPT

#61
post #16
post #6

Good luck to them. AI models are automated plagiarism, top to bottom. None of us gave OpenAI permission to derive their model from our writing, surely billions of dollars worth, but they took it anyway. Copyright hasn't caught up so all that stolen value rests securely with OpenAI. If we're not getting that back, I don't see why AI competitors should have any qualms about borrowing each others' work.

I don't get this sentiment. For some cases sure, if it repurposes your code that ignores the license fine. But it's rarely wholesale copying. It's finding patterns same as anyone studying the code base would do. As for the majority of content written on the internet through reddit or some social media, what's the harm in ingesting that? It's an incredibly useful tool that will add huge value to everyone. It's relativ…

> what's the harm in ingesting that?

It means that large tech companies benefit the most from every incremental piece of content created by humans, in perpetuity.

Re: Google denies training Bard on ChatGPT chats from ShareGPT

#62
post #16

Earlier quoted context omitted.

I don't get this sentiment. For some cases sure, if it repurposes your code that ignores the license fine. But it's rarely wholesale copying. It's finding patterns same as anyone studying the code base would do. As for the majority of content written on the internet through reddit or some social media, what's the harm in ingesting that? It's an incredibly useful tool that will add huge value to everyone. It's relativ…

It's definitely a derived work as far as copyright is concerned: the output would simply not exist without the copyrighted training data. > It's finding patterns same as anyone studying the code base would do. No, it's quite unlike anyone studying data, because it's not a person with legal rights, such as fair use, but an automated algorithm. There is absolutely no legal debate that copyright applies only to human au…

Students in school also will not never learn to read without being exposed to text. Does this mean that teachers who write exercise sheets and school text book publishers now own the copyright of everything students do?

Re: Google denies training Bard on ChatGPT chats from ShareGPT

#63
post #25

Earlier quoted context omitted.

Injunction on which grounds? Even if OpenAI had copyright over ChatGPT output (which is not at all clear), Google isn't distributing those, they just trained a model on them. So from a copyright perspective there's nothing to complain about. Unless OpenAI would want to argue that you need rights to your training data, but something tells me that that's not in their best interest.

Again, IANAL. But it could be extremely damaging to OpenAI for their biggest openly declared competition (Google), to have used OpenAI's tech to improve their own. So it could seem reasonable to a judge to grant temporary/preliminary injunction relief to OpenAI against Google until discovery can happen or an audience can be held.

A judge imposing any penalties or restrictions on Google over Google allegedly—and maximally—scraping data from a third-party site for use as part of Bard's training corpus would be outrageous.

Re: Google denies training Bard on ChatGPT chats from ShareGPT

#64
post #12

Did Google ever agree to these terms of service? Why should they care? From a legal point of view this doesn't matter and from a moral point of view it's hilarious.

If a Google employee working on this thing ever agreed to OpenAI's terms of service, they might be screwed. From OpenAI's terms: (c) Restrictions. You may not (i) use the Services in a way that infringes, misappropriates or violates any person’s rights; (ii) reverse assemble, reverse compile, decompile, translate or otherwise attempt to discover the source code or underlying components of models, algorithms, and syst…

Wouldn't that only apply if that employee was acting as an agent of Google at the time?

Otherwise it would create an interesting dynamic that startups where no-one has created an OpenAI account would have a massive advantage, since they can freely scrape ShareGPT data and train on it while larger companies have enough employees that someone must have signed every TOS.

Re: Google denies training Bard on ChatGPT chats from ShareGPT

#65
post #33

Earlier quoted context omitted.

Sure. If you can get everyone to create an account and agree to those terms before reading your comments, you might have a case. Otherwise, it will be considered public information, at which point it is free to be scraped by anyone (see the precedent set by the LinkedIn/hiQ case).

LinkedIn won that case on appeal, HiQ waas found to be violating the ToS, common misconception I was pointed at a link explaining the case here on HN, after trying to make a similar point, but cannot find the link currently edit, not the one I was pointed at, but similar https://www.fbm.com/publications/what-recent-rulings-in-hiq-...

That's just because they made accounts and so agreed to the terms right?

From your link:

>These rulings suggest that courts are much more comfortable restricting scraping activity where the parties have agreed by contract (whether directly or through agents) not to scrape. But courts remain wary of applying the CFAA and the potential criminal consequences it carries to scraping. The apparent exception is when a company engages in a pattern of intentionally creating fake accounts to collect logged-in data.

Re: Google denies training Bard on ChatGPT chats from ShareGPT

#66
post #58

Earlier quoted context omitted.

Google only has a fraction of the training data. OpenAI had a huge head start and has been collecting training data for years now. ChatGPT is also wildly popular which has given them tons more training data. It's estimated that ChatGPT gained over 100 million users in the first two months alone, and may have over 13 million active users daily. The logs on ShareGPT are merely a drop in the bucket.

> Google only has a fraction of the training data. Uh, what? The same Google that has been crawling, indexing, and letting people search the entire Internet for the last 25 years? They have owned DeepMind for nearly twice as long as OpenAI has been in existence! If anything this is proof that no one at Google can get anything done anymore, and lack of training data ain't the problem.

The alignment portion of training requires you to have upvote/downvote data on many LLM responses. Google’s attempt at that (at least according to the news so far) was asking all employees to volunteer time ranking the responses. Combined with no historical feedback from ChatGPT, they are behind.

Re: Google denies training Bard on ChatGPT chats from ShareGPT

#67

Earlier quoted context omitted.

I am strongly in favor of eliminating copyright completely everywhere, soooo I am pretty fine with that. The other direction should be more enforce-able: stuff derived from open data must also be made open again, like the GPL but for data (and therefore ML stuff).

Right but in a world where copyright does exist we arguably have the worst of both worlds. Small players are not protected at all from scraping and big players are leveraging all of their work and have the legal resources to form a moat.

The smallest player is the user, and they should have real ownership over their computers.

Re: Google denies training Bard on ChatGPT chats from ShareGPT

#68
post #28

How the turn tables. Remember when Google called out Microsoft in 2011 for using Google results? https://googleblog.blogspot.com/2011/02/microsofts-bing-uses... >We look forward to competing with genuinely new search algorithms out there—algorithms built on core innovation, and not on recycled search results from a competitor.

Google: We look forward to [babble babble empty words we don't really mean on principle and more corporate speak that we laugh about having written in the bar.]

Is there even a single free non-bargained soul behind these companies' executive functions?

Re: Google denies training Bard on ChatGPT chats from ShareGPT

#70

According to the article, the story goes this way: This engineer Jacob Devlin raised his concerns on training Bard with ShareGPT data. Then he directly joined OpenAI. He also claims that Google were about to do it, and then they stopped after his warnings. And presumably removed every trace of openai's responses. A couple of things: 1. So, Bard could have been trained on ShareGPT but it's not - according to the same…

For those that don't know, Jacob Devlin was the lead engineer and first publisher of the widely popular BERT model architecture, and initial bert-base models released by Google.

https://www.semanticscholar.org/author/Jacob-Devlin/39172707

Post reply on HN