Live data from Hacker News

Google denies training Bard on ChatGPT chats from ShareGPT

twitter.com

301–310 of 342 posts

Re: Google denies training Bard on ChatGPT chats from ShareGPT

#301
post #300

Earlier quoted context omitted.

forgive me if i have limited sympathy when a burglars house gets robbed

That's not how rights work, though. I'm sure you don't want your rights to be conditional on the sympathy of others.

good thing I'm just some guy expressing an opinion, and not a judge, then

Re: Google denies training Bard on ChatGPT chats from ShareGPT

#302

Regardless of whether this happened or not, would training Bard on ChatGPT output be good or bad for Bard's product quality? I imagine there's a risk of AIs recursively reinforcing bad data in their models. This problem seems unavoidable as more web content becomes AI-generated content and spam.

This is my biggest fear in the space (aside from potential job displacement and the political outcomes), but AI basically eating its own dogfood, and regurgitating its already bad information. It could go south pretty quickly, and it's like a contagion, it can't be easily just removed from the system.

Re: Google denies training Bard on ChatGPT chats from ShareGPT

#303
post #191

Earlier quoted context omitted.

Maybe, but we are fast approaching the point (or more likely have crossed it already) where distinguishing between human and AI generated data isn't really possible. If Google indexes a blog, how does it know whether it was written with AI assistance and therefore should not be used for training? Heck, how does OpenAI itself prevent such a feedback loop from its own output (or that of other LLMs)?

> If Google indexes a blog, how does it know whether it was written with AI assistance and therefore should not be used for training Yes, this is an existential problem for Google and training future LLMs. See also, https://www.theverge.com/23642073/best-printer-2023-brother-... and https://searchengineland.com/verge-best-printer-2023-394709

Or Google can just materialize the expected page into existence at search time.

... it's uncanny how it always finds what you thought you were looking for!

Re: Google denies training Bard on ChatGPT chats from ShareGPT

#304
post #74
post #72

Earlier quoted context omitted.

> Furthermore, the U.S. Copyright Office (USCO) concluded that AI-generated works on their own cannot be copyright, so these ChatGPT logs are free game. Doesn't this depend on where you or the AI live? The US ain't the world.

Microsoft and Google are both US-based companies.

Sure, though they could easily run the training and inference in non-US subsidiaries.

Re: Google denies training Bard on ChatGPT chats from ShareGPT

#305
post #274

Earlier quoted context omitted.

I'm only half joking.... I think we likely will end up with flags for human generated/curated content (and it will have to be that way round, as I can't imagine spammers bothering to put flags on AI-generated stuff), and we probably already should have an equivalent of robots.txt protocol that allows users to specify which parts of their website they would and wouldn't like used in the training of LLMs.

If content with a "human-generated" flag is rated more highly in some way -- e.g. search results -- then of course spammers will automatically add that flag to their AI-generated garbage. How do you propose to prevent them?

And if it's not: why bother.

Re: Google denies training Bard on ChatGPT chats from ShareGPT

#306

Earlier quoted context omitted.

Not to mention it's embarrassing. Google playing second banana to OpenAI.

Many of Google's products were second, or later, to market. Google does not care.

Isn't it generally very hard to be first to market? And even if you are it's more likely that someone coming in later will take your lunch.

Apple wasn't the first one to try to make a successful smartphone, but they had resources, know-how, and tried at a better time with fewer unknowns around.

Re: Google denies training Bard on ChatGPT chats from ShareGPT

#307
post #230
post #176

Earlier quoted context omitted.

> OpenAI and Google both scour the web for human-generated content OpenAI and Google both scour the web for content, period. That content could be human generated or AI generated or a mix of the two. Neither company is respecting copyright or terms of service of every individual bit of data collected. Neither company cares how much effort was put into creating the data, whether humans were paid to do it, or whatever…

And herein is the main problem of AI. Its creators consume knowledge from the commons, and give nothing free and unencumbered back. It's like the guy who never brings anything to the potluck, but after everyone finishes eating, he boxes up the leftovers, and starts selling them out of a food cart.

> after everyone finishes eating, he boxes up the leftovers, and starts selling them out of a food cart.

That particular example doesn't seem all that great, as it gives the impression of not letting food go to waste.

Though sure, the guy could just give the food away, but shrug.

Re: Google denies training Bard on ChatGPT chats from ShareGPT

#308

Earlier quoted context omitted.

> At most this would be a TOS violation And would it be a ShareGPT TOS violation (assuming it had any)? If OpenAI says "you can share these online but don't use them for AI training", people share them on another site, and then someone else comes along to scrape that site for AI training data, there's no relationship between OpenAI and the scraper for the TOS to apply to. Normally I think you'd rely on copyright in t…

Right. And what even is the penalty of that TOS violation and how enforceable is it? I don't have an OpenAI account. I have never agreed to any TOS. I don't see what legal claim they would have to stop me from training an LLM on ShareGPT.

If Google were specifically going to ChatGPT to get its output and train off of it, they could be sued for breach of contract - and OpenAI would likely have a pretty good argument:

- they specifically tried extracting and learning from our model when it says you can't in our TOS

- this makes it easier for them to compete with us via the data they obtain in their breach of contract

- more businesses and enterprises might pass up on renting a shared or dedicated instance from us if they can just get it from Google

Re: Google denies training Bard on ChatGPT chats from ShareGPT

#309

Earlier quoted context omitted.

>Even if they did – so what? Amplification of biases, propagation of errors, echolalia and over-optimization, lack of diverse data, overfitting

Not to mention it's embarrassing. Google playing second banana to OpenAI.

That's okay if they can leapfrog.

Google Maps wasn't the first map product. We all used mapquest way before. But Google Maps was technologically advanced. Ajax made maps usable for the first time.

Gmail wasn't the first webmail. Hotmail had millions of customers already. But Google gave people unlimited space to store old email, whereas email in the old days filled up your inboxes and needed to be deleted.

Question is if they can and will leapfrog.

Google Plus was a sign of desperation and utterly failed. The dozens and dozens of different Messengers (sorry I don't even know what the latest one they're pushing is, RCS?) all failed.

As an organization we will see in the coming months and years if Google can still overtake others when coming in from behind or not.

Post reply on HN