Live data from Hacker News

Google denies training Bard on ChatGPT chats from ShareGPT

twitter.com

221–230 of 342 posts

Re: Google denies training Bard on ChatGPT chats from ShareGPT

#221
post #160

Earlier quoted context omitted.

And what about the terms of service of my blog or code repository? Does OpenAI respect that?

> And what about the terms of service of my blog or code repository? Does OpenAI respect that? Seems to me that’s an issue between you and OpenAI. (Does your blog or code repository actually have published restrictive terms of service? Did it when OpenAI accessed it? Did OpenAI even access it?)

You think OpenAI is going to care unless you have a team of expensive lawyers to back you up?

Microsoft is out there laundering GPL code with Copilot. These companies live firmly in the don't give a fuck region of capitalism. Copyright law for thee, not for me.

Re: Google denies training Bard on ChatGPT chats from ShareGPT

#222

Apart from the open questions of the quality of such once-removed-from-human-generated training data... I can't speak to the legality of the situation, but the morality of using, without their consent, data generated by someone's AI engine... ... that was, itself, trained on other people's data without their consent... ... should be, at the very least, equivalently evil to the original AI's training.

No, it shouldn't. Maybe you should be, at the very least, considered a questionable person. I do not in any way or form consider anything to be wrong with what they're doing, but I question the senses of someone thinking this is immoral or even evil. Keep your subjective nonsense out of this.

[deleted]

Re: Google denies training Bard on ChatGPT chats from ShareGPT

#223
post #57

Earlier quoted context omitted.

its not even that on their own those works cant be copywritten. its that even when you make changes to those works, your changes might qualify for copyright but they do not affect the copyright status of the ai generated portions of the work. if you used ai to design a new superhero and then added pink shoes, yellow hair, and a beard, only those three elements would possibly be able to be protected by copywrite. your…

> if you used ai to design a new superhero and then added pink shoes, yellow hair, and a beard Wouldn't that depend heavily on the prompt used (among other factors such as image to image and ControlNet)? You could be specifying lots of detail about the design in your prompt, and the AI could only be generating concept artwork with little variation from what you already provided. If I'm already providing the pose, the…

According to the copyright board a promot is not anymore than any person commissioning a work from an artist, which does not provide copyright, and the lack of human authorship for the design decisions still stops it from being protected by copyright.

Re: Google denies training Bard on ChatGPT chats from ShareGPT

#224

Earlier quoted context omitted.

its not even that on their own those works cant be copywritten. its that even when you make changes to those works, your changes might qualify for copyright but they do not affect the copyright status of the ai generated portions of the work. if you used ai to design a new superhero and then added pink shoes, yellow hair, and a beard, only those three elements would possibly be able to be protected by copywrite. your…

How could that be ever really be enforceable? If I use an AI tool to design my Superhero, can't I just submit it without disclosing the help I received from an AI. I get that it would be very nice to prevent AI SPAM copyrighting of every possible superhero, but if I use the AI to come up with a concept, then quickly redraw it myself with pen and paper, I feel like it would never be provable that it came from an AI.

you would be committing fraud. what happens if a criminal commits fraud?

Re: Google denies training Bard on ChatGPT chats from ShareGPT

#225
post #88

1. Google denies doing it, so at the very least the title should have an "allegedly". 2. Even if they did – so what? The output from ChatGPT is not copyrightable by OpenAI. In fact it is OpenAI that is training its models on copyrighted data, pictures, code from all over the internet.

OpenAI Terms of service forbid training competitor models via their ML outputs (LoRa alpaca laundering is probably not allowed for commercial use).

Are the TOS even enforceable is AI content can't be copyrighted?

Re: Google denies training Bard on ChatGPT chats from ShareGPT

#226
post #150

Earlier quoted context omitted.

OpenAI Terms of service forbid training competitor models via their ML outputs (LoRa alpaca laundering is probably not allowed for commercial use).

breaking terms of service is not punishable in any way. Facebook tried and lost in court

Correction – breaking terms of service that you have not explicitly agreed to is not punishable in any way. A site cannot enforce a "by using this site you agree to..." clause deep inside some license page that visitors are generally unaware of. If you violate an agreement that you willingly chose to enter, however, you will likely be found liable for it.

Re: Google denies training Bard on ChatGPT chats from ShareGPT

#227

Earlier quoted context omitted.

LinkedIn won that case on appeal, HiQ waas found to be violating the ToS, common misconception I was pointed at a link explaining the case here on HN, after trying to make a similar point, but cannot find the link currently edit, not the one I was pointed at, but similar https://www.fbm.com/publications/what-recent-rulings-in-hiq-...

That's just because they made accounts and so agreed to the terms right? From your link: >These rulings suggest that courts are much more comfortable restricting scraping activity where the parties have agreed by contract (whether directly or through agents) not to scrape. But courts remain wary of applying the CFAA and the potential criminal consequences it carries to scraping. The apparent exception is when a compa…

No, the case did not decide anything, no precedent was set. The point is that you cannot use this case to argue that you can scrape public data free of consequence

Re: Google denies training Bard on ChatGPT chats from ShareGPT

#228
post #191

Earlier quoted context omitted.

Right, but training an LLM on the output of another LLM can certainly exacerbate these issues

Maybe, but we are fast approaching the point (or more likely have crossed it already) where distinguishing between human and AI generated data isn't really possible. If Google indexes a blog, how does it know whether it was written with AI assistance and therefore should not be used for training? Heck, how does OpenAI itself prevent such a feedback loop from its own output (or that of other LLMs)?

> If Google indexes a blog, how does it know whether it was written with AI assistance and therefore should not be used for training

Yes, this is an existential problem for Google and training future LLMs.

See also, https://www.theverge.com/23642073/best-printer-2023-brother-... and https://searchengineland.com/verge-best-printer-2023-394709

Re: Google denies training Bard on ChatGPT chats from ShareGPT

#229
post #191

Earlier quoted context omitted.

Right, but training an LLM on the output of another LLM can certainly exacerbate these issues

Maybe, but we are fast approaching the point (or more likely have crossed it already) where distinguishing between human and AI generated data isn't really possible. If Google indexes a blog, how does it know whether it was written with AI assistance and therefore should not be used for training? Heck, how does OpenAI itself prevent such a feedback loop from its own output (or that of other LLMs)?

> Heck, how does OpenAI itself prevent such a feedback loop from its own output (or that or other LLMs)?

Seems trivial. Only use old data for the bulk? Feed some new data carefully curated?

Re: Google denies training Bard on ChatGPT chats from ShareGPT

#230
post #176

Earlier quoted context omitted.

You can quibble about the ethics of web scraping for ML in general but I think you're conflating issues. OpenAI and Google both scour the web for human-generated content. What Google cares about here is the learnings from OpenAI's proprietary RLHF dataset, for which they had to contract a large sum of human labelers. Finding a roundabout way to extract the value of a direct competitor's purpose-built, costly data fee…

> OpenAI and Google both scour the web for human-generated content OpenAI and Google both scour the web for content, period. That content could be human generated or AI generated or a mix of the two. Neither company is respecting copyright or terms of service of every individual bit of data collected. Neither company cares how much effort was put into creating the data, whether humans were paid to do it, or whatever…

And herein is the main problem of AI. Its creators consume knowledge from the commons, and give nothing free and unencumbered back.

It's like the guy who never brings anything to the potluck, but after everyone finishes eating, he boxes up the leftovers, and starts selling them out of a food cart.

Post reply on HN