Live data from Hacker News

Google denies training Bard on ChatGPT chats from ShareGPT

twitter.com

41–50 of 342 posts

Re: Google denies training Bard on ChatGPT chats from ShareGPT

#41
post #16

Earlier quoted context omitted.

I don't get this sentiment. For some cases sure, if it repurposes your code that ignores the license fine. But it's rarely wholesale copying. It's finding patterns same as anyone studying the code base would do. As for the majority of content written on the internet through reddit or some social media, what's the harm in ingesting that? It's an incredibly useful tool that will add huge value to everyone. It's relativ…

It's definitely a derived work as far as copyright is concerned: the output would simply not exist without the copyrighted training data. > It's finding patterns same as anyone studying the code base would do. No, it's quite unlike anyone studying data, because it's not a person with legal rights, such as fair use, but an automated algorithm. There is absolutely no legal debate that copyright applies only to human au…

The output of human copyrighted work wouldn't exist if it weren't for humans training on the output of other humans.

Humans constantly use cliches in their writing and speech, and most of what they produce is a repackaged version of what someone else has written or said, yet no one's up in arms against this mass of unoriginality as long as it's human-generated.

This is anti-AI bias, pure and simple.

Re: Google denies training Bard on ChatGPT chats from ShareGPT

#42
post #21

Earlier quoted context omitted.

Our writing, our code, our artwork... Furthermore, the U.S. Copyright Office (USCO) concluded that AI-generated works on their own cannot be copyright, so these ChatGPT logs are free game. It would be hypocritical to think that Google is wrong and OpenAI is not.

But clearly everything generated by an AI isn’t automatically in the public domain. That would be a trivial way of copyright laundering. "Sorry, while this looks like a bit for bit copy of a popular Hollywood movie, it was actually entirely dreamt up by our new, sophisticated, definitely AI-using identity function."

No, but the original copyright holder would have to press charges against Bard. OpenAI wouldn't be able to take action there.

Re: Google denies training Bard on ChatGPT chats from ShareGPT

#44
post #16

Earlier quoted context omitted.

I don't get this sentiment. For some cases sure, if it repurposes your code that ignores the license fine. But it's rarely wholesale copying. It's finding patterns same as anyone studying the code base would do. As for the majority of content written on the internet through reddit or some social media, what's the harm in ingesting that? It's an incredibly useful tool that will add huge value to everyone. It's relativ…

It's definitely a derived work as far as copyright is concerned: the output would simply not exist without the copyrighted training data. > It's finding patterns same as anyone studying the code base would do. No, it's quite unlike anyone studying data, because it's not a person with legal rights, such as fair use, but an automated algorithm. There is absolutely no legal debate that copyright applies only to human au…

> It's definitely a derived work as far as copyright is concerned

...in your head. In the US (and most countries) there is no such legal case so far.

Re: Google denies training Bard on ChatGPT chats from ShareGPT

#45
post #6

Good luck to them. AI models are automated plagiarism, top to bottom. None of us gave OpenAI permission to derive their model from our writing, surely billions of dollars worth, but they took it anyway. Copyright hasn't caught up so all that stolen value rests securely with OpenAI. If we're not getting that back, I don't see why AI competitors should have any qualms about borrowing each others' work.

That's the answer to the YC Interview question "What is your unfair competitive advantage" in a nutshell. Morally it might be wrong. From a business building perspective it's access that no one has.

Re: Google denies training Bard on ChatGPT chats from ShareGPT

#46
post #16
post #6

Good luck to them. AI models are automated plagiarism, top to bottom. None of us gave OpenAI permission to derive their model from our writing, surely billions of dollars worth, but they took it anyway. Copyright hasn't caught up so all that stolen value rests securely with OpenAI. If we're not getting that back, I don't see why AI competitors should have any qualms about borrowing each others' work.

I don't get this sentiment. For some cases sure, if it repurposes your code that ignores the license fine. But it's rarely wholesale copying. It's finding patterns same as anyone studying the code base would do. As for the majority of content written on the internet through reddit or some social media, what's the harm in ingesting that? It's an incredibly useful tool that will add huge value to everyone. It's relativ…

> Do you really think it would be a better world in which a large LLM would never be able to be developed?

Maybe. I believe the potential for abuse is far greater than the potential benefits. What is our benefit, a better search engine? Automating some tedious tasks? Increased productivity? What are the downsides? People losing their jobs to AI. Artists/programmers/writers losing value from their work. Fake online personas indistinguishable from real people. Unprecedented amounts of spam and misinformation flooding the internet. Intelligent AIs automatically attacking and hacking systems at unprecedented scale 24/7. Chatbots becoming the new interface for most interactions online and being the moderators of access to information. Chatbots pushing a single viewpoint and influencing public opinion (many people complain today about ChatGPT being too "woke"). And I may just be scratching the surface here.

Re: Google denies training Bard on ChatGPT chats from ShareGPT

#47
post #18

Thankfully archive.org exists, otherwise it would not be possible to get good training data in a few years when the internet is flooded with AI content.

Isn't most of the internet available through common crawl? I don't know what percentage of training data is just that data set but i assume it's enough for anyone with enough compute and ingenuity to create a reasonable LLM

Missed the point - they are saying that, in the future, there will be no human generated content left on the Internet.

Re: Google denies training Bard on ChatGPT chats from ShareGPT

#48
post #33
post #22

Earlier quoted context omitted.

I hereby set a terms of service for everything I post on the internet from now on. OpenAI may not train future GPT models on my words or my code without my express written permission. … Somehow, I don’t think they’ll care.

Sure. If you can get everyone to create an account and agree to those terms before reading your comments, you might have a case. Otherwise, it will be considered public information, at which point it is free to be scraped by anyone (see the precedent set by the LinkedIn/hiQ case).

LinkedIn won that case on appeal, HiQ waas found to be violating the ToS, common misconception

I was pointed at a link explaining the case here on HN, after trying to make a similar point, but cannot find the link currently

edit, not the one I was pointed at, but similar

https://www.fbm.com/publications/what-recent-rulings-in-hiq-...

Re: Google denies training Bard on ChatGPT chats from ShareGPT

#49
post #16
post #6

Good luck to them. AI models are automated plagiarism, top to bottom. None of us gave OpenAI permission to derive their model from our writing, surely billions of dollars worth, but they took it anyway. Copyright hasn't caught up so all that stolen value rests securely with OpenAI. If we're not getting that back, I don't see why AI competitors should have any qualms about borrowing each others' work.

I don't get this sentiment. For some cases sure, if it repurposes your code that ignores the license fine. But it's rarely wholesale copying. It's finding patterns same as anyone studying the code base would do. As for the majority of content written on the internet through reddit or some social media, what's the harm in ingesting that? It's an incredibly useful tool that will add huge value to everyone. It's relativ…

> It's finding patterns same as anyone studying the code base would do.

This is the issue, it's not finding patterns as people do.

If I read someone's code, book, &c, that's extremely lossy. I can only pick up a few things from it in the long term.

But an ML model can store most of what it's given (in a jumbled format) and can do it from billions of sources.

It's essentially corporate piracy, but it's not legally recognized as such because it doesn't store identical reproductions.

This hasn't been an issue before because it's recent and wasn't considered valuable. But now that it's valuable and Microsoft is going to take all our jobs we have to at least consider if it's okay if Microsoft can take our work for free.

Re: Google denies training Bard on ChatGPT chats from ShareGPT

#50
post #6

Good luck to them. AI models are automated plagiarism, top to bottom. None of us gave OpenAI permission to derive their model from our writing, surely billions of dollars worth, but they took it anyway. Copyright hasn't caught up so all that stolen value rests securely with OpenAI. If we're not getting that back, I don't see why AI competitors should have any qualms about borrowing each others' work.

Yeah, I definitely like to see AI companies getting hit with their own medicine. The main problem isn't even "automated plagiarism": the pre-generative era was chock full of AI companies more or less stealing datasets. Clearview AI, for example, trained up its facial recognition technology on your Facebook photos, without asking for and without getting permission.

On the other hand, I genuinely hope copyright never "catches up", because...

1. It is a morally bankrupt system that does not adequately defend the interests of artists. Most artists do not own their own work; publishers demand copyright assignment or extremely broad exclusive licenses as a condition of publication. The bullies know to ask for all their lunch money, not just a couple bucks for themselves. Furthermore, copyright binds noncommercial actors the same as it does commercial ones, which means unconscionably large damage awards for just downloading a couple of songs.

2. The suggested ways to alter copyright to stop AI training would require dramatic expansions of copyright scope. Under current law, the only argument for the AI itself being infringing would be if it memorized training data. You would need to create a new ownership right in artistic styles or techniques. This would inflict unconscionable amounts of psychic and legal damage on all future creators: existing artists would be protected against AI, but no new art could be legally made unless it religiously hewed towards styles already in the public domain. We know this because music companies have already made their domain of copyright effectively work this way[0], and the result is endless bullshit lawsuits on people who write songs that merely "feel" too similar (e.g. Blurred Lines)

3. AI will still be capable of plagiarism. Most plagiarists are not just hoping the AI regurgitates training data, they are actively putting other people's work into the model to be modified. A lot of attention is paid to the sourcing of training data, because it's a weak spot. If we take the training data away then, presumably, there's no generative AI. However, people are working on licensed datasets and training AIs on them. Adobe has Firefly[1], hell even I've tried my hand at training from scratch on public domain images. Such models will still be perfectly capable of doing img2img or being finetuned and thus copying what you tell it to.

If we specifically want to regulate AI, then we need to pass laws that regulate AI, rather than just giving the music labels, movie studios, and book publishers even more power.

[0] Specifically through sampling rights and thin copyright.

[1] I do not consider Adobe Firefly to be ethical: they are training the AI on Adobe Stock images, and they claim this to be licensed because they updated the Adobe Stock agreement to have a license in it. Dropping a contractual roofie into stock photographers' drinks does not an ethical AI make.

Post reply on HN