Live data from Hacker News

Google denies training Bard on ChatGPT chats from ShareGPT

twitter.com

321–330 of 342 posts

Re: Google denies training Bard on ChatGPT chats from ShareGPT

#321

Earlier quoted context omitted.

It's definitely a derived work as far as copyright is concerned: the output would simply not exist without the copyrighted training data. > It's finding patterns same as anyone studying the code base would do. No, it's quite unlike anyone studying data, because it's not a person with legal rights, such as fair use, but an automated algorithm. There is absolutely no legal debate that copyright applies only to human au…

The output of human copyrighted work wouldn't exist if it weren't for humans training on the output of other humans. Humans constantly use cliches in their writing and speech, and most of what they produce is a repackaged version of what someone else has written or said, yet no one's up in arms against this mass of unoriginality as long as it's human-generated. This is anti-AI bias, pure and simple.

> This is anti-AI bias, pure and simple.

^ found the AI-based Chatbot

Re: Google denies training Bard on ChatGPT chats from ShareGPT

#323
post #294

Earlier quoted context omitted.

But the model doesn't have any agency. GPT isn't spitting out novels in the style of J.K. Rowling and sending them to publishers - a human is. GPT being instructed to tell a Harry Potter story itself is no more infringing than a child asking a parent for a made up Harry Potter bed time story. They equally infringe and undermine new or potential markets for copyrighted work. The question is "what do you do with the ma…

This argument is weak. If we agree that the production is infringing, then selling a machine that produces infringing works on demand is also infringing to the rights of the author. For example, if I sell a karaoke machine that comes with pre-recorded famous melodies without the original vocals (thus, derived works), I definitely cannot claim "only the agency of the users matter". No, even the on-demand production it…

> it's a binary compressed representation of the works used to train it, for the purpose of on-demand production of works that mimicry, to a varying, and sometimes strikingly infringing degree, the originals.

To pursue this point further, I have asked ChatGPT to "list the lyrics of a famous Madonna song":

Sure, here are the lyrics of "Like a Prayer" by Madonna, which was released in 1989:

[Verse 1]

Life is a mystery

Everyone must stand alone

I hear you call my name

And it feels like home

[Chorus]

When you call my name

It's like a little prayer

I'm down on my knees

I wanna take you there

...

(the entire lyrics of the song follow)

I think it would be very hard to argue that this model does not embed verbatim copies of the original work, but somehow reconstituted those lyrics via a parallel construction where the cultural impact of Madonna's lyrics were grasped from other fair use sources. Even in that case, it's still a word for word reproduction of the original, therefore not fair use. Therefore, the entire model or service is infringing - even if some of its productions may not be.

The ability of the model to produce copyrighted works is just a proof of the degree on which it relies on the originals; even if that ability would be blocked or somehow filtered by a plagiarism detector in a later model of Chat GPT, it would change nothing to the fundamental nature of the machine: an automated means of generating derivative works without artistic or scientific agency.

Re: Google denies training Bard on ChatGPT chats from ShareGPT

#325

Earlier quoted context omitted.

As a forum moderator, I have transitioned to relying heavily on AI-generated responses to users. These responses can range from short and concise ("Friendly reminder: please ensure that all content posted adheres to our rules regarding hate speech. Let's work together to maintain a safe and inclusive community for everyone") to lengthy explanations of underlying issues. By using AI-generated content, a small moderati…

> Original comment The original comment is much better, please stop rewriting your comments using OpenAI. > In fact, I did it for this comment. Yes, it was obvious from the second sentence. The way ChatGPT structures text by default is very different from how most humans writes. Always the same "By using", "These X can range from" etc. Padding your text with more words doesn't make it better, more words makes it wors…

Interesting, the "By using" was my own addition to shorten a long sentence it had generated that distracted from the example.

To be more clear, using AI to rewrite comments such as this one is not something I often do. My personal use of it for moderation purposes is more prompt based than pasting a long comment for grammar and spelling corrections.

What I did here was an example and that example provided the same criticism that you wrote here as a reply ("which can be tempting but may result in unintentional changes to the original message"). In other words, makes the text more verbose and sanitizes the writing style.

The prompt we use for moderation contains our site's rules and some added context. So using ChatGPT, we can paste in someone's comment and ask the bot to write a short text explaining how that comment does not follow our rules and what the user can do.

"Using the rules above, write a very short message for a user that wrote a rule breaking comment. Show empathy. Use simple English. Explain the rules that were broken. The comment is [comment here]"

Using this saves a lot of time. Is the quality of the comment not as good as it could be if it was written by a human? Absolutely. However, using AI let us change the user:mod ratio in a way. Automoderators are nothing new, what is new is that now the automoderator can take context into account and provide a customized message.

Re: Google denies training Bard on ChatGPT chats from ShareGPT

#326

Earlier quoted context omitted.

It depends if your redrawing is substantially different enough from the original image to earn copyright on its own. Your changes to an image from ChatGPT do not affect the copyrightability of the original content. If you've simply redrawn what the computer designed it may not be substantial enough to earn copyright. If you've made changes, it may only be copyrightable for those changes.

The example was redrawing something by hand that was computer generated originally. It would be pretty much impossible for a hand drawn work of art to not be sufficiently original. Hand drawn art doesn't look the same as what a computer produces. Originality has a very low threshold, simply pointing my camera at something and hitting click is almost always enough to show originality. At any rate it isn't fraud to tak…

If you think you own a design because you hand drew a version of it someone else invented, you're gonna have a bad time. Please redraw a superman picture someone else made and then go to have it copyrighted, and tell me how that goes for you.

Re: Google denies training Bard on ChatGPT chats from ShareGPT

#327

Earlier quoted context omitted.

This argument is weak. If we agree that the production is infringing, then selling a machine that produces infringing works on demand is also infringing to the rights of the author. For example, if I sell a karaoke machine that comes with pre-recorded famous melodies without the original vocals (thus, derived works), I definitely cannot claim "only the agency of the users matter". No, even the on-demand production it…

> it's a binary compressed representation of the works used to train it, for the purpose of on-demand production of works that mimicry, to a varying, and sometimes strikingly infringing degree, the originals. To pursue this point further, I have asked ChatGPT to "list the lyrics of a famous Madonna song": Sure, here are the lyrics of "Like a Prayer" by Madonna, which was released in 1989: [Verse 1] Life is a mystery…

> it would change nothing to the fundamental nature of the machine: an automated means of generating derivative works without artistic or scientific agency.

So is a Xerox machine (or Cannon copier or 'MFD')

The issue there is not "if it can" or even "if it is designed to do so" but rather "can it be used in a way that is not infringing" and "if there is an infringement from its use, the human doing that is the one liable."

And yet, there are non-infringing uses of the Xerox machine.

Even if one was to accept the position that the only thing that GPT can produce is derivative works it doesn't rule out that there are transformative and non-infringing uses of it.

Re: Google denies training Bard on ChatGPT chats from ShareGPT

#328
post #309

Earlier quoted context omitted.

Not to mention it's embarrassing. Google playing second banana to OpenAI.

That's okay if they can leapfrog. Google Maps wasn't the first map product. We all used mapquest way before. But Google Maps was technologically advanced. Ajax made maps usable for the first time. Gmail wasn't the first webmail. Hotmail had millions of customers already. But Google gave people unlimited space to store old email, whereas email in the old days filled up your inboxes and needed to be deleted. Question i…

>But Google gave people unlimited space to store old email

This isn't correct. Gmail launched with 1GB per user, which was way higher than other services, and they did keep doubling the storage space year-after-year, but it was never unlimited until Google Apps offered unlimited storage for businesses and schools.

Re: Google denies training Bard on ChatGPT chats from ShareGPT

#329
post #319

Earlier quoted context omitted.

My understanding is that the Twitter thread author works at OpenAI. Maybe I'm wrong about that.

According to his bio, he works at Vercel. He made a hobby project called ShareGPT[0] and that's probably where the accusation came from. [0] - https://sharegpt.com/

Thanks for the clarification. Aside from the OP, I haven't seen anyone from OpenAI commenting on this, so yeah, unless I've missed something I think you're correct to point that they're not involved so far.

Re: Google denies training Bard on ChatGPT chats from ShareGPT

#330
post #327

Earlier quoted context omitted.

> it's a binary compressed representation of the works used to train it, for the purpose of on-demand production of works that mimicry, to a varying, and sometimes strikingly infringing degree, the originals. To pursue this point further, I have asked ChatGPT to "list the lyrics of a famous Madonna song": Sure, here are the lyrics of "Like a Prayer" by Madonna, which was released in 1989: [Verse 1] Life is a mystery…

> it would change nothing to the fundamental nature of the machine: an automated means of generating derivative works without artistic or scientific agency. So is a Xerox machine (or Cannon copier or 'MFD') The issue there is not "if it can" or even "if it is designed to do so" but rather "can it be used in a way that is not infringing" and "if there is an infringement from its use, the human doing that is the one li…

No xerox machine comes with an embedded copy of Harry Potter that can be reproduced at the push off a button.

That's the crux of the issue, that you can't separate the training data from the derivation ability. If it's just an AI algorithm that could, when trained in a certain way, produce derivative infringing works, nobody would object to it.

Post reply on HN