Live data from Hacker News

Artificial Intelligence and Copyright: Request for comments

federalregister.gov

191–200 of 321 posts

Re: Artificial Intelligence and Copyright: Request for comments

#191

Terms of Use You are prohibited from using the content of this site in "large language models" or any other usage for the purpose of "artificial intelligence". Liquidated Damages

It wouldn’t stop them arguing fair use.

You can’t have a license that restricts fair use because fair use doesn’t require a license.

Re: Artificial Intelligence and Copyright: Request for comments

#192
post #166

Earlier quoted context omitted.

> What stops you from getting a profit? OpenAI and Stable Diffusion not paying for their dataset. I don’t believe GitHub asked for my contribution to Copilot.

But you weren't receiving profit from your works originally? So therefore, why does it matter what someone else was doing?

If someone jacks my car while I'm asleep, races it, and wins a prize, and returns it before I wake up, they haven't deprived me of anything, profited off my property, and it's still wrong.

Profit and deprivation are not and never will be a good tests in determining things like this.

Re: Artificial Intelligence and Copyright: Request for comments

#193
post #82

Earlier quoted context omitted.

I can appreciate how this line of thinking might be attractive. But IMO the human machine comparison doesn't lend itself much credence. We shouldn't assume that just because a human is allowed to do something, a machine is automatically allowed to do the same thing, too. I think some care should be taken when considering if we allow machines to have the same privileges as humans.

A machine is just a tool. It is the creator and the user of the machine that has the privileges he uses the tool with. I think we should be careful not to anthropomorphize, attribute agency, responsibility and autonomy to something that is essential a better photoshop plugin.

I don’t think parent anthropomorphizing anything. The ones who anthropomorphize are saying that machines should be covered by fair use, because they have similarities with humans.

This is not about the rights of a machine but about how one human product is consumed by another human product. This is just a commercial supply chain: if you make a model, you need human data. You generally need to compensate your suppliers of “raw material”.

Re: Artificial Intelligence and Copyright: Request for comments

#194

I believe we first need to answer the question of whether the copyright of the AI model’s source text or images affects the output. My opinion — and note I’m a software engineer, not a lawyer — is that an AI, being a statistical model and not generally intelligent, should not be allowed to disregard the copyright of its source material. This would, I think, require the AI’s creator to secure a license for all of its…

I don’t think it makes sense for both model builders and the model’s users to separately obtain licenses for the same works used in the training set. A model trained on several copyrighted data sources cannot somehow be used in a way depending on a subset of those sources. So all parameters of usage and compensation should be settled by contract between the model builder and copyrighted data supplier, before the copy…

Problem is, how can you determine if the model contains copyrighted material? The laws governs copyright through ownership, so in order to claim copyright infringement you have to be able pinpoint a specific person and prove that their work is somehow embedded in the gradients, which is not practically possible at the point. It's just like how you can't practically enforce copyright on encrypted data unless you ban encryption altogether.

Re: Artificial Intelligence and Copyright: Request for comments

#195
post #166

Earlier quoted context omitted.

But you weren't receiving profit from your works originally? So therefore, why does it matter what someone else was doing?

If someone jacks my car while I'm asleep, races it, and wins a prize, and returns it before I wake up, they haven't deprived me of anything, profited off my property, and it's still wrong. Profit and deprivation are not and never will be a good tests in determining things like this.

> jacks my car

no, they downloaded your car.

Re: Artificial Intelligence and Copyright: Request for comments

#196

AI Jesus chat-bot could claim copyright over biblical content. In theory, a company that owns Christian (c 2023) content could be filing DMCA claims every Sunday. The silliness of digital-racketeers must end at some point. =)

I'm not sure what it's feelings are on the topic but Ask Jesus is a thing on twitch rn. NSFW (It's chat generated content on Twitch)

https://www.twitch.tv/ask_jesus

It takes chat prompts and weaves together scripture, The Internet, odd pronunciation, callbacks to other questions, and hilarious fails when chat gets, uh, playful.

AI Jesus just said there is a time and place for everything and that it might just be the time to buy Starfield. It knew it was a game being only prompted with, 'Jesus, should I buy Starfield?'

Re: Artificial Intelligence and Copyright: Request for comments

#197

Earlier quoted context omitted.

You've described a viable startup business plan.

It is what I enjoy doing, but someone has already launched a product. =) text-with-jesus: https://apps.apple.com/us/app/text-with-jesus/id6446922759

Nope, that's not the plan. The plan is to start an non-practicing entity that copywrites religious texts and sues others who use them.

Re: Artificial Intelligence and Copyright: Request for comments

#198
The entire discussion about AI and copyright strikes me as a bit naive.

Right now, we are in a situation where nobody quite knows what these AI models are useful for. We have some inkling that they might be extraordinarily useful for making money -- but not precisely how, not even the companies that are developing the models themselves.

Once they money starts, the debate over copyright will fall exactly into the economic seams between the major players involved:

- new tech orgs who are monetizing models will say that the model is "exactly as humans are": they see copyrighted works in training, and then produce wholly original outputs. And of course that the model weights themselves are, like the outputs of employees, completely owned by the company.

- incumbents who stand to lose out on the new gold rush will say that every single output of a model belongs to them if just a single image or sentence was seen in training. And that because of that, we really should just shut the whole thing down, because how could you ever prove that a model was not trained on copyrighted material?

The faultlines will entirely rest on who has more power, hard and soft. How much can they influence the legal system, either by spending $ to hire legal talent or by sheer soft politicking, balanced with how favorable they appear to the general public who uses their product (or consumes their media). I suspect that the end result of this debate is a "legal" way of doing things accessible only to the extremely large players, and a small, politically insignificant collection of individuals, hackers, and startups who aim to unseat those large players (or just flat-out train "illegal" models). The worst possible end result is that the legal system is just too fossilized to deal and tries something draconian like not allow datacenter-scale GPU compute.

As an aside, I predict a sizeable space for companies that do "compliance" -- asserting the copyright status of a dataset, perhaps even themselves using ML. That market will carve off and leave rotting a sizeable chunk of the new money's ML profits.

It's fun to talk about this, I guess. But remember that what you or I have to say about what a machine learning model philosophically is has no bearing what-ever when it comes to the actual ability for individuals, startups, or large players to use models.

I will predict though: enjoy Llama2 while it lasts. Like the internet, it will become fully assimilated into the larger intellectual property machine.

Re: Artificial Intelligence and Copyright: Request for comments

#199
post #106

I have never understood the fair use argument when it comes to training data. I publish a copyrighted article. Some LLM ingests it without permission, but since the output of that LLM is sufficiently different from my source article there is no violation. I publish copyrighted code. Some company decides to consume it without purchasing a license. The product they distribute is vastly different from my code itself, bu…

You can actually prove that some company distributed a product with your code but you can't practically prove that LLM contains your source article.

Re: Artificial Intelligence and Copyright: Request for comments

#200
post #2

I’ve made so much money stacking my pitch decks and websites with AI generated media that I don't care if someone copy and pastes it and uses it commercially too People married to their prompt engineering outputs are really missing the forest for the trees

I am having trouble making sense of that last statement - can you rephrase it?

A lot of people would like their AI creations to have to copyright protection, but it isn't necessary to making money
Post reply on HN