Live data from Hacker News

Artificial Intelligence and Copyright: Request for comments

federalregister.gov

201–210 of 321 posts

Re: Artificial Intelligence and Copyright: Request for comments

#201

Earlier quoted context omitted.

I don’t think it makes sense for both model builders and the model’s users to separately obtain licenses for the same works used in the training set. A model trained on several copyrighted data sources cannot somehow be used in a way depending on a subset of those sources. So all parameters of usage and compensation should be settled by contract between the model builder and copyrighted data supplier, before the copy…

I slightly disagree, in that I think the person using the tool should bear the burden of copyright. I.e. if the model outputs something under copywrite it merely can't be republished. In this same way, i can use Photoshop on proprietary data but I can't necessarily sell the results.

I wonder if that analogy represents the same thing. Speaking purely from a non-legal perspective on the ethics in my mind:

When you use Photoshop on propriety data you're providing the original data and choosing what manipulation to make (i.e. what tool) and directly creating the output. It makes sense that if you redistribute this it may be copyright violation.

When you use Copilot or ChatGPT for programming you're typically asking a non-proprietary question or accepting suggestions it's making based on non-proprietary (or proprietary to you) code in the file. You also don't dictate the manipulation process a black box deep learning model does (i.e. I haven't asked it to do something that could be reasonably thought to be a copyright violation).

Am I then responsible for the fact that Copilot is fooling me with effectively copy-pasted copyrighted code when it's being presented to me as generated by the software and I haven't instructed the software to commit a copyright violation? I'm not sure if intent matters for copyright, I assume it doesn't but perhaps that's a missing piece to this.

Diffusion models are gray to me, if you're asking/prompting with "Mickey Mouse riding a horse" I can see the argument that the prompt itself can be interpreted as asking the model to commit copyright violation and the user is just hiding behind a layer of abstraction. If I ask the model to spit out "a picture of a smiling cartoon woman" and it generates a Betty Boop lookalike is that still the users fault?

It seems to me like passing the burden to the user could be reasonable but would need some safe harbor type of exception. It'll be really interesting to see what the courts decide.

Re: Artificial Intelligence and Copyright: Request for comments

#202

I believe we first need to answer the question of whether the copyright of the AI model’s source text or images affects the output. My opinion — and note I’m a software engineer, not a lawyer — is that an AI, being a statistical model and not generally intelligent, should not be allowed to disregard the copyright of its source material. This would, I think, require the AI’s creator to secure a license for all of its…

I don’t think it makes sense for both model builders and the model’s users to separately obtain licenses for the same works used in the training set. A model trained on several copyrighted data sources cannot somehow be used in a way depending on a subset of those sources. So all parameters of usage and compensation should be settled by contract between the model builder and copyrighted data supplier, before the copy…

> I don’t think it makes sense for both model builders and the model’s users to separately obtain licenses for the same works used in the training set.

I'm torn on who should pay, and where and when. In the world of patents, there's often an option/split. Say a chip manufacturer wants to build H265 decoding into their hardware. The chip manufacturer could buy the license. Or the purchaser (who probably is building some sort of board or device around the chip) could pay for the license. Or they could disable that functionality in the end product, and the consumer could pay for a license (or not, if they don't care about that feature).

The most common is usually the middle option: the end-device manufacturer (or brand that eventually sells the product) will pay for the license.

But I'm not sure if this works all that well for an AI model. With hardware, the license is usually paid per unit. It's easy to see that one chip = one license. If the model builder buys a license, that model could be used one time or 100 million times. Tracking use like that probably isn't all that practical, but I think it's safe to say that a 100-million-use model should probably pay more for a license than a single-use model.

So maybe the model builder should be responsible for attaching a comprehensive "copyright history" to the model, and users should have to pay for a license based on their use? Again, not sure how to track that. But I guess general software licensing has similar problems when you can "hide" usage.

Re: Artificial Intelligence and Copyright: Request for comments

#203

Earlier quoted context omitted.

It is what I enjoy doing, but someone has already launched a product. =) text-with-jesus: https://apps.apple.com/us/app/text-with-jesus/id6446922759

Nope, that's not the plan. The plan is to start an non-practicing entity that copywrites religious texts and sues others who use them.

Sounds like you are working with the other gentlemen's organization... you know the one that is likely... will "McKinsey & Company" be the secret behind these WIPO shenanigans too?

There was a 30% drop is chatGPT users in 30 days, so I think we are now past the "early majority" stage of the hype cycle. Perhaps NVIDIA GPUs will be a little less ridiculous next year too.

Have a glorious day =)

Re: Artificial Intelligence and Copyright: Request for comments

#204

Earlier quoted context omitted.

I don’t think it makes sense for both model builders and the model’s users to separately obtain licenses for the same works used in the training set. A model trained on several copyrighted data sources cannot somehow be used in a way depending on a subset of those sources. So all parameters of usage and compensation should be settled by contract between the model builder and copyrighted data supplier, before the copy…

Problem is, how can you determine if the model contains copyrighted material? The laws governs copyright through ownership, so in order to claim copyright infringement you have to be able pinpoint a specific person and prove that their work is somehow embedded in the gradients, which is not practically possible at the point. It's just like how you can't practically enforce copyright on encrypted data unless you ban e…

1. If you know your copyrighted material was in the training dataset is that not sufficient?

2. From a legal perspective do you actually have to prove it's embedded in the gradients? If I draw an exact copy of Mickey Mouse from memory and sell it I didn't think Disney had to prove I've ever actually seen Mickey Mouse before or point to where the image of him is embedded in my brain.

Re: Artificial Intelligence and Copyright: Request for comments

#206

AI Jesus chat-bot could claim copyright over biblical content. In theory, a company that owns Christian (c 2023) content could be filing DMCA claims every Sunday. The silliness of digital-racketeers must end at some point. =)

I'm not sure what it's feelings are on the topic but Ask Jesus is a thing on twitch rn. NSFW (It's chat generated content on Twitch) https://www.twitch.tv/ask_jesus It takes chat prompts and weaves together scripture, The Internet, odd pronunciation, callbacks to other questions, and hilarious fails when chat gets, uh, playful. AI Jesus just said there is a time and place for everything and that it might just be the…

Ah, but can you buy Starfield with the change from within yourself.

Also, the buggy nature of Armored Core 6 could be fun too.

Happy computing =)

Re: Artificial Intelligence and Copyright: Request for comments

#207
post #145

Earlier quoted context omitted.

Some compression, yes, but the analogy oversimplifies. AI rerepresents input information in a transformative way (embedding, say) then creates new, derived and combined output from a new input (e.g prompt). It's not just lossy compression. It's potentially novel.

Phrases like "transformative way" are meaningless woospeak to me. Everything is a transformation. Sulpose I run a linear convolution on ten images and average them. Is the result "new"? Does it not contain the original images? Subspaces and mappings don't create anything "new" any more than SVD does. This is just playing digital Ship of Thesius.

Congratulations, you just discovered that copyright is a weak and ill-defined concept.

Re: Artificial Intelligence and Copyright: Request for comments

#208
post #32

There are three copyright issues here; datasets, model weights, and model outputs. Dataset copyright is pretty well defined and things can often be used under fair use. Fair use decisions are done with a four prong test and really decided by the courts on a case-by-case basis. Model weights cannot currently be copyrighted. They are the output of a mechanical process over the dataset. However, software faced a similar…

You've rather conspicuously failed to mention a fourth, with two parts, and the first listed in the article abstract: "the use of copyrighted works to train AI models, the appropriate levels of transparency and disclosure with respect to the use of copyrighted works".

Outputs are the third mentioned: "the legal status of AI-generated outputs."

Re: Artificial Intelligence and Copyright: Request for comments

#209
Imagine the future where we have instead of HDMI recording cards that circumvent crappy encryption something like a downscale and upscale neural co processor to circumvent copyright protection.

Last time they tried to copyright everything with DMCA it actually turned out to be a useless law.

Why not fix DMCA first and make malicious intent actually prosecutable - and then move on to transformative topics?

Can't fix copyright if copyright itself is already broken in the law. 90 years made sense in a world of pen and paper, but not in a world with the internet.

Re: Artificial Intelligence and Copyright: Request for comments

#210
post #185
post #152

Earlier quoted context omitted.

> if the generated text/image/sound is a nearly identical copy of the original material they don’t recognize how does the industry today deal with artists that "copy" off some other works? This isn't a problem with AI at all - just that AI provides a tool to generate such works faster.

The difference is the artists assertion that it’s either original or a copy from something else. DALLE 2 can’t tell you if it’s original or not. These AI’s have no idea and the company or group that created them doesn’t review individual output so they can’t say either.

> DALLE 2 can’t tell you if it’s original or not

whoever pressed the button to run DALLE will make the assertion, just like whoever was running photoshop to make the image today would make the same assertion.

Post reply on HN