Live data from Hacker News

Artificial Intelligence and Copyright: Request for comments

federalregister.gov

241–250 of 321 posts

Re: Artificial Intelligence and Copyright: Request for comments

#241

Much debate has been had about how existing copyright law applies to AI models. But once you get past that and start asking about how copyright should apply to AI models (as the copyright office is here) the answer in my mind becomes clear. Copyright, as defined in the U.S. Constitution, exists "to promote the Progress of Science and useful Arts"[1]. I can think of no better modern example of "the Progress of Science…

Number 3 doesn't really make sense. I wouldn't get copyright if I told a human artist "draw a dog". Why would that change just because I'm telling an AI to do it?

Re: Artificial Intelligence and Copyright: Request for comments

#242
post #173

Much debate has been had about how existing copyright law applies to AI models. But once you get past that and start asking about how copyright should apply to AI models (as the copyright office is here) the answer in my mind becomes clear. Copyright, as defined in the U.S. Constitution, exists "to promote the Progress of Science and useful Arts"[1]. I can think of no better modern example of "the Progress of Science…

> Copyright, as defined in the U.S. Constitution, exists "to promote the Progress of Science and useful Arts"[1]. I can think of no better modern example of "the Progress of Science and useful Arts" than AI models themselves. This makes no sense. Before AI, it’s clear that copyright itself restricts what can be done, in order to promote overall health of innovation. You can’t just say this is cool so therefore allowe…

AI is the innovation. Trying to misapply copyright here would retard the progress of science and useful arts, not promote them, because it would wrongly restrict that innovation.

The answer to big tech corporations centralizing this is to fund it publically and make it available for free to everyone as a shared summation of our culture, not lobotomize ourselves just to profit a few old dinosaurs that are relying on an outdated idea of copyright.

Re: Artificial Intelligence and Copyright: Request for comments

#243
post #106

I have never understood the fair use argument when it comes to training data. I publish a copyrighted article. Some LLM ingests it without permission, but since the output of that LLM is sufficiently different from my source article there is no violation. I publish copyrighted code. Some company decides to consume it without purchasing a license. The product they distribute is vastly different from my code itself, bu…

It's more like if I look at thousands of articles and produced a big spreadsheet containing interesting facts about those articles, like word frequencies or what words tend to come after what other words. I never before heard anyone suggest that that kind of analysis would be copyright infringing. The new thing is that someone figured out how to organize dumb facts like that in a clever way and use it to create something sometimes useful, but without actually copying parts of the original articles since those were never saved as part of the data.

Re: Artificial Intelligence and Copyright: Request for comments

#244

Earlier quoted context omitted.

Is intelligence really a factor here? Say I use the same training set as one of these LLMs, copyright protected text and all, and use it to derive a compression algorithm that uses very little space to store tokens and token sequences that are common in that huge collection of text. The resulting compression scheme includes some sort of statistical artifact derived from that copyrighted text. Is that allowed? And if…

LLMs are generative though not just compressive

Generation, prediction, and compression are all the same - the only different thing is the intent.

Re: Artificial Intelligence and Copyright: Request for comments

#245
post #145

Earlier quoted context omitted.

Some compression, yes, but the analogy oversimplifies. AI rerepresents input information in a transformative way (embedding, say) then creates new, derived and combined output from a new input (e.g prompt). It's not just lossy compression. It's potentially novel.

Phrases like "transformative way" are meaningless woospeak to me. Everything is a transformation. Sulpose I run a linear convolution on ten images and average them. Is the result "new"? Does it not contain the original images? Subspaces and mappings don't create anything "new" any more than SVD does. This is just playing digital Ship of Thesius.

> Phrases like "transformative way" are meaningless woospeak to me

Fortunately we live in a society that supports specialization where something that is woospeak to a smart person can still be a very well understood topic. AI transformations are methodologically well documented, even if transparency of neural network node activations is yet to be fully formalized.

Re: Artificial Intelligence and Copyright: Request for comments

#246
post #106

I have never understood the fair use argument when it comes to training data. I publish a copyrighted article. Some LLM ingests it without permission, but since the output of that LLM is sufficiently different from my source article there is no violation. I publish copyrighted code. Some company decides to consume it without purchasing a license. The product they distribute is vastly different from my code itself, bu…

The difference is that the software is in active use in your scenario. Consider if you took a copyrighted program and calculated the sha hash of it. You can then use that hash without needing to have a copy of the original program. The hash is also not infringing, because it's a simple fact.

Re: Artificial Intelligence and Copyright: Request for comments

#247
post #85

Earlier quoted context omitted.

> is that an AI, being a statistical model and not generally intelligent, should not be allowed to disregard the copyright of its source material But then you are just shifting the problem forward by an inch. What happens when tomorrow someone declares that their model is generally intelligent and is therefore allowed to disregard copyright when training just like a person can?

Is it your experience that people's facial declarations cary the day in legal disputes? It's not mine. Rather, it seems like the whole thing is designed to provide scrutiny against bare facial declarations that something is true or false. I see this on HN all the time "someone just has to claim" "someone just has to say". Yeah... that's not how it works. People can say whatever they want, that doesn't mean it satisfi…

Intelligence lacks any legal definition, for starters. And if a law like that will provide an arbitrary line in the sand, it will just disincentivize AI research in general.

Re: Artificial Intelligence and Copyright: Request for comments

#248
post #85

Earlier quoted context omitted.

> is that an AI, being a statistical model and not generally intelligent, should not be allowed to disregard the copyright of its source material But then you are just shifting the problem forward by an inch. What happens when tomorrow someone declares that their model is generally intelligent and is therefore allowed to disregard copyright when training just like a person can?

Is it your experience that people's facial declarations cary the day in legal disputes? It's not mine. Rather, it seems like the whole thing is designed to provide scrutiny against bare facial declarations that something is true or false. I see this on HN all the time "someone just has to claim" "someone just has to say". Yeah... that's not how it works. People can say whatever they want, that doesn't mean it satisfi…

[flagged]

Re: Artificial Intelligence and Copyright: Request for comments

#249

Earlier quoted context omitted.

That is why we should halt AI completely and do a more thorough analysis of its societal-level implications before blindingly putting it out there. Because when new technology is introduced, it makes it almost impossible to stop using it due to the way our current society is setup (as a sensitive machine that is very quick to reward any gains in efficieny and economic output as opposed to sustainability).

Simply get every nation on earth to cooperate and ban a vaguely described technology that hundreds of billions of computers can run to varying degrees of efficiency! It's that easy! If we can't get this level of cooperation for global warming, which is largely the result of a few dozen companies, what makes you think that governments across the world can stop everyone with access to a device with a reasonable amount…

As I see it for country to disconnect from AI they'd need to go fully isolationist. Disconnect the internet fully from the rest of the world, block all mail, block all imports, disconnect all financial markets, etc. Otherwise they'd simply become a consumer of the AI output of other nations. For example, if AI can predict stock markets better by efficiently parsing financial documents then eventually foreign investors leveraging AI would dominate. That will work for a while but eventually AI will get cheap and efficient enough to be easily hidden. So now you need the government to police for AI, search people, track everything they do and so on. Criminals using AI will rise to power and prominence until stopped. Essentially prohibition or the war on drugs all over again.

edit: And of course to better understand and deal with the AI threat the government would be given exemptions to the laws. These exemptions would be used more and more widely by the government to exert power while the population is not allowed to even look into what is possible.

Re: Artificial Intelligence and Copyright: Request for comments

#250
post #56

Earlier quoted context omitted.

My opinion as a SWE who is dating a lawyer (joke, not a serious qualification but it does provide some insight): Generative models traverse and interpolate high dimensional state spaces. These state spaces are created from input data. I would argue people do the exact same thing - the first main difference is we can use novel inputs (e.g. we can use images or words to develop our music/temporal state spaces and vice…

The analogy doesn't hold when you consider the sheer scale of the problem. I can outright buy a machine for a few thousand dollars that can crank out a faithful rewrite of every Stephen King novel without the shitty endings and nonsense plot points. It can do it in a few days, maybe a couple of weeks at most. To do that with human labor would take years and cost hundreds of thousands, if not millions of dollars. Inst…

Hello.

Maintaining a system like Netflix or AWS or even Amazon will require insane amount of people and time, if possible at all within a finite time, without all the computers doing work for us in seconds that would take humans ages to do.

Post reply on HN