Live data from Hacker News

An IP attorney’s reading of the Stable Diffusion class action lawsuit

katedowninglaw.com

311–320 of 337 posts

Re: An IP attorney’s reading of the Stable Diffusion class action lawsuit

#311

Earlier quoted context omitted.

That the one uses generic sounds as input and the other uses specific art as input.

Stable Diffusion uses all of the art as input and this is actually incredibly important. In fact, any specific piece of art can be removed and it will still work the same and this is also incredibly important. Both show that there is no intent for individual infringement, with along with no infringing material being produced and the significant non-infringing commercial use like family photo touch-ups, it makes for a…

> In fact, any specific piece of art can be removed and it will still work the same and this is also incredibly important.

This is provably false: if any specific piece can be removed and there is no difference then you can keep doing that until there is no art left. I'd bet a very substantial amount of money that the output at that point will definitely have changed. Induction is a powerful thing.

Re: An IP attorney’s reading of the Stable Diffusion class action lawsuit

#312

Earlier quoted context omitted.

Stable Diffusion uses all of the art as input and this is actually incredibly important. In fact, any specific piece of art can be removed and it will still work the same and this is also incredibly important. Both show that there is no intent for individual infringement, with along with no infringing material being produced and the significant non-infringing commercial use like family photo touch-ups, it makes for a…

> In fact, any specific piece of art can be removed and it will still work the same and this is also incredibly important. This is provably false: if any specific piece can be removed and there is no difference then you can keep doing that until there is no art left. I'd bet a very substantial amount of money that the output at that point will definitely have changed. Induction is a powerful thing.

I didn’t say remove all the art, I said remove a specific piece of art. No one needs to do thought experiments. Remove all the plaintiff’s works, retrain the model, and the model is still just as useful.

Or if we want to do thought experiments, randomly remove any thousand images and the tool is still just as useful. If it’s random, not specific images, that are removed, then it is non-specific images that power the tool… yes, millions of those images, so a specific quantity is required but no specific works.

Re: An IP attorney’s reading of the Stable Diffusion class action lawsuit

#313

Earlier quoted context omitted.

I'm going to hone in on "original art"... What is original to begin with? What's original about Bob Dylan's Blowin in the Wind? Certainly not the form! It's a standard AB folk song. Certainly not the chords! The melody? Sure, but very bounded by Western music theory and containing a number of common American melodic tropes. The words and specific rhymes have all be used before in previous poems and light verse. He us…

> What's original about Bob Dylan's Blowin in the Wind? The fact that he claims he made it, and that this went uncontested for decades is fairly strong proof that it really is his.

So yes, that’s the legal proof of ownership… no one claimed to have written the song otherwise.

But I thought you were interested in originality outside of just the legal perspective? You know, that crimes can be committed without evidence and all?

What makes Blowing in the Wind different from another folk song of the time? What makes them the same? Why are some non-original aspects allowed in a work considered original?

Re: An IP attorney’s reading of the Stable Diffusion class action lawsuit

#314

Earlier quoted context omitted.

Sorry, I’m unfamiliar with civil law societies but here in the Anglosphere if there is no body, there is no murder. Obviously a witness testimony of a person being tossed into a volcano is evidence of a body, etc. So maybe they are locking up and trying people in Europe without evidence of crimes as being committed and having people shift the burden of proof to the accused but here in the United States we really do a…

I'll just leave this here for now. I'm sure you'll have plenty of reasons to say that 'there was evidence after all' but 'no body no murder' is at least to my reading simply not true. https://en.wikipedia.org/wiki/Murder_conviction_without_a_bo... Other observations about how legal systems elsewhere work are ignored on account of the opening sentence.

I stated in plain English that a murder needs a body is not to be taken literally. Obviously evidence is still required. I used an example of witness testimony.

So at this point have you run out of any ammo other than meaningless nitpicks?

Re: An IP attorney’s reading of the Stable Diffusion class action lawsuit

#315

Earlier quoted context omitted.

Are the existing laws written in a way that is favorable to generative AI? No. But the laws do exist. Whether or not one believes those laws apply to generative AI seems to be based on one's belief in how similar that AI software is to humans. I'd argue that systematically ingesting 2.3 billion images is not remotely human (one of a myriad of reasons the comparisons break down), and that it is a long stretch to claim…

There are exactly zero laws about using openly published materials for learning. Human learning but also machine learning. There's implicit assumption that if you can get a hold of a copy and manage to learn from it you are free to use what you learned in your creations.

> There's implicit assumption that if you can get a hold of a copy and manage to learn from it

But there are explicit laws about how you acquire copies of things, and whether they apply seems to be based on what someone believes “learning” to be.

Your claim relies on the belief that a computer ingesting images is similar to a human learning from those images.

Re: An IP attorney’s reading of the Stable Diffusion class action lawsuit

#316

Earlier quoted context omitted.

There are exactly zero laws about using openly published materials for learning. Human learning but also machine learning. There's implicit assumption that if you can get a hold of a copy and manage to learn from it you are free to use what you learned in your creations.

> There's implicit assumption that if you can get a hold of a copy and manage to learn from it But there are explicit laws about how you acquire copies of things, and whether they apply seems to be based on what someone believes “learning” to be. Your claim relies on the belief that a computer ingesting images is similar to a human learning from those images.

Isn't that how people learn art and writing: they study good artists and writers?

In the United States, a legal derivative work, which isn't a parody, needs to make three substantive changes from the original. It's fair to say that creating new works in the same style or 'look-and-feel' of an artist would satisfy that at prima facie.

Re: An IP attorney’s reading of the Stable Diffusion class action lawsuit

#317
post #197

Earlier quoted context omitted.

I think it's not going to be that hard to argue that the company is infringing the copyright of those whose images they are using. Especially once the judge is show how similar the output of SD can be to a particular artist's images with the right prompts (proving that SD has memorized a significant amount of those images).

Some artists images just don't contain much entropy though? If an AI art engine outputs a frame of solid blue, is it infringing the copyright of Yves Klein's solid blue "IKB 79"? I think that some artists' styles can be accurately replicated without training on any of their work: because the artists' style is generic enough that it can be exhaustively encoded via the works of others. This seems like a bad test becaus…

> If an AI art engine outputs a frame of solid blue, is it infringing the copyright of Yves Klein's solid blue "IKB 79"?

Probably not, though even this may be debatable given the specific prompt and specific similarities (for example, if it generated the exact color and exact aspect ratio for a prompt like "Yves Klein IKB 79", I could see an argument for infringement; if it generated the same thing for a prompt like "filled-in 16:5 rectangle with color #0000FF", it would be arguable that it isn't).

> I think that some artists' styles can be accurately replicated without training on any of their work: because the artists' style is generic enough that it can be exhaustively encoded via the works of others.

The important question in copyright is, theoretically, if you actually copied the specific piece or if you happened to create a similar-looking piece by accident. The difficulty of proving one or the other varies with the specific circumstances. For example, it's rather hard to claim you happened to paint a picture almost identical with Picasso's Guernica. It's rather easy to claim that signing an empty piece of canvas was entirely your idea and you weren't copying Dali's signed empty canvases.

Similarly for AI, the copyright discussion will come down to how much of its output is identifiable pieces of other's artworks, especially when prompted for such. I personally haven't played with it, but if it's possible to get SD or other similar models to produce identical or very similar copies of (pieces of) some artist's works using relatively simple prompts*, it should be a pretty slam-dunk case of copyright infringement.

* by this I mean prompts that don't themselves encode the information, like I showed in my "16:7 rectangle with solid color #0000FF" example.

Re: An IP attorney’s reading of the Stable Diffusion class action lawsuit

#318

Earlier quoted context omitted.

> Diffusion models can be compared to a superhumanly talented artist that can be cloned in unlimited fashion by anyone having the software and hardware means. How can you claim with a straight face that this is a better explanation of what an NN is? An NN is simply an approximation of a multi-valued function, whose parameters are adjusted by minimizing the difference between the output of the NN and the output of the…

> An NN is simply an approximation of a multi-valued function, whose parameters are adjusted by minimizing the difference between the output of the NN and the output of the real function for a certain input. Right, but that equally fits a biological NN if you zoom in that close. You'll need more than wikipedia to appreciate what deep-neural-networks are doing here, it's dimensional space that's key. What DNNs do that…

> Right, but that equally fits a biological NN if you zoom in that close.

First of all, we have no idea how biological NNs learn, how they represent information, how they reason etc. Given what we do know, there is no reason to assume any similarity with ANNs on any of those fronts. Just to give one example, we know very well that a single biological neuron encodes significant information and is capable of reasoning on its own. In fact, even non-neuronal biological cells are capable of such - especially looking at single-celled organisms, which display extraordinarily complex behaviors with no NN in sight.

Second of all, we don't exactly understand how the huge models we have actually encode the higher-level representations of the training set that they store. Of course, we can say for sure that they are not literally storing a copy of the data on simple space requirements. But we can also say for sure that their "understanding" of the data, as well as their capacity for inference, is significantly different from our own - since they make certain mistakes that are nearly impossible for a human to make, while showing super human abilities in other aspects. So, if anything, we must conclude that whatever it is they are doing, it is most certainly not a way of understanding the information the way we understand it.

Re: An IP attorney’s reading of the Stable Diffusion class action lawsuit

#319

Earlier quoted context omitted.

> Where is the form to remove my reddit comments from chat gpt training data? Or my blog posts from gpt training data? More pointedly, how do I keep my GPL'd code from spewing, license free, out of CodePilot?

I think that's the point of this blog post: it doesn't matter if the inputs are copyrighted, it matters if the output is infringing. It appears to be almost impossible to directly recreate a source image with SD, but it seems Copilot tends to produce a single input as its output, verbatim. Copilot isn't doing "synthesis" as does SD, it's acting more like a search engine.

Look at these images:

> https://huggingface.co/spaces/stabilityai/stable-diffusion/d...

They were prompted with the text "Mona Lisa Smile". Would you not say that they are an extremely close reproduction of the Mona Lisa, with barely any kind of synthesis?

Re: An IP attorney’s reading of the Stable Diffusion class action lawsuit

#320

Earlier quoted context omitted.

Tone aside, thanks for the feedback. It doesn't excite me to hear my comment came across that way, but I'm trying to clarify what was a misinterpretation of what I was trying to say. > I believe your argument comes down to “computers are super powered compared to humans doing the same thing”? Is that accurate? No, that doesn't really touch it. The speed/power disparity between humans/computers at certain tasks are ce…

> why would it be the same for an automated process? It's perfectly acceptable for a human being to drive a a car, but driving one drunk is completely unacceptable. Conversely, there is no rule against creating or consuming art while intoxicated. So to answer your question, because it is not a matter of life and death. Take your argument and apply it to mass produced goods that were once the realm of only skilled cra…

> It's perfectly acceptable for a human being to drive a a car, but driving one drunk is completely unacceptable. Conversely, there is no rule against creating or consuming art while intoxicated.

You've lost me here. Are you saying that the most important factor when judging whether or not something is appropriate is based on whether or not the activity is dangerous enough to be fatal?

There are plenty of laws and cultural/ethical norms that restrict behavior for many other reasons.

> If a person never leaves and keeps notes, yes, it is exactly the same. I'd call the police for stalking.

You're arguing that a person taking notes with a pen and paper is the same as a video camera recording the same scene?

> The issue here is privacy, which is tangential to AI reproducing the styles of known individuals.

The point is that two forms of "seeing", one mechanical, and one biological, have two very different implications. If you don't believe that, ask the hypothetical person with a notebook to provide you with a 4K rendering of the scene over the last 30 days.

The AI reproducing art is just a single use case. The point of concern has little to do with how innocuous it is to produce images, but whether or not it is acceptable to use arguments about humans when judging what is or is not acceptable in an AI program.

> Completely disagree with you about the nature of learning here. If a person produces art in the style of an individual, they have no idea the internal machinations of the original artist, they just "appear to 'know' what they are doing".

Frankly, this is nonsense. We may not understand all of the underlying processes involved in learning, but we certainly know a lot more than nothing. Even if we knew literally nothing at all about the human brain, there is no standing to conclude that this lack of knowledge must imply that humans use some internal denoising algorithm when imagining what they will draw next.

We know enough to know that human processing of information is subjective, contextual, cultural, emotional, and there are a myriad of factors involved.

We know enough to know that what software like Stable Diffusion is doing looks very little like the human process for achieving a similar outcome, even if there are biologically inspired components inside.

Post reply on HN