Live data from Hacker News

AI Data Laundering

waxy.org

111–120 of 120 posts

Re: AI Data Laundering

#112

I've got a couple examples of Stable Diffusion replicating watermarks along with similar swatches of imagery into scenes from the same prompt [1]. A single case of this should be enough to file a massive lawsuit if the art were recognizable to the creator. [1] https://news.ycombinator.com/item?id=33061707

The model learns all attributes of the images it's trained on, including that some have a watermark. The fact that it generates a watermark in some images doesn't mean that that is a 1:1 image from the training set, it just means to the model some images seem to have a watermark, so it will add it sometimes. Often you can just add "no watermark" (or add it as a negative prompt with some weights) and re-use the same s…

No, it means that it is reproducing the original work and is not producing a new original work. It is basically a really fancy Instagram lense, but it is still 100% derived from the underlying works and therefor derivative instead of a newly created non-derivative work.

Re: AI Data Laundering

#113
post #88

Earlier quoted context omitted.

But literally (and I use that word literally) none of the pictures contain copyrighted material.

I don't know how people can make these strong statements about anything in law. Disney have won cases in court were some artist has drawn their own version of Mickey Mouse, similarly try writing a story about some kids in a wizard school and you need to be extremely careful not to violate (or at least get taken to court) for Harry Potters copyright. I'm pretty certain image production models have produced some images…

You are confusing copyright with trademark. Or, provide a link showing the images case was decided on copyright issues, and I’ll reconsider my position.

Re: AI Data Laundering

#114
post #83
post #63

Earlier quoted context omitted.

The court’s summary also mentions this aspect of differing marketplaces: “… the revelations [i.e. the information served by Google Book Search] do not provide a significant market substitute for the protected aspects of the originals.” This doesn’t apply to AI image generators which are clearly a “market substitute” for the protected originals used to train the system. For this reason I’d expect someone like Getty to…

AI image generators are clearly not a market substitute for images, they are a tool that can be used to create market substitutes, but not themselves one.

Depends on what the product is. For example with openAI Dalle-2 the product is very clearly the generated image. You even pay per image. Also this is kind of what this article is about. Arbitrarily separating the pieces in order to evade copyright.

Re: AI Data Laundering

#115
post #63

Earlier quoted context omitted.

The court’s summary also mentions this aspect of differing marketplaces: “… the revelations [i.e. the information served by Google Book Search] do not provide a significant market substitute for the protected aspects of the originals.” This doesn’t apply to AI image generators which are clearly a “market substitute” for the protected originals used to train the system. For this reason I’d expect someone like Getty to…

Note the "protected aspects of the originals" part. AI generated images don't produce outputs that contain protected aspects.

That’s for a court to decide, ultimately. Something doesn’t have to be a bit for bit copy to be a protected aspect.

Re: AI Data Laundering

#116

I've got a couple examples of Stable Diffusion replicating watermarks along with similar swatches of imagery into scenes from the same prompt [1]. A single case of this should be enough to file a massive lawsuit if the art were recognizable to the creator. [1] https://news.ycombinator.com/item?id=33061707

The model learns all attributes of the images it's trained on, including that some have a watermark. The fact that it generates a watermark in some images doesn't mean that that is a 1:1 image from the training set, it just means to the model some images seem to have a watermark, so it will add it sometimes. Often you can just add "no watermark" (or add it as a negative prompt with some weights) and re-use the same s…

It may or may not be a 1:1 image, but I think it's significant that in both cases, with different seeds, what is directly behind / right of the watermark is a pretty similar building with different distortions applied to it. I'm not sure what the difference is between "learning" from a particular image and encoding that image with a lot of compression, when in either case the usage more or less reliably reconstructs the image algorithmically.

If I have a photographic memory and I memorize the Coca Cola logo and then draw it into a commercial work by decoding the firing of my neurons into muscle movements, the storage and retrieval method I used has no bearing on whether I infringed on their copyright.

Re: AI Data Laundering

#117
post #52

Earlier quoted context omitted.

> Sure but unless you bring down capitalism people will still need to work to eat and most will want to use their hard-earned creative skills to make a living. The concept of a UBI (universal basic income) isn’t inherently in conflict with capitalism. I believe that it is actually in coherence with the idea of Universal Human Rights, as defined by the UN in the 1940s. Perhaps that would be the culmination of anything…

The problem is that UBI is in conflict with arithmetics. Short of near-total redistribution, it's impossible to provide a decent level of UBI for everyone. Total redistribution doesn't work, because economy needs markers as ways of price / demand discovery, and markets apparently lead to power-law distribution, not flat. IMHO, the realistic option is a thick enough safety net for those who is going through a rough sp…

Between unemployment insurance and minimum wage, we already have something like UBI, just mismanaged and with a lot of overhead.

Full-fledged UBI that provides decent living would require highly progressive taxes with the top bracket being in the ballpark of 70%. We could deal that down quite a bit if we start taxing capital gains properly, but even without that, it's neither impossible nor unprecedented.

Re: AI Data Laundering

#118

Earlier quoted context omitted.

>You can also read the HP series and write summaries and reviews about each book as wodenokoto. You can probably create HP looking artwork and write stories that could fit into the HP universe IIRC there have been lawsuits about exactly that. A person wrote (and published) some fandom in the Harry Potter universe (without Harry Potter in it IIRC), he lost the case I believe. This is similar to the fact that you canno…

Probably they used too much reference, I wasn't implying that the universe itself is not protected. But writing something similar that would appeal the fans should be okay.

https://en.wikipedia.org/wiki/Tanya_Grotter

Re: AI Data Laundering

#119
It’s definitely fair use. One question I have though is Mickey Mouse protected by copyright or trademark or ? I assume someone other than Disney can’t sell mickeys likeness or is that wrong in art? And if the AI makes a movie?

Re: AI Data Laundering

#120

Earlier quoted context omitted.

The model learns all attributes of the images it's trained on, including that some have a watermark. The fact that it generates a watermark in some images doesn't mean that that is a 1:1 image from the training set, it just means to the model some images seem to have a watermark, so it will add it sometimes. Often you can just add "no watermark" (or add it as a negative prompt with some weights) and re-use the same s…

No, it means that it is reproducing the original work and is not producing a new original work. It is basically a really fancy Instagram lense, but it is still 100% derived from the underlying works and therefor derivative instead of a newly created non-derivative work.

I'm not sure how you can make this argument just based on the model synthezing a watermark that is has learned about in the original dataset. Don't forget, the model is only 4GB in size, and while it's not out of the question that it could regurgitate an image from its data set, considering the size of the training which is a few magnitudes larger it is highly unlikely.
Post reply on HN