Live data from Hacker News

Search 5.8B images used to train popular AI art models

haveibeentrained.com

121–130 of 187 posts

Re: Search 5.8B images used to train popular AI art models

#121
post #3

I just uploaded a picture of my dog (Bichon Frise) and it showed me a bazillion nearly exact similar dogs. Why isn't this immediately being used as a missing persons (or pet) database service?

Re-identification is harder than just ranking by similarity.

Re: Search 5.8B images used to train popular AI art models

#122
post #37

There is a lot of porn there

I'll get downvoted massively, but porn is by definition art (meant to be viewed and evoke reactions, just as advertising art is), and has been some of the most important creative outputs for as almost as long as humans have been creating art.

It isn't art. But neither is decoration, nor craft.

Things can be "artistic" or contain art within them, but it doesn't make them art.

I don't buy that something done/created for the purpose of being viewed and evoking (even strong or many) reactions is enough of a criteria.

Otherwise things from trolling to 9/11 are "performance art" (an already barely hanging-on category), the latter of which was very "successful" at meeting this minimal set of rules.

Re: Search 5.8B images used to train popular AI art models

#123

Earlier quoted context omitted.

> Again, where in the neural weights does your original work live? That's not the standard for copyright infringement. The standard is (1) access to the original work and (2) producing a work that is "substantial similar" to the original. If the user of an AI produces an output that is substantially similar to one of the AI's training images, then it could potentially infringe. That's the law. In an actual case, a ju…

> producing a work that is "substantial similar" to the original. So the question “where in the model is your original?” is a reasonable and relevant question to ask. If this person can induce the model to produce a work that is reasonably considered a copy of their original, then fair enough. All they have to do is give a prompt and a seed and they can prove copyright infringement very easily because anybody else wi…

If there was an actual case, the question would be whether the defendant had access to the original work. If the original work was used to train the AI model, then the answer is yes. It's not necessary for the model to contain a copy of the original work.

Re: Search 5.8B images used to train popular AI art models

#124
post #9

It's surprising how poorly labeled these images are; who is curating this collection? Can't they crowd-source a proper labeling project - I wonder how much better things like Stable Diffusion would be if its training would include correct, complete labels for the images. I'm sure lots of folks would willingly spend a few minutes here and there to aid with the labeling if it means they get to enjoy the model for free.

I don't think that's feasible for anyone but someone like Google or another tech giant.

They could use Wikimedia Commons to get a smaller collection of better labeled images. Currently it's images from Common Crawl extracted by Laion, not sure if it already includes Wikimedia Commons.

Re: Search 5.8B images used to train popular AI art models

#125

Earlier quoted context omitted.

I'll get downvoted massively, but porn is by definition art (meant to be viewed and evoke reactions, just as advertising art is), and has been some of the most important creative outputs for as almost as long as humans have been creating art.

It isn't art. But neither is decoration, nor craft. Things can be "artistic" or contain art within them, but it doesn't make them art. I don't buy that something done/created for the purpose of being viewed and evoking (even strong or many) reactions is enough of a criteria. Otherwise things from trolling to 9/11 are "performance art" (an already barely hanging-on category), the latter of which was very "successful"…

So you know art when you see it, and there is no intersection between porn and art? Your definition is the generally and legally accepted definition? Is David art? The Rape of the Sabines?

The corpus is of art, not Art.

Re: Search 5.8B images used to train popular AI art models

#126
post #74

Earlier quoted context omitted.

If I compress an image with JPEG, I won't get the exact same pixels values in the original image after decompressing because the compression is lossy, but I'm sure you'll agree that the original image is "there". How are NN weights different from the DCT coefficients in a JPEG file?

The NN weights can't recreate anything without input values - specifically a 512x512 grid of random noise, and a transformed textual prompt. So there's almost a kilobyte of missing data, without which nothing is produced. Would a JPEG file with random noise in it be potentially any given original image? Of course not - even though the decompressor is perfectly capable of recreating one given suitable input data. The…

You can pontificate all you want about what is contained in the model, the fact is a case related to infringement is going to hit an EU court soon enough and any sort of commercialization of those models will be banned. And good fucking riddance to the AI bros, the monkeys of engineering.

Re: Search 5.8B images used to train popular AI art models

#127
post #23

Earlier quoted context omitted.

part of our idea is that artists can help with labelling their own work

The same artists who have been pitching a fit about how the robots are coming for their jobs? Pretty sure they are trying to figure out a way their art can’t be used to train an AI and won’t be willingly participating in making it easier.

I’m skeptical that Al will replace artists. The existence of computer chess and go, though superior in every way, has not killed interest in human competition. I see no reason we couldn’t have the same thing happen with art: people see AI art as a novelty and something worth studying but otherwise preferring human art for its relevance to the human condition.

Re: Search 5.8B images used to train popular AI art models

#129

It would be a stunning twist of irony if this website uploaded images to a proprietary image dataset used for training AI models, pitching "uncorrelated data"

That highlights one of my "concerns" about AI training sets. There seems to be a very real risk of accidentally adding an image that you have no rights to, so what happens if someone finds out and demand their image removed from all models trained on that image?

You can't really back an image out of a model, you can only retrain without that image.

Creative Commons should add a new license that prohibits the use of ones work to train AI.

Re: Search 5.8B images used to train popular AI art models

#130

Earlier quoted context omitted.

> producing a work that is "substantial similar" to the original. So the question “where in the model is your original?” is a reasonable and relevant question to ask. If this person can induce the model to produce a work that is reasonably considered a copy of their original, then fair enough. All they have to do is give a prompt and a seed and they can prove copyright infringement very easily because anybody else wi…

If there was an actual case, the question would be whether the defendant had access to the original work. If the original work was used to train the AI model, then the answer is yes. It's not necessary for the model to contain a copy of the original work.

That wouldn’t be the question because nobody is disputing the access to the original work. You mentioned two factors. That’s the first, which nobody is questioning. The thing people are questioning is the second factor, the reproduction of the original work.

A copy must be made for there to be copyright infringement. Who has shown a copy has been made? If a copy has been observed to occur, it’s trivial to demonstrate copyright infringement. Who has done this?

Post reply on HN