Live data from Hacker News

Search 5.8B images used to train popular AI art models

haveibeentrained.com

151–160 of 187 posts

Re: Search 5.8B images used to train popular AI art models

#151
post #144

Earlier quoted context omitted.

I don't think this analogy holds up.

If someone's copyrighted material was used to train a model, what analogy would you use?

You might need to explain the analogy in a bit more detail because currently I have absolutely no idea what you are getting at.

Re: Search 5.8B images used to train popular AI art models

#153
post #140
post #9

It's surprising how poorly labeled these images are; who is curating this collection? Can't they crowd-source a proper labeling project - I wonder how much better things like Stable Diffusion would be if its training would include correct, complete labels for the images. I'm sure lots of folks would willingly spend a few minutes here and there to aid with the labeling if it means they get to enjoy the model for free.

You're right. I just searched for "milk" and was not expecting the NSFW results...

I wouldn't call this poorly labeled, 'milk' and 'breasts' actually are semantically very close.

Re: Search 5.8B images used to train popular AI art models

#154
post #37

There is a lot of porn there

There really is. I searched for my first and last name, and after a few images of an actor with a vaguely similar name it devolved into shirtless buff dudes in suggestive poses.

On the plus side, i appear to have not (yet) been trained. I can consider myself safe from the AI. For now.

Re: Search 5.8B images used to train popular AI art models

#155
post #4

I'm one of the people building this. Hi HN, AMA :)

Obviously it might be a slightly futile task given the size of the dataset, but would there be value in adding community captions to some of the images? I'd be happy to spend 15 minutes captioning the worst-labelled (maybe cosine similarity between stabdiff im2txt and the prompt?) images in the set / reviewing other people's captions, and if you can get enough people on board you could probably get through a not-insignificant number of new captions. Equally a risk of anti-ai groups sabotaging this process though

Re: Search 5.8B images used to train popular AI art models

#156
post #9

It's surprising how poorly labeled these images are; who is curating this collection? Can't they crowd-source a proper labeling project - I wonder how much better things like Stable Diffusion would be if its training would include correct, complete labels for the images. I'm sure lots of folks would willingly spend a few minutes here and there to aid with the labeling if it means they get to enjoy the model for free.

"Can't they crowd-source a proper labeling project"

I'm very surprised how primitive it is. I put in a few names of uncommon technical objects and some of its images were close and others so far off as to be unrecognizable or totally unrelated/useless.

What it needs is a quick/instant feedback system that allows humans to rate an image against the query word. A rating scale of say 1 to 5 where 1 is an identical match, 2 a close match, 3 marginal, 4 ambiguous, 5 wrong/totally unrelated. If in place then likely thousands would make an effort to optimize the matching.

Re: Search 5.8B images used to train popular AI art models

#157
post #146

Earlier quoted context omitted.

labeling 6 billions pictures, even at 10 labels per second would take 20 years, so I guess nobody

Could be done by reCAPTCHA, although might be more difficult to classify the results (whether a user passed the test) compared to their usual challenges.

Not quite a CAPTCHA, but something similar has actually been done before.

The ESP game[1] paired two random people looking at the same image while a timer ticked down. The players had to enter labels that described the image, and if both players applied the same label, their score increased, and the label was associated more strongly with that image.

Well-known labels for an image were excluded after a while, so you had to guess less and less obvious labels in order to score as time went by. Once in a while the system would also assign test images with known labels to prevent cheating.

Apparently a lot of people played it because it was fun, so there wasn't even the need to pay them for labeling the dataset...

[1]: https://en.wikipedia.org/wiki/ESP_game

Re: Search 5.8B images used to train popular AI art models

#159
post #147
post #16

Earlier quoted context omitted.

This is Laion-5B, you can read more about it here: https://laion.ai/blog/laion-5b/ Imagen and Stable-Diffusion both used subsets of this full 5.8B image set.

Is Imagen actually trained on a subset of Laion-5B and nothing else? I've heard they used huge internal data sets.

They have their own datasets and included Laion-400M, a subset of 5b that was released prior to 5b. You can see a short explanation in imagen's "Limitations and Societal Impact" section at: https://imagen.research.google/.

> While a subset of our training data was filtered to removed noise and undesirable content, such as pornographic imagery and toxic language, we also utilized LAION-400M dataset which is known to contain a wide range of inappropriate content including pornographic imagery, racist slurs, and harmful social stereotypes.

Re: Search 5.8B images used to train popular AI art models

#160

Earlier quoted context omitted.

The same artists who have been pitching a fit about how the robots are coming for their jobs? Pretty sure they are trying to figure out a way their art can’t be used to train an AI and won’t be willingly participating in making it easier.

I’m skeptical that Al will replace artists. The existence of computer chess and go, though superior in every way, has not killed interest in human competition. I see no reason we couldn’t have the same thing happen with art: people see AI art as a novelty and something worth studying but otherwise preferring human art for its relevance to the human condition.

The problem is that competition in chess, in games, has always been for the sake of the game itself and the fun in playing it. 'game playing' is its own reward, there is no other entity interested in paying for results of chess games, even 'interesting' chess games. If there was, you could setup instances of Stockfish to play against each other and dominate that market.

This is not true of the artist market. While yes, many artists get into it for the fun/satisfaction of producing their own artworks, the reason that they can remain in it long term is that there are other entities in paying for the results of their work. See: game studios, film studios, etc. Before AI, the only way to get the assets for these various projects produced was by paying artists/experts for their work. Now, the production of those building blocks can be automated.

Post reply on HN