Live data from Hacker News

Search 5.8B images used to train popular AI art models

haveibeentrained.com

51–60 of 187 posts

Re: Search 5.8B images used to train popular AI art models

#51
post #48

Earlier quoted context omitted.

You mean my intellectual property that they are now charging to buy credits to rip off with a plagiarism algorithm? All rights reserved. I didn't agree for them to use my IP commercially.

What intellectual property? Where in the Stable Diffusion model weights does your intellectual property exist?

The intellectual property they train an algorithm with and charge money for. This is some real mental gymnastics.

Re: Search 5.8B images used to train popular AI art models

#53
post #52

This is great just from a prompt engineering perspective. Now I can see what labels are typically like for the types of images I'm looking for instead of guessing.

For sure! And as artists opt in, you'll be able to use it to see how they describe their work.

Re: Search 5.8B images used to train popular AI art models

#54
Excited to see semantic search getting legs with this and Lexica[1].

Related: we recently released[2] semantic search for custom datasets. You can use it to find all sorts of weird stuff in benchmark datasets like MS COCO[3] used to train many computer vision models.

If folks are interested I can write up a "how we made it" post describing the behind the scenes.

[1] https://lexica.art/

[2] https://blog.roboflow.com/dataset-search/

[3] https://blog.roboflow.com/coco-dataset-image-search/

Re: Search 5.8B images used to train popular AI art models

#55
post #4

I'm one of the people building this. Hi HN, AMA :)

How do you have the rights to use any of the images? They are clearly the same as what a Google image search would result in. However google links to the source.

If that's a rights issue, we'll definitely add a link to the source. For now, you can right click -> open in new tab to see where it came from, but we'll look into this asap.

The goal here is to give people the opportunity to remove images they don't want in this dataset or add images they do want in there.

Re: Search 5.8B images used to train popular AI art models

#57
post #43
post #4

I'm one of the people building this. Hi HN, AMA :)

hi thanks for building this! Could you also enable a simple search by matching exactly words from the caption text? rather than semantic similarity?

Thanks! That is definitely on the list, but might be a few months away. We're focusing on using images to find other images so it will be easy for artists to flag all of their stuff quickly. But, once we have that in a good place, we will definitely be adding more to the text search side.

Re: Search 5.8B images used to train popular AI art models

#58
post #9

It's surprising how poorly labeled these images are; who is curating this collection? Can't they crowd-source a proper labeling project - I wonder how much better things like Stable Diffusion would be if its training would include correct, complete labels for the images. I'm sure lots of folks would willingly spend a few minutes here and there to aid with the labeling if it means they get to enjoy the model for free.

If you aren't Google, manually doing that with 5+ billion images might prove difficult, to put it mildly. Large-scale labeling is typically bootstrapped with smaller models and whatever manual data you have. What's being curated is the bootstrapping process.

Well, imagine a Wikipedia like org/effort.

Re: Search 5.8B images used to train popular AI art models

#59
post #9

It's surprising how poorly labeled these images are; who is curating this collection? Can't they crowd-source a proper labeling project - I wonder how much better things like Stable Diffusion would be if its training would include correct, complete labels for the images. I'm sure lots of folks would willingly spend a few minutes here and there to aid with the labeling if it means they get to enjoy the model for free.

Those images are not really "labeled". They just scraped the alt text. A lot of the recent advances in AI have been by using lower quality large scale web data, instead of hand labeling. The noise will average out. Hand labeled data can be used for finetuning.

Re: Search 5.8B images used to train popular AI art models

#60

It would be a stunning twist of irony if this website uploaded images to a proprietary image dataset used for training AI models, pitching "uncorrelated data"

We don't store any images used for searching.

We are building an opt-in list, because a lot of people do want to be able to prompt AI with something like, "a cat in the style of me" or "me riding a dinosaur". That will be shared publicly, of course.

Post reply on HN