Live data from Hacker News

Search 5.8B images used to train popular AI art models

haveibeentrained.com

11–20 of 187 posts

Re: Search 5.8B images used to train popular AI art models

#11
post #3

I just uploaded a picture of my dog (Bichon Frise) and it showed me a bazillion nearly exact similar dogs. Why isn't this immediately being used as a missing persons (or pet) database service?

Let's think of reasons! 1) Privacy. 2) Lack of geolocation data associated. 3) Privacy. 4) Lack of contact information attached. 5) Privacy. 6) Lack of case numbers for various missing persons cases being attached. 7) Privacy.

Certainly this could be used for evil (tm), but it seems like something could also be built out that enables this for good in ways that don't cause issues with Mr. Fibbonaci.

Re: Search 5.8B images used to train popular AI art models

#13
post #4

I'm one of the people building this. Hi HN, AMA :)

Privacy policy for searched strings?

"Rutkowski" returns a bunch of book covers, repeated a lot. Can you ensure images returned have diverse embeddings? I expected digital art, not detective stories.

Do you use CLIP or just metadata?

What is intended process that starts after you get artist's email (for either of purposes).

Re: Search 5.8B images used to train popular AI art models

#16
post #9

It's surprising how poorly labeled these images are; who is curating this collection? Can't they crowd-source a proper labeling project - I wonder how much better things like Stable Diffusion would be if its training would include correct, complete labels for the images. I'm sure lots of folks would willingly spend a few minutes here and there to aid with the labeling if it means they get to enjoy the model for free.

This is Laion-5B, you can read more about it here: https://laion.ai/blog/laion-5b/

Imagen and Stable-Diffusion both used subsets of this full 5.8B image set.

Re: Search 5.8B images used to train popular AI art models

#17
post #8

It surprises me how many meme images there are. Aren’t they low quality content? I haven’t tried but I don’t see SD making any memes by themselves yet.

Stable Diffusion used an aesthetic filter to train on a subset of the English language images from this full 5.8 billion multi-language set. That probably got a lot of what you're finding.

Re: Search 5.8B images used to train popular AI art models

#18
I put in Donald Trump to see what kind of celebrity images might be in there, and there are a TON of memes / photoshopped versions of him looking like a caricature or otherwise warped. I wonder if the AI will average these into a fair resemblance, or whether prompts using his name will end up more cartoonish than other names due to the source data...

Re: Search 5.8B images used to train popular AI art models

#19
post #13
post #4

I'm one of the people building this. Hi HN, AMA :)

Privacy policy for searched strings? "Rutkowski" returns a bunch of book covers, repeated a lot. Can you ensure images returned have diverse embeddings? I expected digital art, not detective stories. Do you use CLIP or just metadata? What is intended process that starts after you get artist's email (for either of purposes).

It's using clip to match the text to the image, so you can actually prompt it like you might an art generator. Here's "in the style of greg rutkowski": https://haveibeentrained.com/?search_text=in%20the%20style%2...

In the next few weeks we'll be adding the ability to log in and flag or upload your works (if they aren't there). Those lists will have permissions assigned to them, starting with simple opt-in or opt-out.

Re: Search 5.8B images used to train popular AI art models

#20
post #9

It's surprising how poorly labeled these images are; who is curating this collection? Can't they crowd-source a proper labeling project - I wonder how much better things like Stable Diffusion would be if its training would include correct, complete labels for the images. I'm sure lots of folks would willingly spend a few minutes here and there to aid with the labeling if it means they get to enjoy the model for free.

If you aren't Google, manually doing that with 5+ billion images might prove difficult, to put it mildly. Large-scale labeling is typically bootstrapped with smaller models and whatever manual data you have. What's being curated is the bootstrapping process.
Post reply on HN