Live data from Hacker News

Search 5.8B images used to train popular AI art models

haveibeentrained.com

101–110 of 187 posts

Re: Search 5.8B images used to train popular AI art models

#102
post #96

Earlier quoted context omitted.

Certainly this could be used for evil (tm), but it seems like something could also be built out that enables this for good in ways that don't cause issues with Mr. Fibbonaci.

Is that the Prison Break reference I think it is? :D

Lol! I wish I was that good at references.

I was thinking more that 1,2,3,5,7 were all privacy related.

I also have fib burned into my head from pivotal tracker story points...

Re: Search 5.8B images used to train popular AI art models

#103

One thing that sticks out to me is how so much of the the images in the collection have really terrible labels. I uncovered a large collection of pieces by an illustrator who was unsearchable by name, only via image upload. The reason they were unsearchable: the majority of this particular artist's images had labels that were all in the format of: {username}'s profile image

Welcome to the real world of data.

Re: Search 5.8B images used to train popular AI art models

#104
post #35

Earlier quoted context omitted.

Please do sign up and you'll be able to flag these images soon. We'll work to get them removed from this and future datasets built for AI training.

Why did you include them in the 1st place? What legal right did you have to do that?

We (Spawning) did not create the dataset or train the models in question. We're working to make it easy for people to remove themselves from, or add themselves to, this dataset and future models.

Re: Search 5.8B images used to train popular AI art models

#105
post #9

It's surprising how poorly labeled these images are; who is curating this collection? Can't they crowd-source a proper labeling project - I wonder how much better things like Stable Diffusion would be if its training would include correct, complete labels for the images. I'm sure lots of folks would willingly spend a few minutes here and there to aid with the labeling if it means they get to enjoy the model for free.

Could you imagine if the mediocre results we currently get from "AI" were mostly from poorly labeled data in every huge dataset and not lack of a scientific or technological breakthrough?, it wouldn't be the first time that too many academics are blinded by the wishful thinking that what they need is an eureka moment when what is needed is tons of dull repetitive work.

So for instance for photography generation what may be needed is huge amount of clean photos with obsessively detailed labels, maybe just the exact same single-point-lighting (the exact coordinates of the light being a data point, plus strength/lumens), with 8 pictures of each subject in black background: front, back, left, right, top, bottom, 3/4 mostly-front (AKA the corner), 3/4 mostly-back, and then the same 8 ones but with white background, then also include all the info possible: weight, height, width and depth, plus versions of the most common states of each object (ball: inflated or deflated; bird: flying, idle or walking), plus photos adding two subjects together (one dataset of woman with hat, another wearing the same clothes but without the hat, and one of just the hat without the woman), with properly labeled relationships so it's clear they all refer to the exact same hat (or lack of), you get the idea...

Then the most important thing for "image generation AI" may not be computing power but the most boredom-resiliant staff you can hire; of course that's just an example for photography, for things like text you may need an equivalent rigorous effort by a multitude of linguists.

Re: Search 5.8B images used to train popular AI art models

#106
post #55

Earlier quoted context omitted.

If that's a rights issue, we'll definitely add a link to the source. For now, you can right click -> open in new tab to see where it came from, but we'll look into this asap. The goal here is to give people the opportunity to remove images they don't want in this dataset or add images they do want in there.

That's not how this works. Copyright owners have the right to control when, where, how and by whom their content may be used. Not you https://www.law.cornell.edu/uscode/text/17/106

Isn’t this fair use? Copyright holders don’t get to opt out of fair use, correct?

Re: Search 5.8B images used to train popular AI art models

#107
post #63

Earlier quoted context omitted.

So you're trying to assert copyright on something which isn't a copy of something? Again, where in the neural weights does your original work live? EDIT: Let's consider a simpler scenario - I take an MD5 sum of one of your photos, and then hash collide a an image from my phone camera till I find a match. Which part of this process is "stealing"?

Copyright owners have the right to control when, where, how and by whom their content may be used. https://www.law.cornell.edu/uscode/text/17/106

This is very much not what your link says - quoting:

---

Subject to sections 107 through 122, the owner of copyright under this title has the exclusive rights to do and to authorize any of the following:

(1) to reproduce the copyrighted work in copies or phonorecords;

(2) to prepare derivative works based upon the copyrighted work;

(3) to distribute copies or phonorecords of the copyrighted work to the public by sale or other transfer of ownership, or by rental, lease, or lending;

(4) in the case of literary, musical, dramatic, and choreographic works, pantomimes, and motion pictures and other audiovisual works, to perform the copyrighted work publicly;

(5) in the case of literary, musical, dramatic, and choreographic works, pantomimes, and pictorial, graphic, or sculptural works, including the individual images of a motion picture or other audiovisual work, to display the copyrighted work publicly; and

(6) in the case of sound recordings, to perform the copyrighted work publicly by means of a digital audio transmission.

----

This is far from an unlimited grant of power. Of this list, the only plausible grounds is (2) - derivative works.

Which means we're well into then arguing about "Fair Use"[1], which I would encourage people to read the full description of carefully - because the answer isn't whether you can come up with a snippy "gotcha!" it's whether under careful consideration in the court of law anyone would be likely to agree with you.

[1] https://www.copyright.gov/fair-use/

Re: Search 5.8B images used to train popular AI art models

#108
post #9

It's surprising how poorly labeled these images are; who is curating this collection? Can't they crowd-source a proper labeling project - I wonder how much better things like Stable Diffusion would be if its training would include correct, complete labels for the images. I'm sure lots of folks would willingly spend a few minutes here and there to aid with the labeling if it means they get to enjoy the model for free.

Google has been doing this for years with their various iterations of captchas.

Re: Search 5.8B images used to train popular AI art models

#110
post #9

It's surprising how poorly labeled these images are; who is curating this collection? Can't they crowd-source a proper labeling project - I wonder how much better things like Stable Diffusion would be if its training would include correct, complete labels for the images. I'm sure lots of folks would willingly spend a few minutes here and there to aid with the labeling if it means they get to enjoy the model for free.

So stable Diffusion has a img2prompt mode. I wonder if that can be used somehow. The prompts it has yielded for my personal images have been very descriptive and good. It would be interesting to see how different the img2prompt output is for the training images. I would love to measure it but I don't even know how to calculate this distance.
Post reply on HN