Live data from Hacker News

Search 5.8B images used to train popular AI art models

haveibeentrained.com

111–120 of 187 posts

Re: Search 5.8B images used to train popular AI art models

#111

Earlier quoted context omitted.

I'll get downvoted massively, but porn is by definition art (meant to be viewed and evoke reactions, just as advertising art is), and has been some of the most important creative outputs for as almost as long as humans have been creating art.

> but porn is by definition art No, that's why it's called pornography, very much to differentiate it from something "noble" like art, it's vulgar and trivial, by the very definition of the word it isn't art.

Etymology is "writing about prostitutes" [https://www.merriam-webster.com/dictionary/pornography]. So you got me there. In the realm of visual depictions, it's just as much art as cave art. Art is a broad classifier. Still lifes are art. Landscapes are art. Audubon's drawings of birds are art. Porn is film, photography and drawing, with a specific subject matter. If Iron Man comics are art, it is too.

"Noble" is nowhere in any working definition of art. Some academic snob might use it with their in-crowd. Damien Hirsch's shark parts is considered art.

Re: Search 5.8B images used to train popular AI art models

#112
post #75

Earlier quoted context omitted.

> meant to be viewed and evoke reactions, just as advertising art is And when they mix the two things get interesting. Having topless women selling (cow’s) milk was a very strange experience as a “prudish” American.

Sometimes I think a single change that would have an outsized improvement on American culture would be to make the mandatory high school art class include a week of sketch-drawing naked models, as actual artists do, both male and female.

Try modelling for art classes.

Re: Search 5.8B images used to train popular AI art models

#113
post #9

It's surprising how poorly labeled these images are; who is curating this collection? Can't they crowd-source a proper labeling project - I wonder how much better things like Stable Diffusion would be if its training would include correct, complete labels for the images. I'm sure lots of folks would willingly spend a few minutes here and there to aid with the labeling if it means they get to enjoy the model for free.

if there are any researchers reading this, I'd like to introduce you to booru porn galleries like rule34. millions of images meticulously labelled by dedicated enthusiasts

There are already trained models by the community on these datasets. It's called Waifu Diffusion [1] (not 100% if this is the most popular model).

[1] https://huggingface.co/hakurei/waifu-diffusion

Re: Search 5.8B images used to train popular AI art models

#114
post #63

Earlier quoted context omitted.

So you're trying to assert copyright on something which isn't a copy of something? Again, where in the neural weights does your original work live? EDIT: Let's consider a simpler scenario - I take an MD5 sum of one of your photos, and then hash collide a an image from my phone camera till I find a match. Which part of this process is "stealing"?

> Again, where in the neural weights does your original work live? That's not the standard for copyright infringement. The standard is (1) access to the original work and (2) producing a work that is "substantial similar" to the original. If the user of an AI produces an output that is substantially similar to one of the AI's training images, then it could potentially infringe. That's the law. In an actual case, a ju…

> producing a work that is "substantial similar" to the original.

So the question “where in the model is your original?” is a reasonable and relevant question to ask.

If this person can induce the model to produce a work that is reasonably considered a copy of their original, then fair enough. All they have to do is give a prompt and a seed and they can prove copyright infringement very easily because anybody else with similar hardware and software can demonstrate the infringement on demand. If I understand correctly, this has happened with GitHub Copilot, with Copilot reproducing copyrighted works verbatim.

But if they can’t do that… why should anybody take their claims of copyright infringement seriously? If nobody can point to copyright infringement having taken place, what basis is there for believing it has? As you say, the standard is access to the work and producing a work that is substantially similar to the other. The former has been demonstrated. People are asking about the latter.

So “where can we find the original in the model?” is possibly the single most relevant question there is. It’s a clear line that divides infringement from inspiration, and it can be proved definitively if copyright infringement has been observed to occur.

Re: Search 5.8B images used to train popular AI art models

#115
post #4

I'm one of the people building this. Hi HN, AMA :)

This sounds like a great initiative. I can't find anything on the site about a privacy policy, how the email addresses you're collecting will be used. If you're creating this for the AI community, then how will this data be made available to others?

Re: Search 5.8B images used to train popular AI art models

#116
post #45

Earlier quoted context omitted.

This seems no different than the controversy over GPL code being used to train Github Copilot. Just because it's publicly available and allows indexing doesn't say anything about the license it's released under.

When using GitHub, you grant GitHub a license to use your code for Copilot and other such things. This is not necessarily the same license that you might give others separately, such as MIT or GPL.

But that argument doesn't really work, since you can upload someone else's GPL'd code on GitHub. So, it's not really possible for a random person uploading old code to GitHub, to give them new rights that the original authors hadn't granted.

Re: Search 5.8B images used to train popular AI art models

#117
post #9

It's surprising how poorly labeled these images are; who is curating this collection? Can't they crowd-source a proper labeling project - I wonder how much better things like Stable Diffusion would be if its training would include correct, complete labels for the images. I'm sure lots of folks would willingly spend a few minutes here and there to aid with the labeling if it means they get to enjoy the model for free.

Could you imagine if the mediocre results we currently get from "AI" were mostly from poorly labeled data in every huge dataset and not lack of a scientific or technological breakthrough?, it wouldn't be the first time that too many academics are blinded by the wishful thinking that what they need is an eureka moment when what is needed is tons of dull repetitive work. So for instance for photography generation what…

I think for human generation you need to add a bone/skinmesh generator in 3d. So it needs to learn what the position is of the body in a 3d space and compare with the source images.

Re: Search 5.8B images used to train popular AI art models

#118
post #58

Earlier quoted context omitted.

If you aren't Google, manually doing that with 5+ billion images might prove difficult, to put it mildly. Large-scale labeling is typically bootstrapped with smaller models and whatever manual data you have. What's being curated is the bootstrapping process.

Well, imagine a Wikipedia like org/effort.

That exists for anime style images in the form of Danbooru: https://www.gwern.net/Danbooru2021

Other mediums don't have enough people interested in curating and labelling 5 million images with very detailed tag lists for free.

Re: Search 5.8B images used to train popular AI art models

#120

This NSFW! Do not open this while at work my first search was yeah let me check this totally family friendly search term and badabing badabong CTRL + W!

Yup, apparently there's a fetish for every innocent word you can imagine and this thing is completely unfiltered.
Post reply on HN