Live data from Hacker News

Stable Attribution

stableattribution.com

341–350 of 365 posts

Re: Stable Attribution

#341
post #292

Earlier quoted context omitted.

The author of this tool is even aware of that argument and just dismisses it with no real justification: https://twitter.com/atroyn/status/1622360994579357696

I do want to clarify that I think stable diffusion and tools like it can engage in illegal copying. For example it will happily produce infringing images of logos and even somewhat random other images https://arxiv.org/pdf/2212.03860.pdf . It seems like it’s devoting an uneven amount of its weights to different images, but I remain unconvinced that’s all it can do, or at least anymore all it can do than for a human a…

This is what happens when you overtrain a model too. Recent developments have allowed partial sets of model weights called LoRAs to be added to the diffusion model. These models can be fine-tuned independently in under half an hour. If you set the learning rate too high, it will start reproducing the source material with extremely high fidelity. This is what overfitting does.

My conclusion is there is an argument to be made for infringement in some cases, but it's based on degrees instead of absolutes. If infringement is defined as "copyrighted works were used in this dataset", then at a certain point (low enough learning rate) it becomes impossible to tell if infringing data was used. You'd be working with weight amounts that are so miniscule they could be rounding errors, yet by that definition would still be infringing.

And since any arbitrary data can be used with some set of keywords, the standard for what constitutes "infringing" changes with each model. As in, it would probably be hard to have a benchmark test that can definitively state "this model violates copyright." Any number of keywords can be trained on to obfuscate the prompt needed to reproduce the data, assuming there was even a high enough LR for the data to be reproduced similarly enough.

I'm unsure if there can ever be one standard for when a set of a bunch of floating point numbers can pass the threshold for constituting infringement. This is applying an absolute standard to a fuzzy algorithm. It's like compressing a JPEG, at some level of compression on the scale a picture of Mickey Mouse becomes unintelligible. But with JPEGs it isn't really useful to have an unintelligible picture of Mickey Mouse. However, it can be extremely useful to have a LoRA with the weights underfit just enough to where the diffusion gives novel outputs.

Re: Stable Attribution

#342
post #282

Earlier quoted context omitted.

> Art is made with physical motor skills, time and effort Apparently it isn't. Now a computer can do it. > that requires no effort of aquiring the physical skill needed to make the art. So what? This doesn't matter. > If you want to physically copy my art, go for it This is the same thing as if a computer does it.

> Apparently it isn't. Now a computer can do it. Why do we need farmers when you can just buy a burger at McDonald's?

False equivalence.

My AI art generator works right now, on my PC. If everyone stopped making art, the "normal" way, my AI art generator would still work.

This is a false equivalence to your example, because if farmers stopped farming, then it would not be possible to make burgers.

But, thats usually what happens when someone comes up with a pithy, one off meme response, like you just did, instead of actually responding to the substance of the argument. (I expect any future response from you, to instead not be engaging with the substance, and instead coming up with reasons why the silly analogy still works)

Re: Stable Attribution

#343

Earlier quoted context omitted.

> Also 'trained' is a complete misnomer. This here shows that you are speaking without understanding what you are speaking about. AI is absolutely trained. It's a process that is quite literally inspired by the way we understood neurons to work in the 1970s. AI start with a big batch of random numbers. There's a big fancy scientific method used to adjust those numbers in order to cause the system to learn to do some…

I personally care very little for copyright or copycats. I have great disdain for those who would profit off the backs of the labor of others. Saying that art often involves little to no effort just shows how ignorant you are of the subject. I am not threatened by AI and frankly I don't see it as competition, it's not really that good. Besides that the majority of my art is three dimensional. I am just saddened by ho…

> I personally care very little for copyright or copycats. I have great disdain for those who would profit off the backs of the labor of others.

This is actually an anti-capitalist argument, rather than an anti-AI one.

That doesn't weaken the argument, IMO, but it does help to know whose ox you're trying to gore.

Re: Stable Attribution

#344
post #334

Earlier quoted context omitted.

to me it's less about the process and jobs lost and more about some feeling loss related to the excitement about removing expression and experience from a domain for honestly very little real benefit continuing to transform things which in part centered around exploration, discovery and experimentation into something cold and kind of dumb like we will make some interesting things but it's the general trend of modern…

On the other hand, paints and canvases are very very expensive, making art a domain of the rich. You can use the ai tools at the library, making art more proletarian. I think the pushback is around status and elitism, and that people with certain backgrounds are societally expected to be not making art.

I picked up some nice canvases recently at the dollar tree. Masonite, wood and paper can be painted on. Masonite is actually preferable for acryllic paints and it is quite affordable, as are acryllic paints themselves. You can paint with coffee grinds and beet juice...the exploration is endless. I come from a family of seven living in the backwoods and most of my artist friends would not nearly qualify as rich. I had a teacher once who made the most beautiful art out of entirely recycled metal junk. Art is a reflection of culture, top to bottom and ingenuity plays a big role. Ingenuity is accessible to everybody.

Re: Stable Attribution

#345

Earlier quoted context omitted.

> Based on some of the examples they explicitly provided, it is clear to me Stable Diffusion creates novel art. You're jumping to a conclusion that the data doesn't warrant, I suspect because it's a conclusion that suits you. They're doing something like a "reverse image search" from an AI generated image over the original dataset and returning a few examples that have a high degree of similarity. There's no guarante…

Until somebody comes out with a better way to trace back output images to training data, this is the best "data" we have so far. Before this, there were anecdotal examples of SD outputting an image with some Shutterstock watermark or very similar to some artist's work, but the prompts also seemed highly specific or were asking for something in that artist's style. This tool at least lets us start to trace back the av…

> Until somebody comes out with a better way to trace back output images to training data, this is the best "data" we have so far.

There are many better ways, in the sense that they actually do something like estimate the causal effect of a specific training datapoint, like leave-one-out cross-validation training or surrogates to Shapley value, or using nonparametric models to trace backwards. This is a whole subfield of ML research.* (The primary summary is: "it's hard and the easy approaches don't work." Which why he's not doing any of those but an easy incorrect thing.)

* I'm not entirely sure why anyone cared so much... Research topics can be kinda arbitrary. But in this case, I think there was something of a fad around 2017 that there were going to be 'data marketplaces' where you would be trying to estimate the value of each datapoint to price it. This turned out to not exist as a business model: you either used big public data for free for generic model capabilities, or you had small proprietary data you'd die rather than sell to a competitor.

Re: Stable Attribution

#346
post #45

Uploading works by real human artists gives you a batch of results that resemble a reference board (mood board, inspiration board, etc) the artist could have been looking at while creating their original work of art. Obviously it’s not the actual reference board, the only way to get that is to ask the artist yourself, but it sure looks like what you’d expect their reference board to look like. This site is grift, of…

The AI does NOT build or use reference boards for specific prompts. The only reference it has is the prompt itself, which gets distilled down into a list of 512 numbers, each one of which the AI associates with a particular image feature (or set of features). The only reference material it has is the training set. Most images are not actually retained in the model; but certain statistically significant or repeated im…

> The AI does NOT build or use reference boards for specific prompts.

Interestingly, even if it did build reference boards, that doesn't mean that the generated image would 'copy' or 'look like' the references all that much. We know this because there are diffusion generative models which do exactly that: they do retrievals first, and include the 'examples' along with the prompt before generating. Here is a visualization: https://arxiv.org/pdf/2204.02849.pdf#page=19 (the generated sample on the left is clearly semantically correct and high-quality but also very different from all the exemplars it was generated using).

Re: Stable Attribution

#348

Earlier quoted context omitted.

So it is the opposite of attribution. It just makes up a plausible-looking story of attribution and passes it off as the truth. If you passed it a hand-painted image from 1850 which was not in SD's training dataset, it can happily declare that it was inspired by someone's piece from 2019. There's no actual causal inference going on. It has more in common with using an AI language model as a "bullshit generator" than…

Why do you close the space for people who are pro-ai AND pro-attribution? The entire ai space wreaks of this divisiveness, and is likely why it will continue to die out as another "art-fad". There is seemingly little willingness to integrate into the existing art world in good faith.

I close the space for liars, whether they are pro or anti, and this service lies about attribution. Lying about attribution will just make everything worse.

Re: Stable Attribution

#349
post #104

Earlier quoted context omitted.

The slight adjustment of where I'm looking is minor. The instrument is uncommon and honestly difficult to get stable diffusion to generate correctly at all; though, from my experience playing with this (I spent a lot of time trying to figure out why it knew who I was before discovering the CLIP database browser), I'm going to argue that the reason that hand is showing up in the position it is is because of my hand ho…

> Clearly seeing at least one photo of me (and AFAIK it was only trained on thousands of copies of that single photo) was absolutely crucial to the construction of this image, and yet this website isn't finding any I think there are two separate ideas that are being conflated here. The first idea is that there is a mapping between text input and a joint text/image embedding space. For that mapping, yes, your profile…

I think a key question in attribution is whether the model would have been able to generate the same result without access to the input, and then how much it would have lost having been restricted from that input. If you remove from the mechanism all of the copies of my profile picture (and there are a lot of them...), I guess I am willing to believe that it might still have enough text descriptions of saurik to come up with sort of what I look like, but I doubt it? On the other side, if you remove any random handful of these profile pictures of fat hairy nerds (aka, people who look a bit like me) I doubt it really needed all of them to figure out what it needed to know.

To take your example: let's say its knowledge that "throwaway1851" looks like Ben Affleck comes from your one profile picture--the only time the multi-modal embedding model was ever able to associate that word and Ben Affleck's photo into a similar location in the vector space--then if it wasn't allowed to see that photo during its training there is no way that would have happened. It simply doesn't matter if it seems to mostly rely on its knowledge of Ben Affleck to conjure up a photo of you: it has so many photos of Ben Affleck that none of them really matter anymore, but that single photo of you that made it even realize you looked like Ben Affleck in the first place is absolutely critical and probably deserves most of the attribution.

Re: Stable Attribution

#350

You ever feel like this specific propaganda war is actually unwinnable? Many people are extremely motivated to bullshit the public (usually sincerely though I kind of doubt it in this case), and from I've seen, the public are far more willing to believe the 3 extremely online artists who they've heard an opinion on the topic from than the 1 software engineer/data scientist who actually knows half a thing about machin…

It is a plagiarism machine - software engineer with years of ML experience.
Post reply on HN