Live data from Hacker News

No elephants: Breakthroughs in image generation

oneusefulthing.org

211–220 of 373 posts

Re: No elephants: Breakthroughs in image generation

#211
post #170

Earlier quoted context omitted.

I love using LLMs to generate pictures. I'd call myself rather creative, but absolutely useless in any artistic craft. Now, I can just describe any image I can imagine and get 90% accurate results, which is good enough for the presentations I hold, online pet projects (created a squirrel-themed online math-learning game for which I previously would have needed a designer to create squirrel highschool themed imagery)…

If you use this technology, you're actively harming creative labor.

All labor is bad.

Re: No elephants: Breakthroughs in image generation

#212

Earlier quoted context omitted.

I love using LLMs to generate pictures. I'd call myself rather creative, but absolutely useless in any artistic craft. Now, I can just describe any image I can imagine and get 90% accurate results, which is good enough for the presentations I hold, online pet projects (created a squirrel-themed online math-learning game for which I previously would have needed a designer to create squirrel highschool themed imagery)…

My problem with finding enjoyment in this is the same problem I have when using cheat codes in games: the doing part is the fun part, getting to the end or just permutations of the end gets really boring.

Trying to draw a squirrel when you have no artistic talents or experience is not the fun part.

I've produced my own music recordings in the past and I've hired musicians to play the instruments that I cannot. Having exasperated recording engineers watch my 5,000th take on a drum fill that I absolutely cannot play is not the fun part. Sitting behind the glass and watching my vision come to life from a really good drummer is absolutely the fun part.

Re: No elephants: Breakthroughs in image generation

#213

Earlier quoted context omitted.

It took a truly colossal amount of human time and effort to build AI systems. It takes significant amount of energy to run those AI systems. I don’t see any meaningful difference at all between the system of a human, a computer and a corpus of images producing new images, and the system of a human, a paintbrush, an easel, a canvas and a corpus of images producing new images. Emphasis on the new — copying is still cop…

>It took a truly colossal amount of human time and effort to build AI systems. It takes significant amount of energy to run those AI systems. Those people and effort aren't at all tied to the people who are making and using the art. In the past every individual person would have to individually study art and some style and practice for years of their life to be able to replicate it really well. And for each piece of…

I do wonder what the outcome would be for a model trained only on truly non copyright work, and derivatives from there. I'm no AI expert, but from what I understand they use some models to generate data with which to train further models. I'd be interested in the output, whether it would eventually just match what we have now anyway, so the copyright question may end up moot. I wonder how the argument would shift at that point?

I think in reality, it is probably too late for that, because the internet is now polluted with AI generated images which would be consumed by any "ethical" model anyway.

Re: No elephants: Breakthroughs in image generation

#214

Earlier quoted context omitted.

I think, Studio Ghibli will be affected, as well, since their "trademark style" (as we used to say), formerly a welcome sight and indicative for a certain type of story telling, will be devaluated as an indicator for slop. (Much like there are certain traits of an image, which we associate with soap operas and assume to be indicative of a low-value production.)

I doubt that. "Which movie does the slop belong to?" "Oh none of them? Ok" Is a pretty easy search term

I doubt that, when confronted with an image that you've learned to associate with a plethora of low-quality / low-effort productions, you'd search for the possible origin, in the first place.

(After all, it's yet another ephemeral image in "that AI style", with no apparent thought having gone into it, just some name dropping, at best. Or some generated, senseless story, you would be glad, the algorithm hadn't pointed your kids at. Why should you?)

Re: No elephants: Breakthroughs in image generation

#215
post #3

This is a before/after moment for image generation. A simple example is the background images on a ton of (mediocre) music youtube channels. They almost all use AI generated images that are full of nonsense the closer you look. Jazz channels will feature coffee shops with garbled text on the menu and furniture blending together. I bet all of that disappears over the next few months. On another note, and perhaps other…

I love using LLMs to generate pictures. I'd call myself rather creative, but absolutely useless in any artistic craft. Now, I can just describe any image I can imagine and get 90% accurate results, which is good enough for the presentations I hold, online pet projects (created a squirrel-themed online math-learning game for which I previously would have needed a designer to create squirrel highschool themed imagery)…

> For many, many websites this is going to be good enough.

It was largely a solved problem though. Companies did not seem to have an issue with using stock photos. My current company's website is full of them.

For business use cases, those galleries were already so extensive before AI image generation, that what you wanted was almost always there. They seemingly looked at people's search queries, and added images to match previously failed queries. Even things you wouldn't think would have a photo like "man in business suit jump kicking a guy while screaming", have plenty of results.

Re: No elephants: Breakthroughs in image generation

#216
post #68

Earlier quoted context omitted.

That is a very interesting point about how little use of AI most of us making day to day, despite the potential utility that seems to be lurking. I think it just takes time for people and economies to adapt to new technology. Even if technological progress on AI were to stop today, and the best models that exist in 2030 are the same models we have now, there would still be years of social and economic change as peopl…

Unless I'm doing something simple like writing out some basic shell script or python program, it's often easier to just do something myself than take the time to explain what I want to an LLM. There's something to be said about taking the time to formulate your plan in clear steps ahead of time, but for many problems it just doesn't feel like it's worth the time to write it all out.

I find that if a problem doesn't require planning it's probably simple enough that the LLM can handle it with little input. if it does require planning, I might as well dump it into an LLM as another evaluator and then to drive the implementation.

Re: No elephants: Breakthroughs in image generation

#217
post #170

Earlier quoted context omitted.

If you use this technology, you're actively harming creative labor.

All labor is bad.

Interesting philosophy, what is this predicated on? Do you mean that people should not have to work for a living, ie labor versus play?

Re: No elephants: Breakthroughs in image generation

#218
post #171

Earlier quoted context omitted.

Very impressive. This is going to result in an explosion of content creation by pro studios, just as CG with cel-shading renderers did. I greatly prefer the hand-drawn + AI tweened look to the current low-budget CG 3D models look.

Most of the professionals in this industry actively despise this technology.

That's true, but it will likely be newer studios with younger professionals who are going to be using it, much as Miyazaki doesn't like CGI either yet it's widely used now in anime. The young drive the advances while the older eschew them, that's generally how human progress has been.

Re: No elephants: Breakthroughs in image generation

#219
post #3

This is a before/after moment for image generation. A simple example is the background images on a ton of (mediocre) music youtube channels. They almost all use AI generated images that are full of nonsense the closer you look. Jazz channels will feature coffee shops with garbled text on the menu and furniture blending together. I bet all of that disappears over the next few months. On another note, and perhaps other…

> They almost all use AI generated images that are full of nonsense the closer you look. Jazz channels will feature coffee shops with garbled text on the menu and furniture blending together. Noticed tht. Maybe it's my algorithm but YouTube is seemingly filled with these videos now.

Probably is your algorithm as mine is pretty good in not showing me those low effort channels. Check out extensions like PocketTube, SponsorBlock, and DeArrow to manage your YouTube feeds better.

Re: No elephants: Breakthroughs in image generation

#220

Earlier quoted context omitted.

But it’s not copying and reselling — it’s imitation. Copying is controlled by copyrights. And imitation isn’t controlled by anything. As for a company: a company is just a group of people acting together.

#1, it’s extremely easy to coerce direct copies out of models that artists could be pursued for infringement if they drew, but companies reselling said copyrighted artwork face no penalty #2, yes, it’s a group of people who came together to build an algorithm that learns to extract features learned from images made by other people in order to generate images somewhere between these images in a high dimensional space.…

> it’s extremely easy to coerce direct copies out of models that artists could be pursued for infringement if they drew, but companies reselling said copyrighted artwork face no penalty

The purpose of the model isn't to make exact reproductions. It's like saying you can use the internet for copyright infringement. You can, but it's the user who chooses the use, so is that on AT&T and Microsoft or is it on the users doing the infringement?

> They sell these images and give no credit or cash to the images being “interpolated” between.

A big part of the problem is that machines aren't qualified to be judges.

Suppose the image you request is Gollum but instead of the One Ring he wants PewDiePie. Obviously this is using a character from the LOTR films by Warner Bros. If you're PewDiePie and you want this image to use in an ad for your channel, you might be in trouble.

But Warner Bros. got into a scandal for paying YouTubers to promote Promote Middle Earth: Shadow of Mordor without disclosing the payments. If you're creating the image to criticize the company's behavior, it's likely fair use.

The service has no way to tell why you want the image, so what is it supposed to do? A law that requires them to deny you in the second case is restricting a right of the public. But it's the same image.

Meanwhile in the first case you don't really need the company generating the image to do anything because Warner Bros. could then go after PewDiePie for using the character in commercial advertising without permission.

> Notice this doesn’t extend to open source, it’s the commercial aspect that represents theft.

It's also not really clear how this works. For example, Stable Diffusion is published. You can run it locally. If you buy a GPU from Nvidia or AMD in order to do that, is that now commercial use? Is the GPU manufacturer in trouble? What if you pay a cloud provider like AWS to use one of their GPUs to do it? You can also pay for the cloud service from Stability AI, the makers of Stable Diffusion. Is it different in that case than the others? How?

Post reply on HN