Live data from Hacker News

No elephants: Breakthroughs in image generation

oneusefulthing.org

171–180 of 373 posts

Re: No elephants: Breakthroughs in image generation

#171
post #31

Earlier quoted context omitted.

That blog post is a year old. There has been a lot of progress since then: https://doubiiu.github.io/projects/ToonCrafter/

Very impressive. This is going to result in an explosion of content creation by pro studios, just as CG with cel-shading renderers did. I greatly prefer the hand-drawn + AI tweened look to the current low-budget CG 3D models look.

Most of the professionals in this industry actively despise this technology.

Re: No elephants: Breakthroughs in image generation

#172

I am waiting for when I could provide these a scene snippet from "Hitchhiker's Guide To Galaxy" (or any book) and it could draw that for me. The gold planets, the waking up on the beach, total perspective vortex etc. I like the book, but there are quite a few scenes which are quite hard to visualize and make sense. An image generator that can follow that language and detail will be amazing. Even more awesome will be…

I've seen stuff that echoes this sentiment before, and I have to say I don't understand this desire at all. Why would I need a computer to show me what something in a book looks like? I already have an imagination for that!

Books are fundamentally a collaborative artform between the author and the reader. The author provides the blueprint, but it's up to the reader to construct the scene in their own head. And every reader is going to have slightly different interpretations based on how they imagine the events of a book. This act of imagination and re-interpertation is one of the things I love about reading books.

Having a computer do the visualization for you completely destroys what makes books engaging and interesting. If you don't want to visualize the book yourself, I have to wonder why the hell you're reading a book in the first place.

If you need that visual component, just watch a movie or read a comic book or something. This isn't a slight against movies or comics! They're fantastic mediums and are able to utilize their visual element to communicate ideas in ways that books can struggle with. And these visuals will form a much more cohesive artistic vision then whatever an AI outputs, since they're an integrated and intentional part of the work.

Re: No elephants: Breakthroughs in image generation

#173

Earlier quoted context omitted.

But it’s not copying and reselling — it’s imitation. Copying is controlled by copyrights. And imitation isn’t controlled by anything. As for a company: a company is just a group of people acting together.

#1, it’s extremely easy to coerce direct copies out of models that artists could be pursued for infringement if they drew, but companies reselling said copyrighted artwork face no penalty #2, yes, it’s a group of people who came together to build an algorithm that learns to extract features learned from images made by other people in order to generate images somewhere between these images in a high dimensional space.…

I guess I don't understand what you mean when you say companies reselling said copyrighted artwork face no penalty. Why wouldn't they? If I was to make a copy of a Studio Ghibli movie and sell it, I would absolutely face a penalty if I was caught.

Re: No elephants: Breakthroughs in image generation

#174

There is circumstantial evidence out there that 4o image manipulation isn't done within the 4o image generator in one shot but is a workflow done by an agentic system. Meaning this, user inputs prompt "create an image with no elephants in the room" > prompt goes to an llm which preprocesses the human prompt > outputs a a prompt that it knows works withing this image generator well > create an image of a room > and th…

The prompt enrichment thing is pretty standard. Everyone does that bit, though some make it user-visible. On Grok it used to populate to the frontend via the download name on the image. The image editing is interesting.

All the stable diffusion software I've used names the files after some form of the prompt, and probably because SD weights the first tokens higher than the last tokens, probably as a side effect of the way the CLIP/BLIP works.

I doubt any of these companies have rolled their own interface to stable diffusion / transformers. It's copy and paste from huggingface all the way down.

I'm still waiting for a confirmed Diffusion Language Model to be released as gguf that works with llama.cpp

Re: No elephants: Breakthroughs in image generation

#175
post #170

Earlier quoted context omitted.

I love using LLMs to generate pictures. I'd call myself rather creative, but absolutely useless in any artistic craft. Now, I can just describe any image I can imagine and get 90% accurate results, which is good enough for the presentations I hold, online pet projects (created a squirrel-themed online math-learning game for which I previously would have needed a designer to create squirrel highschool themed imagery)…

If you use this technology, you're actively harming creative labor.

Creative labor is not entitled to the work parent comment is describing. We employ labor because it is beneficial to us, not merely because it exists as an option. Creative labor’s responsibility is to adapt to a changing world and find roles where their labor is not simply produced / exceeded by a computer system.

Practically speaking, the work described would most likely never have been done, rather than been done by an artist if that were the only option - it’s uncommon to employ artists to help with incidental tasks relative to side projects, etc.

Re: No elephants: Breakthroughs in image generation

#176
post #33

> Is it okay to reproduce the hard-won style of other artists using AI? Who owns the resulting art? Who profits from it? Which artists are in the training data for AI, and what is the legal and ethical status of using copyrighted work for training? These were important questions before multimodal AI, but now developing answers to them is increasingly urgent. I have to disagree with the conclusion. This was an importa…

I don't think there's consensus around that idea. Lots of people (myself included) feel that copyright is already vastly overreaching, and that AI represents forward progress for the proliferation of art in society (its crap today, but digital cameras were crap in 2007 and look where they are now). Its also not clear for example that Studio Ghibli lost by having their art style plastered all over the internet. I went…

Nearly every artist I've spoken to or have seen talk about this technology says it's evil, so at least among the victims of this corporate abuse of the creative community, there's wide consensus that it's bad.

> but I certainly don't think that AI without restrictions is going to lead to fewer people with art jobs.

It's great that you think that but in reality a lot of artists are saying they're getting less work these days. Maybe that's the result of a shitty economy but I find it very difficult to believe this technology isn't actively stealing work from people.

Re: No elephants: Breakthroughs in image generation

#177
post #3

This is a before/after moment for image generation. A simple example is the background images on a ton of (mediocre) music youtube channels. They almost all use AI generated images that are full of nonsense the closer you look. Jazz channels will feature coffee shops with garbled text on the menu and furniture blending together. I bet all of that disappears over the next few months. On another note, and perhaps other…

I have the Gemini app on my phone and you can interact with it with voice only and I was like oh this is really cool I can use it while I'm driving instead of listening to music.

I can never think of anything to talk to an AI about. I run LM local, as well

Re: No elephants: Breakthroughs in image generation

#178

Earlier quoted context omitted.

Right the difference is that it’s a large company looking at it then copying it and reselling it without credit, which basically everyone would understand as bad without the indirection of a model. Edit: the key words here are “company” and “reselling”

But it’s not copying and reselling — it’s imitation. Copying is controlled by copyrights. And imitation isn’t controlled by anything. As for a company: a company is just a group of people acting together.

You need to copy the work to use it for AI training.

Re: No elephants: Breakthroughs in image generation

#179
post #170

Earlier quoted context omitted.

I love using LLMs to generate pictures. I'd call myself rather creative, but absolutely useless in any artistic craft. Now, I can just describe any image I can imagine and get 90% accurate results, which is good enough for the presentations I hold, online pet projects (created a squirrel-themed online math-learning game for which I previously would have needed a designer to create squirrel highschool themed imagery)…

If you use this technology, you're actively harming creative labor.

Whatever. I wrote and co-wrote ten albums and my total take was $3.

The market is saturated and the way it works means ten get rich for every million artists. I feel as though this has been pretty constant throughout history.

Of course there's a lot of talent out there, "wasted", but I think that's always been the case. How many William Shakesmans did we lose with all the war, famine, disease?

I actually decided I'd probably never write music again after 1-shot making a song about the south Korea coup attempt several months ago. I had the song done before the news really even hit the US. Why would I destroy my own hearing writing music anymore when I can prompt an AI to do it for me, with the same net result - no one cares.

here's the 3-shot remix, the triangle cracks me up so much that i had to upload it https://soundcloud.com/djoutcold/coup-detat-symphony-remix

the "original" "1-shot" is on my soundcloud page as well. https://soundcloud.com/djoutcold/i-aint-even-writing-music-a...

it's in lojban. That's why you can't understand it. Yes. Lojban. Brings a tear to my eye every time i hear it. fkin AI

[0] more my style - hold music for our PBX https://soundcloud.com/djoutcold/bew-hold-music also all my stuff is CC licensed, mostly CC0 at this point.

Re: No elephants: Breakthroughs in image generation

#180

> Is it okay to reproduce the hard-won style of other artists using AI? Who owns the resulting art? Who profits from it? Which artists are in the training data for AI, and what is the legal and ethical status of using copyrighted work for training? These were important questions before multimodal AI, but now developing answers to them is increasingly urgent. I have to disagree with the conclusion. This was an importa…

“Fair” doesn’t matter. The only consensus that matters is what is legal and profitable. The former seems to be pretty much decided in favor of AI, with some open question about whether large media companies enjoy protections that smaller artists don’t. (The legal battle when some AI company finally decides to let their model imitate Disney stuff is going to be epic.) Profitable remains to be seen, but doesn’t matter…

> The former seems to be pretty much decided in favor of AI

None of the cases against AI companies have been decided afaik. There's a ton of ongoing litigation.

> but doesn’t matter much while investors’ money is so plentiful.

More and more people are realizing how wasteful this stuff is every day.

Post reply on HN