Earlier quoted context omitted.
That blog post is a year old. There has been a lot of progress since then: https://doubiiu.github.io/projects/ToonCrafter/
Very impressive. This is going to result in an explosion of content creation by pro studios, just as CG with cel-shading renderers did. I greatly prefer the hand-drawn + AI tweened look to the current low-budget CG 3D models look.
No elephants: Breakthroughs in image generation
171–180 of 373 posts
Re: No elephants: Breakthroughs in image generation
#172I am waiting for when I could provide these a scene snippet from "Hitchhiker's Guide To Galaxy" (or any book) and it could draw that for me. The gold planets, the waking up on the beach, total perspective vortex etc. I like the book, but there are quite a few scenes which are quite hard to visualize and make sense. An image generator that can follow that language and detail will be amazing. Even more awesome will be…
Books are fundamentally a collaborative artform between the author and the reader. The author provides the blueprint, but it's up to the reader to construct the scene in their own head. And every reader is going to have slightly different interpretations based on how they imagine the events of a book. This act of imagination and re-interpertation is one of the things I love about reading books.
Having a computer do the visualization for you completely destroys what makes books engaging and interesting. If you don't want to visualize the book yourself, I have to wonder why the hell you're reading a book in the first place.
If you need that visual component, just watch a movie or read a comic book or something. This isn't a slight against movies or comics! They're fantastic mediums and are able to utilize their visual element to communicate ideas in ways that books can struggle with. And these visuals will form a much more cohesive artistic vision then whatever an AI outputs, since they're an integrated and intentional part of the work.
Re: No elephants: Breakthroughs in image generation
#173Earlier quoted context omitted.
But it’s not copying and reselling — it’s imitation. Copying is controlled by copyrights. And imitation isn’t controlled by anything. As for a company: a company is just a group of people acting together.
#1, it’s extremely easy to coerce direct copies out of models that artists could be pursued for infringement if they drew, but companies reselling said copyrighted artwork face no penalty #2, yes, it’s a group of people who came together to build an algorithm that learns to extract features learned from images made by other people in order to generate images somewhere between these images in a high dimensional space.…
Re: No elephants: Breakthroughs in image generation
#174There is circumstantial evidence out there that 4o image manipulation isn't done within the 4o image generator in one shot but is a workflow done by an agentic system. Meaning this, user inputs prompt "create an image with no elephants in the room" > prompt goes to an llm which preprocesses the human prompt > outputs a a prompt that it knows works withing this image generator well > create an image of a room > and th…
The prompt enrichment thing is pretty standard. Everyone does that bit, though some make it user-visible. On Grok it used to populate to the frontend via the download name on the image. The image editing is interesting.
I doubt any of these companies have rolled their own interface to stable diffusion / transformers. It's copy and paste from huggingface all the way down.
I'm still waiting for a confirmed Diffusion Language Model to be released as gguf that works with llama.cpp
Re: No elephants: Breakthroughs in image generation
#175Earlier quoted context omitted.
I love using LLMs to generate pictures. I'd call myself rather creative, but absolutely useless in any artistic craft. Now, I can just describe any image I can imagine and get 90% accurate results, which is good enough for the presentations I hold, online pet projects (created a squirrel-themed online math-learning game for which I previously would have needed a designer to create squirrel highschool themed imagery)…
If you use this technology, you're actively harming creative labor.
Practically speaking, the work described would most likely never have been done, rather than been done by an artist if that were the only option - it’s uncommon to employ artists to help with incidental tasks relative to side projects, etc.
Re: No elephants: Breakthroughs in image generation
#176> Is it okay to reproduce the hard-won style of other artists using AI? Who owns the resulting art? Who profits from it? Which artists are in the training data for AI, and what is the legal and ethical status of using copyrighted work for training? These were important questions before multimodal AI, but now developing answers to them is increasingly urgent. I have to disagree with the conclusion. This was an importa…
I don't think there's consensus around that idea. Lots of people (myself included) feel that copyright is already vastly overreaching, and that AI represents forward progress for the proliferation of art in society (its crap today, but digital cameras were crap in 2007 and look where they are now). Its also not clear for example that Studio Ghibli lost by having their art style plastered all over the internet. I went…
> but I certainly don't think that AI without restrictions is going to lead to fewer people with art jobs.
It's great that you think that but in reality a lot of artists are saying they're getting less work these days. Maybe that's the result of a shitty economy but I find it very difficult to believe this technology isn't actively stealing work from people.
Re: No elephants: Breakthroughs in image generation
#177This is a before/after moment for image generation. A simple example is the background images on a ton of (mediocre) music youtube channels. They almost all use AI generated images that are full of nonsense the closer you look. Jazz channels will feature coffee shops with garbled text on the menu and furniture blending together. I bet all of that disappears over the next few months. On another note, and perhaps other…
I can never think of anything to talk to an AI about. I run LM local, as well
Re: No elephants: Breakthroughs in image generation
#178Earlier quoted context omitted.
Right the difference is that it’s a large company looking at it then copying it and reselling it without credit, which basically everyone would understand as bad without the indirection of a model. Edit: the key words here are “company” and “reselling”
But it’s not copying and reselling — it’s imitation. Copying is controlled by copyrights. And imitation isn’t controlled by anything. As for a company: a company is just a group of people acting together.
Re: No elephants: Breakthroughs in image generation
#179Earlier quoted context omitted.
I love using LLMs to generate pictures. I'd call myself rather creative, but absolutely useless in any artistic craft. Now, I can just describe any image I can imagine and get 90% accurate results, which is good enough for the presentations I hold, online pet projects (created a squirrel-themed online math-learning game for which I previously would have needed a designer to create squirrel highschool themed imagery)…
If you use this technology, you're actively harming creative labor.
The market is saturated and the way it works means ten get rich for every million artists. I feel as though this has been pretty constant throughout history.
Of course there's a lot of talent out there, "wasted", but I think that's always been the case. How many William Shakesmans did we lose with all the war, famine, disease?
I actually decided I'd probably never write music again after 1-shot making a song about the south Korea coup attempt several months ago. I had the song done before the news really even hit the US. Why would I destroy my own hearing writing music anymore when I can prompt an AI to do it for me, with the same net result - no one cares.
here's the 3-shot remix, the triangle cracks me up so much that i had to upload it https://soundcloud.com/djoutcold/coup-detat-symphony-remix
the "original" "1-shot" is on my soundcloud page as well. https://soundcloud.com/djoutcold/i-aint-even-writing-music-a...
it's in lojban. That's why you can't understand it. Yes. Lojban. Brings a tear to my eye every time i hear it. fkin AI
[0] more my style - hold music for our PBX https://soundcloud.com/djoutcold/bew-hold-music also all my stuff is CC licensed, mostly CC0 at this point.
Re: No elephants: Breakthroughs in image generation
#180> Is it okay to reproduce the hard-won style of other artists using AI? Who owns the resulting art? Who profits from it? Which artists are in the training data for AI, and what is the legal and ethical status of using copyrighted work for training? These were important questions before multimodal AI, but now developing answers to them is increasingly urgent. I have to disagree with the conclusion. This was an importa…
“Fair” doesn’t matter. The only consensus that matters is what is legal and profitable. The former seems to be pretty much decided in favor of AI, with some open question about whether large media companies enjoy protections that smaller artists don’t. (The legal battle when some AI company finally decides to let their model imitate Disney stuff is going to be epic.) Profitable remains to be seen, but doesn’t matter…
None of the cases against AI companies have been decided afaik. There's a ton of ongoing litigation.
> but doesn’t matter much while investors’ money is so plentiful.
More and more people are realizing how wasteful this stuff is every day.