A: Your face is pressed up against the ceiling!
No elephants: Breakthroughs in image generation
161–170 of 373 posts
Re: No elephants: Breakthroughs in image generation
#162Re: No elephants: Breakthroughs in image generation
#163This is a before/after moment for image generation. A simple example is the background images on a ton of (mediocre) music youtube channels. They almost all use AI generated images that are full of nonsense the closer you look. Jazz channels will feature coffee shops with garbled text on the menu and furniture blending together. I bet all of that disappears over the next few months. On another note, and perhaps other…
That is a very interesting point about how little use of AI most of us making day to day, despite the potential utility that seems to be lurking. I think it just takes time for people and economies to adapt to new technology. Even if technological progress on AI were to stop today, and the best models that exist in 2030 are the same models we have now, there would still be years of social and economic change as peopl…
Re: No elephants: Breakthroughs in image generation
#164Earlier quoted context omitted.
Yeah, this is in my opinion the biggest limitation of the current gen GPT 4o image generation: it is incapable of editing only parts of an image. I assume what it does every time is tokenizing the source image, then transforming it according to the prompt and then giving you the final result. For some use cases that’s fine but if you really just want a small edit while keeping the rest of the image intact you’re out…
I thought the selection tool allows you to limit the area of the image that a revision will make changes to, but I tested it and I still see changes outside of the selected area which is good to know. As an example the tape spindles, among other changes, are different: https://chatgpt.com/share/67f53965-9480-800a-a166-a6c1faa87c... https://help.openai.com/en/articles/9055440-editing-your-ima...
Re: No elephants: Breakthroughs in image generation
#165Earlier quoted context omitted.
I was wondering yesterday how AI is coming along for tweening animation frames. I just did a quick search and apparently last year the state of the art was garbage: https://yosefk.com/blog/the-state-of-ai-for-hand-drawn-anima... Maybe this multimodal thing can fix that?
That blog post is a year old. There has been a lot of progress since then: https://doubiiu.github.io/projects/ToonCrafter/
Re: No elephants: Breakthroughs in image generation
#166Wonderful to be alive for these step changes in human capability.
Re: No elephants: Breakthroughs in image generation
#167Diagrams are still a big unsolved problem. Making diagrams for a talk or paper is an extremely tedious process and I am still waiting for a good multimodal LLM solution for this. It should take a sketch and/or text description of what you want and in a few iterations you should get what you want. GPT4o tries hard but results are still bad.
I've had the best luck at getting it to produce diagrams-as-code like mermaid or plantuml.
Re: No elephants: Breakthroughs in image generation
#168> Is it okay to reproduce the hard-won style of other artists using AI? Who owns the resulting art? Who profits from it? Which artists are in the training data for AI, and what is the legal and ethical status of using copyrighted work for training? These were important questions before multimodal AI, but now developing answers to them is increasingly urgent. I have to disagree with the conclusion. This was an importa…
I don't think there's consensus around that idea. Lots of people (myself included) feel that copyright is already vastly overreaching, and that AI represents forward progress for the proliferation of art in society (its crap today, but digital cameras were crap in 2007 and look where they are now). Its also not clear for example that Studio Ghibli lost by having their art style plastered all over the internet. I went…
This sounds like a rewording of "You won't get paid, but this is a great opportunity for you because you'll get exposure".
Re: No elephants: Breakthroughs in image generation
#169My understanding is it’s a meta-LLM approach, using multiple models and having them interact. I feel like it’s also evidence that OpenAI is not seriously pursuing AGI (just my opinion, I know there’s some on here who would aggressively disagree), but rather market use cases. It feels like an acceptance that any given model, at least now, has its own limitations but can get more useful in combination.
Re: No elephants: Breakthroughs in image generation
#170This is a before/after moment for image generation. A simple example is the background images on a ton of (mediocre) music youtube channels. They almost all use AI generated images that are full of nonsense the closer you look. Jazz channels will feature coffee shops with garbled text on the menu and furniture blending together. I bet all of that disappears over the next few months. On another note, and perhaps other…
I love using LLMs to generate pictures. I'd call myself rather creative, but absolutely useless in any artistic craft. Now, I can just describe any image I can imagine and get 90% accurate results, which is good enough for the presentations I hold, online pet projects (created a squirrel-themed online math-learning game for which I previously would have needed a designer to create squirrel highschool themed imagery)…