Live data from Hacker News

No elephants: Breakthroughs in image generation

oneusefulthing.org

161–170 of 373 posts

Re: No elephants: Breakthroughs in image generation

#163
post #68
post #3

This is a before/after moment for image generation. A simple example is the background images on a ton of (mediocre) music youtube channels. They almost all use AI generated images that are full of nonsense the closer you look. Jazz channels will feature coffee shops with garbled text on the menu and furniture blending together. I bet all of that disappears over the next few months. On another note, and perhaps other…

That is a very interesting point about how little use of AI most of us making day to day, despite the potential utility that seems to be lurking. I think it just takes time for people and economies to adapt to new technology. Even if technological progress on AI were to stop today, and the best models that exist in 2030 are the same models we have now, there would still be years of social and economic change as peopl…

Unless I'm doing something simple like writing out some basic shell script or python program, it's often easier to just do something myself than take the time to explain what I want to an LLM. There's something to be said about taking the time to formulate your plan in clear steps ahead of time, but for many problems it just doesn't feel like it's worth the time to write it all out.

Re: No elephants: Breakthroughs in image generation

#164
post #26

Earlier quoted context omitted.

Yeah, this is in my opinion the biggest limitation of the current gen GPT 4o image generation: it is incapable of editing only parts of an image. I assume what it does every time is tokenizing the source image, then transforming it according to the prompt and then giving you the final result. For some use cases that’s fine but if you really just want a small edit while keeping the rest of the image intact you’re out…

I thought the selection tool allows you to limit the area of the image that a revision will make changes to, but I tested it and I still see changes outside of the selected area which is good to know. As an example the tape spindles, among other changes, are different: https://chatgpt.com/share/67f53965-9480-800a-a166-a6c1faa87c... https://help.openai.com/en/articles/9055440-editing-your-ima...

Yeah, I'm not sure what the selection brush actually does. Is it just a hint to the LLM?

Re: No elephants: Breakthroughs in image generation

#165
post #31

Earlier quoted context omitted.

I was wondering yesterday how AI is coming along for tweening animation frames. I just did a quick search and apparently last year the state of the art was garbage: https://yosefk.com/blog/the-state-of-ai-for-hand-drawn-anima... Maybe this multimodal thing can fix that?

That blog post is a year old. There has been a lot of progress since then: https://doubiiu.github.io/projects/ToonCrafter/

Very impressive. This is going to result in an explosion of content creation by pro studios, just as CG with cel-shading renderers did. I greatly prefer the hand-drawn + AI tweened look to the current low-budget CG 3D models look.

Re: No elephants: Breakthroughs in image generation

#166
Anyone stuck claiming AI isn’t useful - there are so many useful things it can now do. With text that makes sense you can generate invitations for your next picnic. That wasn’t possible mere weeks ago.

Wonderful to be alive for these step changes in human capability.

Re: No elephants: Breakthroughs in image generation

#167

Diagrams are still a big unsolved problem. Making diagrams for a talk or paper is an extremely tedious process and I am still waiting for a good multimodal LLM solution for this. It should take a sketch and/or text description of what you want and in a few iterations you should get what you want. GPT4o tries hard but results are still bad.

I've had the best luck at getting it to produce diagrams-as-code like mermaid or plantuml.

I know but those diagrams often don’t adequately capture what I want. Think of diagrams in nice technical talks or papers. I’ve even tried having the LLM (Claude) generate SVGs. They all fall short.

Re: No elephants: Breakthroughs in image generation

#168
post #33

> Is it okay to reproduce the hard-won style of other artists using AI? Who owns the resulting art? Who profits from it? Which artists are in the training data for AI, and what is the legal and ethical status of using copyrighted work for training? These were important questions before multimodal AI, but now developing answers to them is increasingly urgent. I have to disagree with the conclusion. This was an importa…

I don't think there's consensus around that idea. Lots of people (myself included) feel that copyright is already vastly overreaching, and that AI represents forward progress for the proliferation of art in society (its crap today, but digital cameras were crap in 2007 and look where they are now). Its also not clear for example that Studio Ghibli lost by having their art style plastered all over the internet. I went…

> Its also not clear for example that Studio Ghibli lost by having their art style plastered all over the internet. I went home and watched a Ghibli film that week, as I'm sure many others did as well. Their revenue is probably up quite a bit right now?

This sounds like a rewording of "You won't get paid, but this is a great opportunity for you because you'll get exposure".

Re: No elephants: Breakthroughs in image generation

#169
I’ve seen a few YouTube thumbnail generation examples on Reddit (I’m on vacation so not gonna search for a link) that show multimodal with inline text giving specific instructions. It’s impressed me in a way that I haven’t been with LLMs for 2 years, IE it’s not just getting better at what it already does, but a totally new and intuitive way of working with generative AI.

My understanding is it’s a meta-LLM approach, using multiple models and having them interact. I feel like it’s also evidence that OpenAI is not seriously pursuing AGI (just my opinion, I know there’s some on here who would aggressively disagree), but rather market use cases. It feels like an acceptance that any given model, at least now, has its own limitations but can get more useful in combination.

Re: No elephants: Breakthroughs in image generation

#170
post #3

This is a before/after moment for image generation. A simple example is the background images on a ton of (mediocre) music youtube channels. They almost all use AI generated images that are full of nonsense the closer you look. Jazz channels will feature coffee shops with garbled text on the menu and furniture blending together. I bet all of that disappears over the next few months. On another note, and perhaps other…

I love using LLMs to generate pictures. I'd call myself rather creative, but absolutely useless in any artistic craft. Now, I can just describe any image I can imagine and get 90% accurate results, which is good enough for the presentations I hold, online pet projects (created a squirrel-themed online math-learning game for which I previously would have needed a designer to create squirrel highschool themed imagery)…

If you use this technology, you're actively harming creative labor.
Post reply on HN