Live data from Hacker News

No elephants: Breakthroughs in image generation

oneusefulthing.org

181–190 of 373 posts

Re: No elephants: Breakthroughs in image generation

#181
post #158

Earlier quoted context omitted.

>I love using LLMs to generate pictures. I'd call myself rather creative, but absolutely useless in any artistic craft. Now, I can just describe any image I can imagine and get 90% accurate results May I ask what you use? I'm not yet even a paid subscriber to any of the models, because my company offer a corporate internal subscription chatbot and code integration that works well enough for what I've been doing so fa…

I was generating pictures to use for a little game I made with my six and ten year old kids. They were so excited to see us go from idea to execution so quickly, they were laughing and we had a ton of fun. The only thing that disappointed me was I got throttled. We’d need to pay for API image gen to get it even faster. I made a logo for an internal product that wouldn’t have had a logo otherwise at our company. I als…

Feels like you didn't answer the question. i know you weren't who was asked, but still.

Re: No elephants: Breakthroughs in image generation

#182

Earlier quoted context omitted.

The prompt enrichment thing is pretty standard. Everyone does that bit, though some make it user-visible. On Grok it used to populate to the frontend via the download name on the image. The image editing is interesting.

All the stable diffusion software I've used names the files after some form of the prompt, and probably because SD weights the first tokens higher than the last tokens, probably as a side effect of the way the CLIP/BLIP works. I doubt any of these companies have rolled their own interface to stable diffusion / transformers. It's copy and paste from huggingface all the way down. I'm still waiting for a confirmed Diffu…

Auto1111 and co are using the prompt in the filename because it's convenient, not due to some inherent CLIP mechanism.

If you think that companies like OpenAI (for all the criticisms they deserve) don't use their own inference harness and image models I have a bridge to sell to you.

Re: No elephants: Breakthroughs in image generation

#183
post #31

Earlier quoted context omitted.

That blog post is a year old. There has been a lot of progress since then: https://doubiiu.github.io/projects/ToonCrafter/

Very impressive. This is going to result in an explosion of content creation by pro studios, just as CG with cel-shading renderers did. I greatly prefer the hand-drawn + AI tweened look to the current low-budget CG 3D models look.

Yeah it will be much better than the low budget 3D models in anime, hopefully there will be a production ready product that works at a high enough resolution and studios will probably adopt it instead of using cheap labor.

Re: No elephants: Breakthroughs in image generation

#184
post #3

This is a before/after moment for image generation. A simple example is the background images on a ton of (mediocre) music youtube channels. They almost all use AI generated images that are full of nonsense the closer you look. Jazz channels will feature coffee shops with garbled text on the menu and furniture blending together. I bet all of that disappears over the next few months. On another note, and perhaps other…

> I am finding myself surprised at how little use I have for this stuff

I think this will change as more practical use cases begin to emerge as this is all brand new. For example, the photos you take with your smartphone can tell a story or be annotated so you can see things in the photos you didn't think about but your profile thinks you might. Things will get more sophisticated soon.

Re: No elephants: Breakthroughs in image generation

#185
post #110
post #99

Earlier quoted context omitted.

A friend who is heavily inked has gone on at length to me about understanding skin elasticity--particularly how it changes over a lifetime--as well as the way joints and muscles change and distort visual lines, etc. It sure seems like a skilled trade to me. And, I don't know, depth of penetration of a needle in flesh and sanitation don't strike me as minor things to get right.

People love to make things seem harder than they are. I tattoo people, I am aware about skin types, usually thats not a big issue unless its heavily scarred. The quality of your tattoo machine matters most, as my 70€ eBay makeshift one wasnt nearly as good as a proper one. Amount of ink matters, needle depth, skin type, sweat. But thats stuff you have figured put after your 20th tattoo. Its like knowing datatypes in…

> People love to make things seem harder than they are.

In my experience people tend to underestimate or downplay how difficult something will be or how complex it is. This happens in people who know only a little about something, but also in people who are highly experienced because it becomes normal and easy for them and they can quickly evaluate a situation and know which considerations don't apply.

Re: No elephants: Breakthroughs in image generation

#186

Earlier quoted context omitted.

I love the idea, but I feel like I have to say that I’ve got a pretty solid idea of what the total perspective vortex would look like for someone being subjected to it. When I first read the books I immediately had a visual and that has never changed when I’ve read them again (and again…). I’m not sure what that says about either of us, but I would say that your definitive “quite hard to visualise” statement is very…

Vortex may be not so much but there are other hard to visualize things. I am on the third book, and I have no idea what Beeblebrox's two heads look like. Second head is often mentioned in passing. Sometimes its mentioned as its always there, other times it feels like its just pops out of somewhere, otherwise, it's like it doesn't exist. There is the scene when they see themselves on the beach on first rescue by the s…

We have alt text for images, you want alt images for text.

You can see other people's interpretation of Zaphod's two heads by watching the BBC HHGTTG show (Mark Wing-Davey) or the movie (Sam Rockwell), among other renditions, which offer completely different interpretations, none of them canonical (not the least of which is because there was no canonical version of HHGTTG according to DA). I'm sure there are multitudes of fan art for HHGTTG on deviantart. Having AI generate an image doesn't offer any more "official" visualization.

Zaphod's second head is mentioned just as much is warranted. If a character has a limp or a crazy haircut it is not mentioned every time, because it has nothing to do with what is going on. And the book mentions that one head is often distracted/asleep, so it sounds like you do have a good visual of what his two heads are like.

While I understand that people think differently and some people are more visual thinkers, a good portion of the concepts expressed through writing are meant to be mindfucks that are difficult to express visually. A picture may be worth a thousand words, but the meat of writing is usually not the visual representation of its concepts. That's a great thing about writing: you can fill in the visuals yourself and it's fodder for fans to discuss.

(BTW, Hotblack Desiato's ship would just be black. Your eyes couldn't focus on it. Even the controls were black labels on a black background. There is nothing here to visualize other than, well, blackness).

Re: No elephants: Breakthroughs in image generation

#187
post #33

> Is it okay to reproduce the hard-won style of other artists using AI? Who owns the resulting art? Who profits from it? Which artists are in the training data for AI, and what is the legal and ethical status of using copyrighted work for training? These were important questions before multimodal AI, but now developing answers to them is increasingly urgent. I have to disagree with the conclusion. This was an importa…

I don't think there's consensus around that idea. Lots of people (myself included) feel that copyright is already vastly overreaching, and that AI represents forward progress for the proliferation of art in society (its crap today, but digital cameras were crap in 2007 and look where they are now). Its also not clear for example that Studio Ghibli lost by having their art style plastered all over the internet. I went…

> "How can we monetize art" remains an open question for society

Yet much of the best art imho is in the wild to the element while being at home at some random place. Or perhaps in someone's collection forgot and displaced. Art's worth will always be an open question.

Re: No elephants: Breakthroughs in image generation

#188
post #3

This is a before/after moment for image generation. A simple example is the background images on a ton of (mediocre) music youtube channels. They almost all use AI generated images that are full of nonsense the closer you look. Jazz channels will feature coffee shops with garbled text on the menu and furniture blending together. I bet all of that disappears over the next few months. On another note, and perhaps other…

> They almost all use AI generated images that are full of nonsense the closer you look. Jazz channels will feature coffee shops with garbled text on the menu and furniture blending together. Noticed tht. Maybe it's my algorithm but YouTube is seemingly filled with these videos now.

All my music cover images are AI generated. At the same time I refuse to listen to AI music. We're all going to sink alone on this one.

What's frustrating me is if I tell the Youtube algo 'don't recommend' to AI music video channels it stops giving me any music video channels. That's not what I want, I just don't want the AI. They need to seperate the two. But of course they need to not do that with AI cover images because otherwise it would harm me. :)

Re: No elephants: Breakthroughs in image generation

#189
post #64
post #62

Earlier quoted context omitted.

I keep waiting for physical objects to become important again. AI isn't coming for the ceramic folks.

I think that would be great. Traditional art market is not nearly as big as the digital space. I'd love if people valued traditional art again as thats the only stuff I do.

Depends on the niche. Original physical art for trading card games or comics is a significant chunk of the income of your typical artist. Digital art in those niches does not have this source of income. But then again digital art has other niches where the actual commission rates are high enough to not make this a problem.

Re: No elephants: Breakthroughs in image generation

#190

Earlier quoted context omitted.

#1, it’s extremely easy to coerce direct copies out of models that artists could be pursued for infringement if they drew, but companies reselling said copyrighted artwork face no penalty #2, yes, it’s a group of people who came together to build an algorithm that learns to extract features learned from images made by other people in order to generate images somewhere between these images in a high dimensional space.…

I guess I don't understand what you mean when you say companies reselling said copyrighted artwork face no penalty. Why wouldn't they? If I was to make a copy of a Studio Ghibli movie and sell it, I would absolutely face a penalty if I was caught.

A common joke is to type in a description of some corporate IP and have ChatGPT generate it without ever saying it directly. Plenty of people have paid a subscription to do that, generate corporate IP that an artist could be sued over, but I don’t believe OpenAI has faced any legal issues if I’m correct, just as an example.
Post reply on HN