Live data from Hacker News

No elephants: Breakthroughs in image generation

oneusefulthing.org

91–100 of 373 posts

Re: No elephants: Breakthroughs in image generation

#91
post #26

Earlier quoted context omitted.

Yeah, this is in my opinion the biggest limitation of the current gen GPT 4o image generation: it is incapable of editing only parts of an image. I assume what it does every time is tokenizing the source image, then transforming it according to the prompt and then giving you the final result. For some use cases that’s fine but if you really just want a small edit while keeping the rest of the image intact you’re out…

It just means that you comp it together manually. That's still much better than having to set up some inpainting pipeline or whatever.

100%. Multimodal images surpass ComfyUI and inpainting (for now). It's a step function improvement in image generation.

I'm hoping we see an open weights or open source model with these capabilities soon, because good tools need open models.

As has happened in the past, once an open implementation of DallE or whatever comes out, the open source community pushes the capabilities much further by writing lots of training, extensions, and pipelines. The results look significantly better than closed SaaS models.

Re: No elephants: Breakthroughs in image generation

#92

Looking at the example where the coffee table is swapped, I notice every time the image is reprocessed it mutates, based on the previous iteration, and objects become more bizarre each time, like chinese whispers. * The weird-ass basket decoration on the table originally has some big chain links (maybe anchor chain, to keep the theme with the beach painting). By the third version, they're leathery and are merging wit…

It's kind of clear that for every request, it generates a new image entirely. Some people are speculating a diffusion decoder but i think it's more likely an implementation of VAR - https://arxiv.org/abs/2404.02905 . So rather than predicting each patch at the target resolution right away, it starts with the image (as patches) at a very small resolution and increasingly scales up. I guess that could make it hard for…

BUT it's doing a stunningly better job replicating previous scenes than it did before. I asked it just now for a selfie of two biker buddies on a Nevada highway, but one is a quokka and one is a hyrax. It did it. Then I asked for the same photo with late afternoon lighting, and it did a pretty amazing job of preserving the context where just a few months ago it would have had no idea what it had done before.

Also, sweet jesus, after more than a year of hilarious frustration, it now knows that a flying squirrel is a real animal and not just a tree squirrel with butterfly wings.

Re: No elephants: Breakthroughs in image generation

#93
post #88

Earlier quoted context omitted.

> some just heavily overcharge for mediocre results if people are paying, then they aren't "overcharging"

In the tattoo business people have no other places to go. Charging 1k€+ for half a sleeve is extremely overcharged. If people are paying they often simply don't have enough alternatives.

Given the time commitment and network of basic biological, anatomical, and health knowledge required, that doesn't strike me as an insane price, assuming an artist who is able to create the requested art.

Re: No elephants: Breakthroughs in image generation

#94
post #92

Earlier quoted context omitted.

It's kind of clear that for every request, it generates a new image entirely. Some people are speculating a diffusion decoder but i think it's more likely an implementation of VAR - https://arxiv.org/abs/2404.02905 . So rather than predicting each patch at the target resolution right away, it starts with the image (as patches) at a very small resolution and increasingly scales up. I guess that could make it hard for…

BUT it's doing a stunningly better job replicating previous scenes than it did before. I asked it just now for a selfie of two biker buddies on a Nevada highway, but one is a quokka and one is a hyrax. It did it. Then I asked for the same photo with late afternoon lighting, and it did a pretty amazing job of preserving the context where just a few months ago it would have had no idea what it had done before. Also, sw…

I agree. I'm not saying it's a different model generating the images. 4o is clearly generating the images itself rather than sending a prompt to some other model. I'm speculating about the mechanism for generation in the model itself.

Re: No elephants: Breakthroughs in image generation

#95
post #63

Earlier quoted context omitted.

What about the half of the remaining artists that are below the new median?

Should be good enough to alrdy have established some customers.

The problem with that is that people aren't asking for AI generated images in the style of Raven from Topeka with an Etsy shop. They're asking for Ghibli. So the people whose livelihoods are most directly impacted are (assuming they're not centuries dead) the famous, talented, and trend-making artists, not the lower tier making bad Precious Moments knockoffs. Society's problem is understanding that not wanting to pay for bad Precious Moments knockoffs is rational, while not wanting to pay, say, a Studio Ghibli for quality, professional creativity is insane.

Re: No elephants: Breakthroughs in image generation

#96
post #92

Earlier quoted context omitted.

BUT it's doing a stunningly better job replicating previous scenes than it did before. I asked it just now for a selfie of two biker buddies on a Nevada highway, but one is a quokka and one is a hyrax. It did it. Then I asked for the same photo with late afternoon lighting, and it did a pretty amazing job of preserving the context where just a few months ago it would have had no idea what it had done before. Also, sw…

I agree. I'm not saying it's a different model generating the images. 4o is clearly generating the images itself rather than sending a prompt to some other model. I'm speculating about the mechanism for generation in the model itself.

Oh, no, I wasn't taking issue with what you said, just reacting that, yes, it's not editing the same image but redrawing from scratch every time, BUT it's doing a much better job of that, with some understanding of the context of the previous image so that it can tweak it, even if it's never bit for bit identical.

Re: No elephants: Breakthroughs in image generation

#97

Earlier quoted context omitted.

It’s interesting that out of all the aquatic animals it could have used, it chose one that perhaps looks the most like an elephant.

That's just luck of the draw, in a new session with the same prompt it outputted a bird: https://i.imgur.com/SLT8cYe.png

I'm not convinced. I tried it and it showed me a swimming hippopotamus, which is even more elephant-like than the turtle. I tried again and it gave me a pelican, which is not generally very elephantish, but this particular one has a gray body with a texture that looks a lot like elephant skin.

Re: No elephants: Breakthroughs in image generation

#98
post #93
post #88

Earlier quoted context omitted.

In the tattoo business people have no other places to go. Charging 1k€+ for half a sleeve is extremely overcharged. If people are paying they often simply don't have enough alternatives.

Given the time commitment and network of basic biological, anatomical, and health knowledge required, that doesn't strike me as an insane price, assuming an artist who is able to create the requested art.

Tbh you barely have to know anything, most important one is how deep to go with the needle and sanitizing. Everything else is not rlly important.

Re: No elephants: Breakthroughs in image generation

#99
post #98
post #93

Earlier quoted context omitted.

Given the time commitment and network of basic biological, anatomical, and health knowledge required, that doesn't strike me as an insane price, assuming an artist who is able to create the requested art.

Tbh you barely have to know anything, most important one is how deep to go with the needle and sanitizing. Everything else is not rlly important.

A friend who is heavily inked has gone on at length to me about understanding skin elasticity--particularly how it changes over a lifetime--as well as the way joints and muscles change and distort visual lines, etc. It sure seems like a skilled trade to me.

And, I don't know, depth of penetration of a needle in flesh and sanitation don't strike me as minor things to get right.

Re: No elephants: Breakthroughs in image generation

#100
post #3

This is a before/after moment for image generation. A simple example is the background images on a ton of (mediocre) music youtube channels. They almost all use AI generated images that are full of nonsense the closer you look. Jazz channels will feature coffee shops with garbled text on the menu and furniture blending together. I bet all of that disappears over the next few months. On another note, and perhaps other…

> They almost all use AI generated images that are full of nonsense the closer you look. Jazz channels will feature coffee shops with garbled text on the menu and furniture blending together. Noticed tht. Maybe it's my algorithm but YouTube is seemingly filled with these videos now.

Youtube Studio is has build in AI thumbnail functionality. Google actively encourages use of AI to clickbait and to generate automatic AI replies to comments ala onlyfaps giving your viewers that feeling of interaction without reading their comments.
Post reply on HN