Live data from Hacker News

No elephants: Breakthroughs in image generation

oneusefulthing.org

111–120 of 373 posts

Re: No elephants: Breakthroughs in image generation

#111
post #95
post #63

Earlier quoted context omitted.

Should be good enough to alrdy have established some customers.

The problem with that is that people aren't asking for AI generated images in the style of Raven from Topeka with an Etsy shop. They're asking for Ghibli. So the people whose livelihoods are most directly impacted are (assuming they're not centuries dead) the famous, talented, and trend-making artists, not the lower tier making bad Precious Moments knockoffs. Society's problem is understanding that not wanting to pay…

As if the Ghibli trend wasnt just a short trend people will have forgotten about in 4 weeks... Also I couldnt care less about big studios, they print money anyway.

Re: No elephants: Breakthroughs in image generation

#112
4o, despite OpenAI's practically draconian content policies, is a pretty big leap forward. I put together a comparison of some of the most competitive generative models (Imagen, 4o, Flux, and MJ7) where I prioritized increasingly difficult prompt adherence. If Imagen 3 had 4o's multimodal capabilities (being able to make constant adjustments against a generated image by prompting) I would say its nearly on-par with 4o.

https://genai-showdown.specr.net

Re: No elephants: Breakthroughs in image generation

#115
post #26

Looking at the example where the coffee table is swapped, I notice every time the image is reprocessed it mutates, based on the previous iteration, and objects become more bizarre each time, like chinese whispers. * The weird-ass basket decoration on the table originally has some big chain links (maybe anchor chain, to keep the theme with the beach painting). By the third version, they're leathery and are merging wit…

Yeah, this is in my opinion the biggest limitation of the current gen GPT 4o image generation: it is incapable of editing only parts of an image. I assume what it does every time is tokenizing the source image, then transforming it according to the prompt and then giving you the final result. For some use cases that’s fine but if you really just want a small edit while keeping the rest of the image intact you’re out…

I thought the selection tool allows you to limit the area of the image that a revision will make changes to, but I tested it and I still see changes outside of the selected area which is good to know.

As an example the tape spindles, among other changes, are different: https://chatgpt.com/share/67f53965-9480-800a-a166-a6c1faa87c...

https://help.openai.com/en/articles/9055440-editing-your-ima...

Re: No elephants: Breakthroughs in image generation

#116
post #23

Earlier quoted context omitted.

> it's unfair for artists to have their works sucked up I never thought it was unfair to artists for others to look at their work and imitate it. That seems to me to be what artists have been doing since the second caveman looked at a hand painting on a cave wall and thought, ‘huh, that’s pretty neat! I’d like to try my hand at that!’

You don't see a massive difference in the shear number of images that the AI can look at and the speed at which it can imitate it as a fundamental difference between AI and a human copying works or styles? For a human it took a lot of practice and a lot of time and effort. But now it takes practically no time or effort at all.

Well yeah, but copyright infringement isn't a function of how quickly you can view and create works.

Copyright is meant to secure distribution of works you create. It's not a tool to stop people from creating art because it looks like your art. That has been a thing for centuries, we even categorize art by it's style. Imagine anime was had to adhere to a copyright interpretation of "it's my style!".

Re: No elephants: Breakthroughs in image generation

#117
post #29

Hmmm isn't stable diffusion already doing that?

SD has a very primitive conceptual model. Basically “bag of words nudging pixels around for a while”. Words near each other influence each other. But, there’s nearly no understanding of grammar.

Midjourney is similar with text prompts. But, with image prompts it is able to understand content separately from style. You can give it a photo of two people and it can return many images of recognizable approximations of those people in different poses.

SD can only start from pixels, blur and deblur those pixels in place.

MJ image prompts probably works via image-to-tokens added on to your text-to-tokens-to-image.

Re: No elephants: Breakthroughs in image generation

#118
post #34

It's interesting to hear people side with the artists when in previous discussions on this forum I've gotten significant approval/agreement arguing that copyright is far too long. As I've argued in the past, I think copyright should last maybe five years: in this modern era, monetizing your work doesn't (usually) have to take more than a short time. I'd happily concede to some sort of renewal process to extend that p…

As you grow older and run through more cycles of general opinions, you realize that pretty much everyone is in it for themselves, what serves them best, and support what narrative aligns with that.

2007: Copyright is garbage and must be abolished (so I can get music/movies free)

2025: Copyright needs to be strengthened (so my artistic abilities retain value)

Re: No elephants: Breakthroughs in image generation

#120
post #3

This is a before/after moment for image generation. A simple example is the background images on a ton of (mediocre) music youtube channels. They almost all use AI generated images that are full of nonsense the closer you look. Jazz channels will feature coffee shops with garbled text on the menu and furniture blending together. I bet all of that disappears over the next few months. On another note, and perhaps other…

> They almost all use AI generated images that are full of nonsense the closer you look. Jazz channels will feature coffee shops with garbled text on the menu and furniture blending together. Noticed tht. Maybe it's my algorithm but YouTube is seemingly filled with these videos now.

They insist on feeding me AI generated videos about "HOA Karens" for some odd reason.

True, I do enjoy watching the LawTubers and sometimes they talk about HOAs but that is a far stretch from someone taking a reddit post and laundering it through the robots.

Post reply on HN