Live data from Hacker News

4o Image Generation

openai.com

371–380 of 629 posts

Re: 4o Image Generation

#371
post #288

Visual internet content is completely over. Pack it up

So I spent a good few hours investigating the current state of the art a few weeks ago. I would like to generate a collection of images for the art in a video game. It is incredibly difficult to develop an art style, then get the model to generate a collection of different images in that unique art style. I couldn't work out how to do it. I also couldn't work out how to illustrate the same characters or objects in di…

Even with custom LoRas, controlnets, etc. we're still a pretty long ways from being able to one-click generate thematically consistent images especially in the context of a video game where you really need the ability to generate seamless tiles, animation based spritesheets, etc.

Re: 4o Image Generation

#373

Earlier quoted context omitted.

That's expected with any image generating models because they aren't trained with an alpha channel. It's more pragmatic to pipeline the results to a background removal model. EDIT: It appears GPT-4o is different as there is a video demo dedicated to transparancy.

There's a mod for stable diffusion webui forge/automatic1111/ComfyUI which enables this for all diffusion models (except these closed source ones).

SD extensions like rembg are post-processing effects - with their video transparency demo I'd be curious if 4o actually did training with an alpha channel.

Re: 4o Image Generation

#374

Garbage compared to Midjourney. I don't even know why you'd market this. It's takes a minute or more and the results are what I'd say Midjourney looked like 1.5 years ago.

Prompt adherence is far, far ahead of midjourney. FAR.

I don't have an hour to work an image. It's slow as hell.

Re: 4o Image Generation

#375
post #237

Earlier quoted context omitted.

That's the point. With the old models they all failed to produce a wine glass that is completley to the brim full. Because you can't find that a lot in the data they used for training.

The old models were doing it correct also. There is no one correct way to interpert 'full'. If you go to a wine bar and ask for a full glass of wine, they'll probably interpert that as a double. But you could also interpert it the way a friend would at home, which is about 2-3cm from the rim. Personally I would call a glass of wine filled to the brim 'overfilled', not 'full'.

I think you're missing the context everyone else has - this video is where the "AI can't draw a full glass of wine" meme got traction https://www.youtube.com/watch?v=160F8F8mXlo

The prompts (some generated by ChatGPT itself, since it's instructing DALL-E behind the scenes) include phrases like "full to the brim" and "almost spilling over" that are not up to interpretation at all.

Re: 4o Image Generation

#376
post #194
post #185

Earlier quoted context omitted.

https://i.imgur.com/xsFKqsI.png "Draw a picture of a full glass of wine, ie a wine glass which is full to the brim with red wine and almost at the point of spilling over... Zoom out to show the full wine glass, and add a caption to the top which says "HELL YEAH". Keep the wine level of the glass exactly the same."

Can't replicate. Maybe the rollout is staggered? Using Plus from Europe, it's consistently giving me a half full glass.

I am using Plus from Australia, and while I am not getting a full glass, nor am I getting a half full glass. The glass I'm getting is half empty.

Re: 4o Image Generation

#377
Really liked the fact that the team shared all the shortcomings of the model in the post. Sometimes products just highlights the best results and isn't forthcoming in areas that need improvement. Kudos to the OpenAI team on that.

Re: 4o Image Generation

#378
post #235

Earlier quoted context omitted.

> What's important about this new type of image generation that's happening with tokens rather than with diffusion, is that this is effectively reasoning in pixel space. I do not think that this is correct. Prior to this release, 4o would generate images by calling out to a fully external model (DALL-E). After this release, 4o generates images by calling out to a multi-modal model that was trained alongside it. You c…

I think this is actually correct even if the evidence is not right. See this chat for example: https://chatgpt.com/share/67e355df-9f60-8000-8f36-874f8c9a08...

Honest question, do you believe something just because the bot tells you that?

Re: 4o Image Generation

#379
Just curious if it works for creating a comic strip? I.e. will it maintain the consistency of the characters? I watched a video somewhere they demo'ed it creating comic panels, but I want to create the panels one by one.

Re: 4o Image Generation

#380

Earlier quoted context omitted.

I think this is actually correct even if the evidence is not right. See this chat for example: https://chatgpt.com/share/67e355df-9f60-8000-8f36-874f8c9a08...

Honest question, do you believe something just because the bot tells you that?

No, did you look at my link?
Post reply on HN