Live data from Hacker News

4o Image Generation

openai.com

141–150 of 629 posts

Re: 4o Image Generation

#141

Earlier quoted context omitted.

Are you sure you are using the new 4o image generation? https://imgur.com/a/wGkBa0v

That is an unexpectedly literal definition of "full glass".

Generating an image of a completely full glass of wine has been one of the popular limitations of image generators, the reason being neural networks struggling to generalise outside of their training data (there are almost no pictures on the internet of a glass "full" of wine). It seems they implemented some reasoning over images to overcome that.

Re: 4o Image Generation

#142

Earlier quoted context omitted.

Are you sure you are using the new 4o image generation? https://imgur.com/a/wGkBa0v

That is an unexpectedly literal definition of "full glass".

That's the point. With the old models they all failed to produce a wine glass that is completley to the brim full. Because you can't find that a lot in the data they used for training.

Re: 4o Image Generation

#143

Earlier quoted context omitted.

Are you sure you are using the new 4o image generation? https://imgur.com/a/wGkBa0v

That is an unexpectedly literal definition of "full glass".

This is another cool example from their blog

https://imgur.com/a/Svfuuf5

Re: 4o Image Generation

#144

Earlier quoted context omitted.

That's expected with any image generating models because they aren't trained with an alpha channel. It's more pragmatic to pipeline the results to a background removal model. EDIT: It appears GPT-4o is different as there is a video demo dedicated to transparancy.

There's an entire video in the post dedicated to how well it does transparency: https://openai.com/index/introducing-4o-image-generation/?vi... I suspect we're getting a flood of comment from people who are using Dall-E.

The video was helpful. I started with the prompt "Generate a transparent image. "

And that created the isolated image on a transparent background.

Thank-you.

Re: 4o Image Generation

#146
post #95

I’ll just be happy with not everything having that over saturated cg/cartoon style that you cant prompt your way out of.

Is that an artifact of the training data? Where are all these original images with that cartoony look that it was trained on?

A large part of deviantart.com would fit that description. There are also a lot of cartoony or CG images in communities dedicated to fanart. Another component in there is probably the overly polished and clean look of stock images, like the front page results of shutterstock.

"Typical" AI images are this blend of the popular image styles of the internet. You always have a bit of digital drawing + cartoon image + oversaturated stock image + 3d render mixed in. Models trained on just one of these work quite well, but for a generalist model this blend of styles is an issue

Re: 4o Image Generation

#147

Earlier quoted context omitted.

[flagged]

[flagged]

> That is a great right, as long as it's not programmers.

You realize that almost weekly we have new AI models coming out that are better and better at programming? It just happened that the image generation is an easier problem than programming. But make no mistake, AI is coming for us too.

That's the price of automating everything.

Re: 4o Image Generation

#148
post #104

What's important about this new type of image generation that's happening with tokens rather than with diffusion, is that this is effectively reasoning in pixel space. Example: Ask it to draw a notepad with an empty tic-tac-toe, then tell it to make the first move, then you make a move, and so on. You can also do very impressive information-conserving translations, such as changing the drawing style, but also stuff l…

It still can't generate a full glass of wine. Even in follow up questions it failed to manipulate the image correctly.

Yeah, it seems like somewhere in the semantic space (which then gets turned into a high resolution image using a specialized model probably) there is not enough space to hold all this kind of information. It becomes really obvious when you try to meaningfully modify a photo of yourself, it will lose your identity.

For Gemini it seems to me there's some kind of "retain old pixels" support in these models since simple image edits just look like a passthrough, in which case they do maintain your identity.

Re: 4o Image Generation

#149

Earlier quoted context omitted.

It still can't generate a full glass of wine. Even in follow up questions it failed to manipulate the image correctly.

Are you sure you are using the new 4o image generation? https://imgur.com/a/wGkBa0v

Looks amazing,can you please also create a unconventional image like the clock at 2:35 , I tried it something like this with gemini when some redditor asked it and it failed so wondering if 4o does do it
Post reply on HN