Live data from Hacker News

4o Image Generation

openai.com

191–200 of 629 posts

Re: 4o Image Generation

#191

Earlier quoted context omitted.

That's the point. With the old models they all failed to produce a wine glass that is completley to the brim full. Because you can't find that a lot in the data they used for training.

Imagine if they just actually trained the model on a bunch of photographs of a full glass of wine, knowing of this litmus test

A lot of other things are rare in datasets, let alone correctly labeled. Overturned cars (showing the underside), views from under the table, people walking on the ceiling with plausible upside down hair, clothes, and facial features etc etc

Re: 4o Image Generation

#192
well it failed on me, after many tries:

...Once the wait time is up, I can generate the corrected version with exactly eight characters: five mice, one elephant, one polar bear, and one giraffe in a green turtleneck. Let me know if you'd like me to try again later!

Re: 4o Image Generation

#193
post #85

Earlier quoted context omitted.

Agreed. It seems totally unnatural that a couple of nerds high-five awkwardly.

Not awkward. Anatomically uncanny and physically impossible.

While drawing hands is difficult (because the surface morphs in a variety of ways), the shapes and relative proportions are quite simple. That’s how you can have tools like Metahuman[0]

[0]: https://www.unrealengine.com/en-US/metahuman

Re: 4o Image Generation

#194
post #185

Earlier quoted context omitted.

It still can't generate a full glass of wine. Even in follow up questions it failed to manipulate the image correctly.

https://i.imgur.com/xsFKqsI.png "Draw a picture of a full glass of wine, ie a wine glass which is full to the brim with red wine and almost at the point of spilling over... Zoom out to show the full wine glass, and add a caption to the top which says "HELL YEAH". Keep the wine level of the glass exactly the same."

Can't replicate. Maybe the rollout is staggered? Using Plus from Europe, it's consistently giving me a half full glass.

Re: 4o Image Generation

#195
post #178

To avoid confusion, why not always use a general AI model upfront, then depending on the user's prompt, redirect it to a specific model?

The models are noticeably different — for example, o1 and o3 have reasoning, and some users (eg. me) want to tell the model when to use reasoning, and when not.

As to why they don't automatically detect when reasoning could be appropriate and then switch to o3, I don't know, but I'd assume it's about cost (and for most users the output quality is negligible). 4o can do everything, it's just not great at "logic".

Re: 4o Image Generation

#196
post #175

Earlier quoted context omitted.

You know the images themselves don’t get shared in links like that, right? (It even tells you so when you make the link.)

I created a shared link just now, was not presented with any such warning, and have the same problem with the image not showing up: https://chatgpt.com/share/67e319dd-bd08-8013-8f9b-6f5140137f...

The image shows for me.

Re: 4o Image Generation

#197
post #100

Is there any way to see whether a given prompt was serviced by 4o or Dall-E? Currently, my prompts seem to be going to the latter still, based on e.g. my source image being very obviously looped through a verbal image description and back to an image, compared to gemini-2.0-flash-exp-image-generation. A friend with a Plus plan has been getting responses from either. The long-term plan seems to be to move to 4o comple…

If you don't have access to it on ChatGPT yet, you can try Sora, which already has access for me.

Re: 4o Image Generation

#199
post #104

What's important about this new type of image generation that's happening with tokens rather than with diffusion, is that this is effectively reasoning in pixel space. Example: Ask it to draw a notepad with an empty tic-tac-toe, then tell it to make the first move, then you make a move, and so on. You can also do very impressive information-conserving translations, such as changing the drawing style, but also stuff l…

That's very interesting. I would have assumed that 4o is internally using a single seed for the entire conversation, or something analogous to that, to control randomness across image generation requests. Can you share the technical name for this reasoning process so I could look up research about it?

Re: 4o Image Generation

#200

Earlier quoted context omitted.

Can you do this with the prompt of a cow jumping over the moon? I can’t ever seem to get it to make the cow appear to be above the moon. Always literally covering it or to the side etc.

https://chatgpt.com/share/67e31a31-3d44-8011-994e-b7f8af7694... got it on the second try.

To be clear, that is DALL-E, not 4o image generation. (You can see the prompt that 4o generated to give to DALL-E.)
Post reply on HN