Live data from Hacker News

4o Image Generation

openai.com

131–140 of 629 posts

Re: 4o Image Generation

#131
post #78
post #72

Earlier quoted context omitted.

It's a nitpick about the repetitive phrasing for announcements : Our most yet|ever.

I hate modern marketing trends. This one isn't even my biggest gripe. If I could eliminate any word from the English language forever, it would be "effortlessly".

Idk, right now I think I'd eliminate "blazingly fast" from software engineering vocabulary.

Re: 4o Image Generation

#132
post #85

Earlier quoted context omitted.

The high five looks extremely unnatural. Their wrists are aligned, but their fingers aren't, somehow? If that's best of 8, I'd love to see the outtakes.

Agreed. It seems totally unnatural that a couple of nerds high-five awkwardly.

Not awkward. Anatomically uncanny and physically impossible.

Re: 4o Image Generation

#133
post #100

Is there any way to see whether a given prompt was serviced by 4o or Dall-E? Currently, my prompts seem to be going to the latter still, based on e.g. my source image being very obviously looped through a verbal image description and back to an image, compared to gemini-2.0-flash-exp-image-generation. A friend with a Plus plan has been getting responses from either. The long-term plan seems to be to move to 4o comple…

4o generates top down (picture goes from mostly blurry to clear starting from the top). If it's not generating like that for you then you don't have it yet.

That's useful, thank you! But it also highlights my point: Why do I have to observe minor details about how the result is being presented to me to know which model was used?

I get the intent to abstract it all behind a chat interface, but this seems a bit too much.

Re: 4o Image Generation

#135
post #104

What's important about this new type of image generation that's happening with tokens rather than with diffusion, is that this is effectively reasoning in pixel space. Example: Ask it to draw a notepad with an empty tic-tac-toe, then tell it to make the first move, then you make a move, and so on. You can also do very impressive information-conserving translations, such as changing the drawing style, but also stuff l…

It still can't generate a full glass of wine. Even in follow up questions it failed to manipulate the image correctly.

Are you sure you are using the new 4o image generation?

https://imgur.com/a/wGkBa0v

Re: 4o Image Generation

#136

Earlier quoted context omitted.

It still can't generate a full glass of wine. Even in follow up questions it failed to manipulate the image correctly.

Are you sure you are using the new 4o image generation? https://imgur.com/a/wGkBa0v

That is an unexpectedly literal definition of "full glass".

Re: 4o Image Generation

#137
post #124
post #105

Earlier quoted context omitted.

Pretty sure the modern Gemini image models can already do token based image generation/editing and are significantly better and faster.

Yeah Gemini has had this for a few weeks, but much lower resolution. Not saying 4o is perfect, but my first few images with it are much more impressive than my first few images with Gemini.

weeks, ya'll, weeks!

Re: 4o Image Generation

#138
post #95

I’ll just be happy with not everything having that over saturated cg/cartoon style that you cant prompt your way out of.

Is that an artifact of the training data? Where are all these original images with that cartoony look that it was trained on?

Wild speculation: video game engines. You want your model to understand what a car looks like from all angles, but it’s expensive to get photos of real cars from all angles, so instead you render a car model in UE5, generating hundreds of pictures of it, from many different angles, in many different colors and styles.

Re: 4o Image Generation

#139
post #133

Earlier quoted context omitted.

4o generates top down (picture goes from mostly blurry to clear starting from the top). If it's not generating like that for you then you don't have it yet.

That's useful, thank you! But it also highlights my point: Why do I have to observe minor details about how the result is being presented to me to know which model was used? I get the intent to abstract it all behind a chat interface, but this seems a bit too much.

Oh I agree 100%. Open AI roll outs leave much to be desired. Sometimes there isn't even a clear difference like there is for this.

Re: 4o Image Generation

#140

Earlier quoted context omitted.

Are you sure you are using the new 4o image generation? https://imgur.com/a/wGkBa0v

That is an unexpectedly literal definition of "full glass".

Except this is correct in this context. None of existing Diffusion models could, apparently.
Post reply on HN