Live data from Hacker News

4o Image Generation

openai.com

151–160 of 629 posts

Re: 4o Image Generation

#151
post #104

What's important about this new type of image generation that's happening with tokens rather than with diffusion, is that this is effectively reasoning in pixel space. Example: Ask it to draw a notepad with an empty tic-tac-toe, then tell it to make the first move, then you make a move, and so on. You can also do very impressive information-conserving translations, such as changing the drawing style, but also stuff l…

> Finally, once these models become faster, you can imagine a truly generative UI, where the model produces the next frame of the app you are using based on events sent to the LLM

With current GPU technology, this system would need its own Dyson sphere.

Re: 4o Image Generation

#153
post #80
post #49

Earlier quoted context omitted.

LLMs are autoregressive, so they can't be (multi-modality) integrated with diffusion image models, only with autoregressive image models (which generate an image via image tokens). Historically those had lower image fidelity than diffusion models. OpenAI now seems to have solved this problem somehow. More than that, they appear far ahead of any available diffusion model, including Midjourney and Imagen 3. Gemini "int…

I expect the Chinese to have an open source answer for this soon. They haven't been focusing attention on images because the most used image models have been open source. Now they might have a target to beat.

ByteDance has been working on autoregressive image generation for a while (see VAR, NeurIPS 2024 best paper). Traditionally they weren't in the open-source gang though.

Re: 4o Image Generation

#156

Earlier quoted context omitted.

That is an unexpectedly literal definition of "full glass".

That's the point. With the old models they all failed to produce a wine glass that is completley to the brim full. Because you can't find that a lot in the data they used for training.

Imagine if they just actually trained the model on a bunch of photographs of a full glass of wine, knowing of this litmus test

Re: 4o Image Generation

#157
post #124
post #105

Earlier quoted context omitted.

Pretty sure the modern Gemini image models can already do token based image generation/editing and are significantly better and faster.

Yeah Gemini has had this for a few weeks, but much lower resolution. Not saying 4o is perfect, but my first few images with it are much more impressive than my first few images with Gemini.

[deleted]

Re: 4o Image Generation

#158
post #63
post #49

Earlier quoted context omitted.

LLMs are autoregressive, so they can't be (multi-modality) integrated with diffusion image models, only with autoregressive image models (which generate an image via image tokens). Historically those had lower image fidelity than diffusion models. OpenAI now seems to have solved this problem somehow. More than that, they appear far ahead of any available diffusion model, including Midjourney and Imagen 3. Gemini "int…

Is this the same for their gemini-2.0-flash-exp-image-generation model?

No that seems to be indeed a native part of the multimodal Gemini model. I didn't know this existed, it's not available in the normal Gemini interface.

Re: 4o Image Generation

#159
post #104

What's important about this new type of image generation that's happening with tokens rather than with diffusion, is that this is effectively reasoning in pixel space. Example: Ask it to draw a notepad with an empty tic-tac-toe, then tell it to make the first move, then you make a move, and so on. You can also do very impressive information-conserving translations, such as changing the drawing style, but also stuff l…

It still can't generate a full glass of wine. Even in follow up questions it failed to manipulate the image correctly.

I don't buy the meme or w/e that they can't produce an image with the full glass of wine. Just takes a little prompt engineering.

Using Dall-e / old model without too much effort (I'd call this "full".)

https://imgur.com/a/J2bCwYh

Post reply on HN