Live data from Hacker News

4o Image Generation

openai.com

271–280 of 629 posts

Re: 4o Image Generation

#271
post #235

Earlier quoted context omitted.

> What's important about this new type of image generation that's happening with tokens rather than with diffusion, is that this is effectively reasoning in pixel space. I do not think that this is correct. Prior to this release, 4o would generate images by calling out to a fully external model (DALL-E). After this release, 4o generates images by calling out to a multi-modal model that was trained alongside it. You c…

Models are famously good at understanding themselves.

I hope you're joking. Sometimes they don't even know which company developed them. E.g. DeepSeek was claiming it was developed by OpenAI.

Re: 4o Image Generation

#273

Earlier quoted context omitted.

I tried and while the clock it generated was very well done and high quality, it showed the time as the analog clock default of 10:10.

The problem now is we don't know if people mistake dall-e for the new multimodal gpt4o output, they really should've made that clearer.

I’m using 4o and it gets time wrong a decent chunk but doesn’t get anything else in the prompt incorrect. I asked for the clock to be 4:30 but got 10:10. OpenAI pro account.

Re: 4o Image Generation

#274
post #104

What's important about this new type of image generation that's happening with tokens rather than with diffusion, is that this is effectively reasoning in pixel space. Example: Ask it to draw a notepad with an empty tic-tac-toe, then tell it to make the first move, then you make a move, and so on. You can also do very impressive information-conserving translations, such as changing the drawing style, but also stuff l…

> writing the code to reproduce it

I'm super excited for all the free money and data our new AI written apps will be giving away.

Re: 4o Image Generation

#275
post #194

Earlier quoted context omitted.

Can't replicate. Maybe the rollout is staggered? Using Plus from Europe, it's consistently giving me a half full glass.

Is it drawing the image from top to bottom very slowly over the course of at least 30 seconds? If not, then you're using DALL-E, not 4o image generation.

This top to bottom drawing – does this tell us anything about the underlying model architecture? AFAIK diffusion models do not work like that. They denoise the full frame over many steps. In the past there used to be attempts to slowly synthetize a picture by predicting the next pixel, but I wasn't aware whether there has been a shift to that kind of architecture within OpenAI.

Re: 4o Image Generation

#276
post #254
post #233

Earlier quoted context omitted.

We're in the middle of a massive and unprecedented boom in AI capabilities. It is hard to be upset about this phrasing - it is literally true and extremely accurate.

If that's so then there's no need to be hyperbolic about it. Why would they publish a model that is not their most advanced model?

(Shrug) It's common for less-than-foundation-level models to be released every so often. This is done in order to provide new options, features, pricing, service levels, APIs or whatever that aren't yet incorporated into the main model, or that are never intended to be.

Just a consequence of how much time and money it takes to train a new foundation model. It's not going to happen every other week. When it does, it is reasonable to announce it with "Announcing our most powerful model yet."

Re: 4o Image Generation

#277
post #254
post #233

Earlier quoted context omitted.

We're in the middle of a massive and unprecedented boom in AI capabilities. It is hard to be upset about this phrasing - it is literally true and extremely accurate.

If that's so then there's no need to be hyperbolic about it. Why would they publish a model that is not their most advanced model?

Most things aren't in a massive boom and most people aren't that involved in AI. This is a rare example of great communication in marketing - they're telling people who might not be across this field what is going on.

> Why would they publish a model that is not their most advanced model?

I dunno, I'm not sitting in the OpenAI meetings. That is why they need to tell us what they are doing - it is easy to imagine them releasing something that isn't their best model ever and so they clarify that this is, in fact, the new hotness.

Re: 4o Image Generation

#278

Earlier quoted context omitted.

Are you sure you are using the new 4o image generation? https://imgur.com/a/wGkBa0v

Can you do this with the prompt of a cow jumping over the moon? I can’t ever seem to get it to make the cow appear to be above the moon. Always literally covering it or to the side etc.

Here you go: https://imgur.com/a/QJlj4I9

Re: 4o Image Generation

#279
post #230
post #104

What's important about this new type of image generation that's happening with tokens rather than with diffusion, is that this is effectively reasoning in pixel space. Example: Ask it to draw a notepad with an empty tic-tac-toe, then tell it to make the first move, then you make a move, and so on. You can also do very impressive information-conserving translations, such as changing the drawing style, but also stuff l…

Hmmm, I wanted to do that tic tac toe example, and it failed to create a 3x3 grid, instead creating a 5x5 (?) grid with two first moves marked. https://chatgpt.com/share/67e32d47-eac0-8011-9118-51b81756ec...

Tried it myself on the new model, worked out pretty well:

https://chatgpt.com/share/67e34558-5244-8004-933a-23896c738b...

Post reply on HN