Live data from Hacker News

4o Image Generation

openai.com

251–260 of 629 posts

Re: 4o Image Generation

#252
> we see the photographer's reflection

Am I the only one immediately looking past the amazing text generation, the excellent direction following, the wonderful reflection, and screaming inside my head, "That's not how reflection works!"

I know it's super nitpicky when it's so obviously a leap forward on multiple other metrics, but still, that reflection just ain't right.

Re: 4o Image Generation

#253
post #63

Earlier quoted context omitted.

Is this the same for their gemini-2.0-flash-exp-image-generation model?

No that seems to be indeed a native part of the multimodal Gemini model. I didn't know this existed, it's not available in the normal Gemini interface.

That's pretty disappointing, it has been out for a while, and we still get top comments like (https://news.ycombinator.com/item?id=43475043) where people clearly think native image generation capability is new. Where do you usually get your updates from for this kind of thing?

Re: 4o Image Generation

#254
post #233
post #57

> Introducing 4o Image Generation: [...] our most advanced image generator yet Then google: > Gemini 2.5: Our most intelligent AI model > Introducing Gemini 2.0 | Our most capable AI model yet I could go on forever. I hope this trend dies and apple starts using something effective so all the other companies can start copying a new lexicon.

We're in the middle of a massive and unprecedented boom in AI capabilities. It is hard to be upset about this phrasing - it is literally true and extremely accurate.

If that's so then there's no need to be hyperbolic about it. Why would they publish a model that is not their most advanced model?

Re: 4o Image Generation

#255
I try this on every new generation:

Generate a photo of a lake taken by a mobile phone camera. No hands or phones in the photo, just the lake.

The hand holding a phone is always there :D

Re: 4o Image Generation

#256

Earlier quoted context omitted.

Imagine if they just actually trained the model on a bunch of photographs of a full glass of wine, knowing of this litmus test

I obviously have no idea if they added real or synthetic data to the training set specifically regarding the full-to-the-brim wineglass test, but I fully expect that this prompt is now compromised in the sense that because it is being discussed in the public sphere, it's has inherently become part of the test suite. Remember the old internet adage that the fastest way to get a correct answer online is to post an inco…

Humans don’t train on the entire contents of the Internet, so i’d wager that they do learn differently

Re: 4o Image Generation

#257
post #194
post #185

Earlier quoted context omitted.

https://i.imgur.com/xsFKqsI.png "Draw a picture of a full glass of wine, ie a wine glass which is full to the brim with red wine and almost at the point of spilling over... Zoom out to show the full wine glass, and add a caption to the top which says "HELL YEAH". Keep the wine level of the glass exactly the same."

Can't replicate. Maybe the rollout is staggered? Using Plus from Europe, it's consistently giving me a half full glass.

You might still be on DALL-E. My account is if you use ChatGPT.

I switched over to the sora.com domain and now I have access to it.

Re: 4o Image Generation

#258
post #194
post #185

Earlier quoted context omitted.

https://i.imgur.com/xsFKqsI.png "Draw a picture of a full glass of wine, ie a wine glass which is full to the brim with red wine and almost at the point of spilling over... Zoom out to show the full wine glass, and add a caption to the top which says "HELL YEAH". Keep the wine level of the glass exactly the same."

Can't replicate. Maybe the rollout is staggered? Using Plus from Europe, it's consistently giving me a half full glass.

[deleted]

Re: 4o Image Generation

#259
post #235
post #104

What's important about this new type of image generation that's happening with tokens rather than with diffusion, is that this is effectively reasoning in pixel space. Example: Ask it to draw a notepad with an empty tic-tac-toe, then tell it to make the first move, then you make a move, and so on. You can also do very impressive information-conserving translations, such as changing the drawing style, but also stuff l…

> What's important about this new type of image generation that's happening with tokens rather than with diffusion, is that this is effectively reasoning in pixel space. I do not think that this is correct. Prior to this release, 4o would generate images by calling out to a fully external model (DALL-E). After this release, 4o generates images by calling out to a multi-modal model that was trained alongside it. You c…

You're incorrect. 4o was not trained on knowledge of itself so literally can't tell you that. What 4o is doing isn't even new either, Gemini 2.0 has the same capability.
Post reply on HN