Earlier quoted context omitted.
> What's important about this new type of image generation that's happening with tokens rather than with diffusion, is that this is effectively reasoning in pixel space. I do not think that this is correct. Prior to this release, 4o would generate images by calling out to a fully external model (DALL-E). After this release, 4o generates images by calling out to a multi-modal model that was trained alongside it. You c…
Models are famously good at understanding themselves.
4o Image Generation
271–280 of 629 posts
Re: 4o Image Generation
#272Re: 4o Image Generation
#273Earlier quoted context omitted.
I tried and while the clock it generated was very well done and high quality, it showed the time as the analog clock default of 10:10.
The problem now is we don't know if people mistake dall-e for the new multimodal gpt4o output, they really should've made that clearer.
Re: 4o Image Generation
#274What's important about this new type of image generation that's happening with tokens rather than with diffusion, is that this is effectively reasoning in pixel space. Example: Ask it to draw a notepad with an empty tic-tac-toe, then tell it to make the first move, then you make a move, and so on. You can also do very impressive information-conserving translations, such as changing the drawing style, but also stuff l…
I'm super excited for all the free money and data our new AI written apps will be giving away.
Re: 4o Image Generation
#275Earlier quoted context omitted.
Can't replicate. Maybe the rollout is staggered? Using Plus from Europe, it's consistently giving me a half full glass.
Is it drawing the image from top to bottom very slowly over the course of at least 30 seconds? If not, then you're using DALL-E, not 4o image generation.
Re: 4o Image Generation
#276Earlier quoted context omitted.
We're in the middle of a massive and unprecedented boom in AI capabilities. It is hard to be upset about this phrasing - it is literally true and extremely accurate.
If that's so then there's no need to be hyperbolic about it. Why would they publish a model that is not their most advanced model?
Just a consequence of how much time and money it takes to train a new foundation model. It's not going to happen every other week. When it does, it is reasonable to announce it with "Announcing our most powerful model yet."
Re: 4o Image Generation
#277Earlier quoted context omitted.
We're in the middle of a massive and unprecedented boom in AI capabilities. It is hard to be upset about this phrasing - it is literally true and extremely accurate.
If that's so then there's no need to be hyperbolic about it. Why would they publish a model that is not their most advanced model?
> Why would they publish a model that is not their most advanced model?
I dunno, I'm not sitting in the OpenAI meetings. That is why they need to tell us what they are doing - it is easy to imagine them releasing something that isn't their best model ever and so they clarify that this is, in fact, the new hotness.
Re: 4o Image Generation
#278Earlier quoted context omitted.
Are you sure you are using the new 4o image generation? https://imgur.com/a/wGkBa0v
Can you do this with the prompt of a cow jumping over the moon? I can’t ever seem to get it to make the cow appear to be above the moon. Always literally covering it or to the side etc.
Re: 4o Image Generation
#279What's important about this new type of image generation that's happening with tokens rather than with diffusion, is that this is effectively reasoning in pixel space. Example: Ask it to draw a notepad with an empty tic-tac-toe, then tell it to make the first move, then you make a move, and so on. You can also do very impressive information-conserving translations, such as changing the drawing style, but also stuff l…
Hmmm, I wanted to do that tic tac toe example, and it failed to create a 3x3 grid, instead creating a 5x5 (?) grid with two first moves marked. https://chatgpt.com/share/67e32d47-eac0-8011-9118-51b81756ec...
https://chatgpt.com/share/67e34558-5244-8004-933a-23896c738b...
Re: 4o Image Generation
#280I try this on every new generation: Generate a photo of a lake taken by a mobile phone camera. No hands or phones in the photo, just the lake. The hand holding a phone is always there :D