Earlier quoted context omitted.
It still can't generate a full glass of wine. Even in follow up questions it failed to manipulate the image correctly.
The question remains: why would you generate a full glass of wine? Is that something really that common?
4o Image Generation
221–230 of 629 posts
Re: 4o Image Generation
#222Re: 4o Image Generation
#223If the subject matter is paywalled, I feel that the post should include some explanation of what is newsworthy behind the link.
Re: 4o Image Generation
#224Re: 4o Image Generation
#225Anyone else frightened by this? Seeing meant believing, and now that isnt the case anymore...
Re: 4o Image Generation
#226The garbled text on these things always just makes them basically useless, especially it often text without being told to like previous models.
Re: 4o Image Generation
#227Re: 4o Image Generation
#228I’ve just tried it and oh wow it’s really good. I managed to create a birthday invitation card for my daughter in basically 1-shot, it nailed exactly the elements and style I wanted. Then I asked to retain everything but tweak the text to add more details about the date, venue etc. And it did. I’m in shock. Previous models would not be even halfway there.
Re: 4o Image Generation
#229Earlier quoted context omitted.
> What's important about this new type of image generation that's happening with tokens rather than with diffusion That sounds really interesting. Are there any write-ups how exactly this works?
Would be interested to know as well. As far as I know there is no public information about how this works exactly. This is all I could find: > The system uses an autoregressive approach — generating images sequentially from left to right and top to bottom, similar to how text is written — rather than the diffusion model technique used by most image generators (like DALL-E) that create the entire image at once. Goh sp…
Also wonder if you'd get better results in generating something like blender files and using its engine to render the result.
Re: 4o Image Generation
#230What's important about this new type of image generation that's happening with tokens rather than with diffusion, is that this is effectively reasoning in pixel space. Example: Ask it to draw a notepad with an empty tic-tac-toe, then tell it to make the first move, then you make a move, and so on. You can also do very impressive information-conserving translations, such as changing the drawing style, but also stuff l…
https://chatgpt.com/share/67e32d47-eac0-8011-9118-51b81756ec...