am I dumb or every time they release something I can never find out how to actually use it and forget about it. take this for instance I wanted to try out their newton "an infographic explaining newton's prism experiment in great detail" example, but it generated a very bad result but maybe it's because I'm not using the right model? every release of theirs is not really a release, it's like a trailer. right?
4o Image Generation
161–170 of 629 posts
Re: 4o Image Generation
#162What's important about this new type of image generation that's happening with tokens rather than with diffusion, is that this is effectively reasoning in pixel space. Example: Ask it to draw a notepad with an empty tic-tac-toe, then tell it to make the first move, then you make a move, and so on. You can also do very impressive information-conserving translations, such as changing the drawing style, but also stuff l…
That sounds really interesting. Are there any write-ups how exactly this works?
Re: 4o Image Generation
#163Re: 4o Image Generation
#164Earlier quoted context omitted.
I expect the Chinese to have an open source answer for this soon. They haven't been focusing attention on images because the most used image models have been open source. Now they might have a target to beat.
ByteDance has been working on autoregressive image generation for a while (see VAR, NeurIPS 2024 best paper). Traditionally they weren't in the open-source gang though.
Re: 4o Image Generation
#165Re: 4o Image Generation
#166Earlier quoted context omitted.
Is this the same for their gemini-2.0-flash-exp-image-generation model?
No that seems to be indeed a native part of the multimodal Gemini model. I didn't know this existed, it's not available in the normal Gemini interface.
The (no longer, I guess) industry-leading features people actually want are hidden away in some obscure “AI studio” with horrible usability, while the headline Gemini app still often refuses to do anything useful for me. (Disclaimer: I last checked a couple of months ago, after several more of mild amusement/great frustration.)
Re: 4o Image Generation
#167I’ll just be happy with not everything having that over saturated cg/cartoon style that you cant prompt your way out of.
Is that an artifact of the training data? Where are all these original images with that cartoony look that it was trained on?
Re: 4o Image Generation
#168> Introducing 4o Image Generation: [...] our most advanced image generator yet Then google: > Gemini 2.5: Our most intelligent AI model > Introducing Gemini 2.0 | Our most capable AI model yet I could go on forever. I hope this trend dies and apple starts using something effective so all the other companies can start copying a new lexicon.
Re: 4o Image Generation
#169Earlier quoted context omitted.
Are you sure you are using the new 4o image generation? https://imgur.com/a/wGkBa0v
Looks amazing,can you please also create a unconventional image like the clock at 2:35 , I tried it something like this with gemini when some redditor asked it and it failed so wondering if 4o does do it