Can anyone tell me when this will be available in the API? Or is it already available? I couldn't find anything on the pricing page.
4o Image Generation
501–510 of 629 posts
Re: 4o Image Generation
#502Earlier quoted context omitted.
We're in the middle of a massive and unprecedented boom in AI capabilities. It is hard to be upset about this phrasing - it is literally true and extremely accurate.
If that's so then there's no need to be hyperbolic about it. Why would they publish a model that is not their most advanced model?
And no, not all models are intended to push the frontier in terms of benchmark performance, some are just fast and cheap.
Re: 4o Image Generation
#503Earlier quoted context omitted.
share prompt minus identifying details?
> Draw a birthday invitation for a 4 year old girl [name here]. It should be whimsical, look like its hand-drawn with little drawings on the sides of stuff like dinosaurs, flowers, hearts, cats. The background should be light and the foreground elements should be red, pink, orange and blue. Then I asked for some changes: > That's almost perfect! Retain this style and the elements, but adjust the text to read: > [refi…
Re: 4o Image Generation
#504Earlier quoted context omitted.
>You can also do very impressive information-conserving translations, such as changing the drawing style, but also stuff like "change day to night", or "put a hat on him", and so forth. You can do that with diffusion, too. Just lock the parameters in ComfyUi.
Yeah I wasn’t very imaginative in my examples, with 4o you can also perform transformations like “rotate the camera 10 degrees to the left” which would be hard without a specialized model. Basically you can run arbitrary functions on the exact image contents but in latent space.
Re: 4o Image Generation
#505It's very impressive. It feels like the text is a bit of a hack where they're somehow rendering the text separately and interpolating it into the image. Not always, I got it to render calligraphy with flourishes, but only for a handful of words. For example, I asked it to render a few lines of text on a medieval scroll, and it basically looked like a picture of a gothic font written onto a background image of a scrol…
Re: 4o Image Generation
#506Re: 4o Image Generation
#507What's important about this new type of image generation that's happening with tokens rather than with diffusion, is that this is effectively reasoning in pixel space. Example: Ask it to draw a notepad with an empty tic-tac-toe, then tell it to make the first move, then you make a move, and so on. You can also do very impressive information-conserving translations, such as changing the drawing style, but also stuff l…
> truly generative UI, where the model produces the next frame of the app Please sir step away from the keyboard now! That is an absurd proposition and I hope I never get to use an app that dreams of the next frame. Apps are buggy as they are, I don't need every single action to be interpreted by LLM. An existing example of this is that AI Minecraft demo and it's a literal nightmare.
> “DLSS Multi Frame Generation generates up to three additional frames per traditionally rendered frame, working in unison with the complete suite of DLSS technologies to multiply frame rates by up to 8X over traditional brute-force rendering. This massive performance improvement on GeForce RTX 5090 graphics cards unlocks stunning 4K 240 FPS fully ray-traced gaming.”
Re: 4o Image Generation
#508Earlier quoted context omitted.
> truly generative UI, where the model produces the next frame of the app Please sir step away from the keyboard now! That is an absurd proposition and I hope I never get to use an app that dreams of the next frame. Apps are buggy as they are, I don't need every single action to be interpreted by LLM. An existing example of this is that AI Minecraft demo and it's a literal nightmare.
This argument could be made for every level of abstraction we've added to software so far... yet here we are commenting about it from our buggy apps!
No. Not at all. Those levels of abstractions – whether good, bad, everything in between – were fully understood through-and-through by humans. Having an LLM somewhere in the stack of abstractions is radically different, and radically stupid.
Re: 4o Image Generation
#509I think it is too biased to use heuristics discovered in the first response to apply the same level of compute to subsequent requests.
It makes me kind of want to rewrite an interface that builds appropriate context and starts new chats for every request issued..
Re: 4o Image Generation
#510Can it draw the notorious glass of wine filled to the brim yet?