I was blown away when they showed this many months ago, and found it strange that more people weren't talking about it.
This is much more precise than the Gemini one that just came out recently.
321–330 of 629 posts
I was blown away when they showed this many months ago, and found it strange that more people weren't talking about it.
This is much more precise than the Gemini one that just came out recently.
Visual internet content is completely over. Pack it up
It is incredibly difficult to develop an art style, then get the model to generate a collection of different images in that unique art style. I couldn't work out how to do it.
I also couldn't work out how to illustrate the same characters or objects in different contexts.
AI seems great for one off images you don't care much about, but when you need images to communicate specific things, I think we are still a long way away.
Earlier quoted context omitted.
Is it drawing the image from top to bottom very slowly over the course of at least 30 seconds? If not, then you're using DALL-E, not 4o image generation.
This top to bottom drawing – does this tell us anything about the underlying model architecture? AFAIK diffusion models do not work like that. They denoise the full frame over many steps. In the past there used to be attempts to slowly synthetize a picture by predicting the next pixel, but I wasn't aware whether there has been a shift to that kind of architecture within OpenAI.
Edit: Please ignore. They hadn't rolled the new model out to my account yet. The announcement blog post is a bit misleading saying you can try it today. -- Comparison with Leonardo.Ai. ChatGPT: https://chatgpt.com/share/67e2fb21-a06c-8008-b297-07681dddee... ChatGPT again (direct one shot): https://chatgpt.com/share/67e2fc44-ecc8-8008-a40f-e1368d306e... ChatGPT again (using word "photorealistic instead of "photo"): ht…
Earlier quoted context omitted.
Is the ”2D animation style" part you put at the beginning and then changed an attempt to see how well the AI responds to gas lighting?
My bad, I was trying the conversational aspect, but that's not an apples to apples conparison. I have put a direct one shot example in the original post as well.
Earlier quoted context omitted.
It still can't generate a full glass of wine. Even in follow up questions it failed to manipulate the image correctly.
https://i.imgur.com/xsFKqsI.png "Draw a picture of a full glass of wine, ie a wine glass which is full to the brim with red wine and almost at the point of spilling over... Zoom out to show the full wine glass, and add a caption to the top which says "HELL YEAH". Keep the wine level of the glass exactly the same."
Garbage compared to Midjourney. I don't even know why you'd market this. It's takes a minute or more and the results are what I'd say Midjourney looked like 1.5 years ago.
Otherwise impressive.
What's important about this new type of image generation that's happening with tokens rather than with diffusion, is that this is effectively reasoning in pixel space. Example: Ask it to draw a notepad with an empty tic-tac-toe, then tell it to make the first move, then you make a move, and so on. You can also do very impressive information-conserving translations, such as changing the drawing style, but also stuff l…
Hmmm, I wanted to do that tic tac toe example, and it failed to create a 3x3 grid, instead creating a 5x5 (?) grid with two first moves marked. https://chatgpt.com/share/67e32d47-eac0-8011-9118-51b81756ec...