Live data from Hacker News

4o Image Generation

openai.com

321–330 of 629 posts

Re: 4o Image Generation

#321
The new model in the drop down says something like "4o Create Image (Updated)". It is truly incredible. Far better than any other image generator as far as understanding and following complex prompts.

I was blown away when they showed this many months ago, and found it strange that more people weren't talking about it.

This is much more precise than the Gemini one that just came out recently.

Re: 4o Image Generation

#322
post #288

Visual internet content is completely over. Pack it up

So I spent a good few hours investigating the current state of the art a few weeks ago. I would like to generate a collection of images for the art in a video game.

It is incredibly difficult to develop an art style, then get the model to generate a collection of different images in that unique art style. I couldn't work out how to do it.

I also couldn't work out how to illustrate the same characters or objects in different contexts.

AI seems great for one off images you don't care much about, but when you need images to communicate specific things, I think we are still a long way away.

Re: 4o Image Generation

#323
post #275

Earlier quoted context omitted.

Is it drawing the image from top to bottom very slowly over the course of at least 30 seconds? If not, then you're using DALL-E, not 4o image generation.

This top to bottom drawing – does this tell us anything about the underlying model architecture? AFAIK diffusion models do not work like that. They denoise the full frame over many steps. In the past there used to be attempts to slowly synthetize a picture by predicting the next pixel, but I wasn't aware whether there has been a shift to that kind of architecture within OpenAI.

apparently it's not diffusion, but tokens

Re: 4o Image Generation

#324
post #40

Edit: Please ignore. They hadn't rolled the new model out to my account yet. The announcement blog post is a bit misleading saying you can try it today. -- Comparison with Leonardo.Ai. ChatGPT: https://chatgpt.com/share/67e2fb21-a06c-8008-b297-07681dddee... ChatGPT again (direct one shot): https://chatgpt.com/share/67e2fc44-ecc8-8008-a40f-e1368d306e... ChatGPT again (using word "photorealistic instead of "photo"): ht…

[deleted]

Re: 4o Image Generation

#325
post #51

Earlier quoted context omitted.

Is the ”2D animation style" part you put at the beginning and then changed an attempt to see how well the AI responds to gas lighting?

My bad, I was trying the conversational aspect, but that's not an apples to apples conparison. I have put a direct one shot example in the original post as well.

I'm my test a few months ago, I found that just starting a new prompt would not clear GPT's memory about what I had asked for in previous conversations. You might be stuck with 2D animation style for a while. :)

Re: 4o Image Generation

#326
post #185

Earlier quoted context omitted.

It still can't generate a full glass of wine. Even in follow up questions it failed to manipulate the image correctly.

https://i.imgur.com/xsFKqsI.png "Draw a picture of a full glass of wine, ie a wine glass which is full to the brim with red wine and almost at the point of spilling over... Zoom out to show the full wine glass, and add a caption to the top which says "HELL YEAH". Keep the wine level of the glass exactly the same."

The head of foam on that glass of wine is perfect!

Re: 4o Image Generation

#329
post #230
post #104

What's important about this new type of image generation that's happening with tokens rather than with diffusion, is that this is effectively reasoning in pixel space. Example: Ask it to draw a notepad with an empty tic-tac-toe, then tell it to make the first move, then you make a move, and so on. You can also do very impressive information-conserving translations, such as changing the drawing style, but also stuff l…

Hmmm, I wanted to do that tic tac toe example, and it failed to create a 3x3 grid, instead creating a 5x5 (?) grid with two first moves marked. https://chatgpt.com/share/67e32d47-eac0-8011-9118-51b81756ec...

Your images say "Created with DALL-E", so you have not tried out the new model yet. I think they are gradually rolling it out.
Post reply on HN