Visual internet content is completely over. Pack it up
So I spent a good few hours investigating the current state of the art a few weeks ago. I would like to generate a collection of images for the art in a video game. It is incredibly difficult to develop an art style, then get the model to generate a collection of different images in that unique art style. I couldn't work out how to do it. I also couldn't work out how to illustrate the same characters or objects in di…
4o Image Generation
371–380 of 629 posts
Re: 4o Image Generation
#372So Google released Gemini 2.5 and one hour later OpenAI comes with this. It’s almost childish at this point.
Re: 4o Image Generation
#373Earlier quoted context omitted.
That's expected with any image generating models because they aren't trained with an alpha channel. It's more pragmatic to pipeline the results to a background removal model. EDIT: It appears GPT-4o is different as there is a video demo dedicated to transparancy.
There's a mod for stable diffusion webui forge/automatic1111/ComfyUI which enables this for all diffusion models (except these closed source ones).
Re: 4o Image Generation
#374Garbage compared to Midjourney. I don't even know why you'd market this. It's takes a minute or more and the results are what I'd say Midjourney looked like 1.5 years ago.
Prompt adherence is far, far ahead of midjourney. FAR.
Re: 4o Image Generation
#375Earlier quoted context omitted.
That's the point. With the old models they all failed to produce a wine glass that is completley to the brim full. Because you can't find that a lot in the data they used for training.
The old models were doing it correct also. There is no one correct way to interpert 'full'. If you go to a wine bar and ask for a full glass of wine, they'll probably interpert that as a double. But you could also interpert it the way a friend would at home, which is about 2-3cm from the rim. Personally I would call a glass of wine filled to the brim 'overfilled', not 'full'.
The prompts (some generated by ChatGPT itself, since it's instructing DALL-E behind the scenes) include phrases like "full to the brim" and "almost spilling over" that are not up to interpretation at all.
Re: 4o Image Generation
#376Earlier quoted context omitted.
https://i.imgur.com/xsFKqsI.png "Draw a picture of a full glass of wine, ie a wine glass which is full to the brim with red wine and almost at the point of spilling over... Zoom out to show the full wine glass, and add a caption to the top which says "HELL YEAH". Keep the wine level of the glass exactly the same."
Can't replicate. Maybe the rollout is staggered? Using Plus from Europe, it's consistently giving me a half full glass.
Re: 4o Image Generation
#377Re: 4o Image Generation
#378Earlier quoted context omitted.
> What's important about this new type of image generation that's happening with tokens rather than with diffusion, is that this is effectively reasoning in pixel space. I do not think that this is correct. Prior to this release, 4o would generate images by calling out to a fully external model (DALL-E). After this release, 4o generates images by calling out to a multi-modal model that was trained alongside it. You c…
I think this is actually correct even if the evidence is not right. See this chat for example: https://chatgpt.com/share/67e355df-9f60-8000-8f36-874f8c9a08...
Re: 4o Image Generation
#379Re: 4o Image Generation
#380Earlier quoted context omitted.
I think this is actually correct even if the evidence is not right. See this chat for example: https://chatgpt.com/share/67e355df-9f60-8000-8f36-874f8c9a08...
Honest question, do you believe something just because the bot tells you that?