Live data from Hacker News

4o Image Generation

openai.com

111–120 of 629 posts

Re: 4o Image Generation

#111
post #100

Is there any way to see whether a given prompt was serviced by 4o or Dall-E? Currently, my prompts seem to be going to the latter still, based on e.g. my source image being very obviously looped through a verbal image description and back to an image, compared to gemini-2.0-flash-exp-image-generation. A friend with a Plus plan has been getting responses from either. The long-term plan seems to be to move to 4o comple…

4o generates top down (picture goes from mostly blurry to clear starting from the top). If it's not generating like that for you then you don't have it yet.

Re: 4o Image Generation

#113
I think the biggest problem I still see is the models awareness of the images it generated itself.

The glaring issue for the older image generators is how it would proudly proclaim to have presented an image with a description that has almost no relation to the image it actually provided.

I'm not sure if this update improves on this aspect. It may create the illusion of awareness of the picture by having better prompt adherence.

Re: 4o Image Generation

#114

I wanted to use this to generate funny images of myself. Recently I was playing around with Gemini Image Generation to dress myself up as different things. Gemini Image Generation is surprisingly good, although the image quality quickly degrades as you add more changes. Nothing harmful, just silly things like dressing me up as a wizard or other typical RPG roles. Trying out 4o image generation... It doesn't seem to s…

It's not actually out for everyone yet. You can tell by the generation style. 4o generates top down (picture goes from mostly blurry to clear starting from the top).

Re: 4o Image Generation

#115
post #105
post #104

What's important about this new type of image generation that's happening with tokens rather than with diffusion, is that this is effectively reasoning in pixel space. Example: Ask it to draw a notepad with an empty tic-tac-toe, then tell it to make the first move, then you make a move, and so on. You can also do very impressive information-conserving translations, such as changing the drawing style, but also stuff l…

Pretty sure the modern Gemini image models can already do token based image generation/editing and are significantly better and faster.

It's faster but it's definitely not better than what's being showcased here. The quality of Flash 2 Image gens are generally pretty meh.

Re: 4o Image Generation

#116
post #13

Earlier quoted context omitted.

Flux 1.1 Pro has good prompt adherence, but some of these (admittingly cherry-picked) GPT-4o generated image demos are beyond what you would get with Flux without a lot of iteration, particularly the large paragraphs of text. I'm excited to see what a Flux 2 can do if it can actually use a modern text encoder.

Structural editing and control nets are much more powerful than text prompting alone. The image generators used by creatives will not be text-first. "Dragon with brown leathery scales with an elephant texture and 10% reflectivity positioned three degrees under the mountain, which is approximately 250 meters taller than the next peak, ..." is not how you design. Creative work is not 100% dice rolling in a crude and in…

It can do in-context learning from images you upload. So you can just upload a depth map or mark up an image with the locations of edits you want and it should be able to handle that. I guess my point is that since its the same model that understands how to see images and how to generate them you aren't restricted from interacting with it via text only.

Re: 4o Image Generation

#117

Earlier quoted context omitted.

Prompt adherence and additional tricks such as ControlNet/ComfyUI pipelines are not mutually exclusive. Both are very important to get good image generation results.

It is when it's kept behind an API. You cannot use Controlnet/ComfyUI and especially not the best stuff like regional prompting with this model. You can't do it with Gemini, and that's by design because otherwise coomers are going to generate 999999 anime waifus like they do on Civit.ai.

That just elicits a cheeky refusal I'm afraid:

"""

That's a fun idea—but generating an image with 999,999 anime waifus in it isn't technically possible due to visual and processing limits. But we can get creative.

Want me to generate:

1. A massive crowd of anime waifus (like a big collage or crowd scene)?

2. A stylized representation of “999999 anime waifus” (maybe with a few in focus and the rest as silhouettes or a sea of colors)?

3. A single waifu with a visual reference to the number 999999 (like a title, emblem, or digital counter in the background)?

Let me know your vibe—epic, funny, serious, chaotic?

"""

Re: 4o Image Generation

#118

Is it live yet? Have been trying it out and am still getting poor results on text generation.

This is OpenAI's bread and butter - announce something as though it's being launched and then proceed to slowly roll it out after a couple of days.

Truly infuriating, especially when it's something like this that makes it tough to tell if the feature is even enabled.

Re: 4o Image Generation

#120

am I dumb or every time they release something I can never find out how to actually use it and forget about it. take this for instance I wanted to try out their newton "an infographic explaining newton's prism experiment in great detail" example, but it generated a very bad result but maybe it's because I'm not using the right model? every release of theirs is not really a release, it's like a trailer. right?

You're not dumb. They do this for nearly every single major release. I can't really understand why considering it generates negative sentiment about the release, but it's something to be expected from OpenAI at this point.
Post reply on HN