Live data from Hacker News

4o Image Generation

openai.com

201–210 of 629 posts

Re: 4o Image Generation

#201
post #98

To quote myself from a comment on sora: Iterations are the missing link. With ChatGPT, you can iteratively improve text (e.g., "make it shorter," "mention xyz"). However, for pictures (and video), this functionality is not yet available. If you could prompt iteratively (e.g., "generate a red car in the sunset," "make it a muscle car," "place it on a hill," "show it from the side so the sun shines through the windshie…

You can do that with Gemini's image model, flash 2.0 (image generation) exp.[1] It's not perfect but it does mostly maintain likeness between generations. [1] https://aistudio.google.com/prompts/new_chat

Whisk I think is possibly the best at it. No idea what it uses under the hood though.

https://labs.google/fx/tools/whisk

Re: 4o Image Generation

#202

Earlier quoted context omitted.

Looks amazing,can you please also create a unconventional image like the clock at 2:35 , I tried it something like this with gemini when some redditor asked it and it failed so wondering if 4o does do it

I tried and it failed repeatedly (like actual error messages): > It looks like there was an error when trying to generate the updated image of the clock showing 5:03. I wasn’t able to create it. If you’d like, you can try again by rephrasing or repeating the request. A few times it did generate an image but it never showed the right time. It would frequently show 10:10 for instance.

If it tried and failed repeatedly, then it was prompting DALL-E, looking at the results, then prompting DALL-E again, not doing direct image generation.

Re: 4o Image Generation

#203
post #100

Is there any way to see whether a given prompt was serviced by 4o or Dall-E? Currently, my prompts seem to be going to the latter still, based on e.g. my source image being very obviously looped through a verbal image description and back to an image, compared to gemini-2.0-flash-exp-image-generation. A friend with a Plus plan has been getting responses from either. The long-term plan seems to be to move to 4o comple…

I've generated (and downloaded) a couple of images. All filenames start with `DALL·E`, so I guess that's a safe way to tell how the images were generated.

Re: 4o Image Generation

#204
post #187

Earlier quoted context omitted.

I created a shared link just now, was not presented with any such warning, and have the same problem with the image not showing up: https://chatgpt.com/share/67e319dd-bd08-8013-8f9b-6f5140137f...

Interesting. I see this: https://imgur.com/a/QNWeEoZ

Aha! I see different messages in the Android app vs. web app.

In the web app I see:

Your name, custom instructions, and any messages you add after sharing stay private. Learn more

Re: 4o Image Generation

#207
post #199
post #104

What's important about this new type of image generation that's happening with tokens rather than with diffusion, is that this is effectively reasoning in pixel space. Example: Ask it to draw a notepad with an empty tic-tac-toe, then tell it to make the first move, then you make a move, and so on. You can also do very impressive information-conserving translations, such as changing the drawing style, but also stuff l…

That's very interesting. I would have assumed that 4o is internally using a single seed for the entire conversation, or something analogous to that, to control randomness across image generation requests. Can you share the technical name for this reasoning process so I could look up research about it?

multimodal chain of thought / generation of thought

Nobody has really decided on a name.

Also chain of thought is somewhat different from chain of thought reasoning so mb throw in multimodal chain of thought reasoning

Re: 4o Image Generation

#208
post #205

Anyone else frightened by this? Seeing meant believing, and now that isnt the case anymore...

Look closer at the fingers. These models still don’t have a firm handle on them. The right elbow on the second picture also doesn’t quite look anatomically possible.

Re: 4o Image Generation

#209

Earlier quoted context omitted.

Looks amazing,can you please also create a unconventional image like the clock at 2:35 , I tried it something like this with gemini when some redditor asked it and it failed so wondering if 4o does do it

I tried and while the clock it generated was very well done and high quality, it showed the time as the analog clock default of 10:10.

The problem now is we don't know if people mistake dall-e for the new multimodal gpt4o output, they really should've made that clearer.

Re: 4o Image Generation

#210

Earlier quoted context omitted.

Speaking as someone who'd love to not speak that way in my own marketing - it's an unfortunate necessity in a world where people will give you literal milliseconds of their time. Marketing isn't there to tell you about the thing, it's there to get you to want to know more about the thing.

A term for people giving only milliseconds of their attention is: uninterested people. If I’m not looking for a project planner, or interested in the space, there’s no wording that can make me stay on an announcement for one. If I am, you can be sure I’m going to read the whole feature page.

Idealistic and wrong, marketing does work in a lot of cases and that's why everybody does it
Post reply on HN