Live data from Hacker News

4o Image Generation

openai.com

91–100 of 629 posts

Re: 4o Image Generation

#91

This works great for many purposes. One area where it does not work well at all is modifying photographs of people's faces.* Completely fumbles if you take a selfie and ask it to modify your shirt, for example. * = unless the people are in the training set

> We’re aware of a bug where the model struggles with maintaining consistency of edits to faces from user uploads but expect this to be fixed within the week.

Sounds like it may be a safety thing that's still getting figured out

Re: 4o Image Generation

#92

This works great for many purposes. One area where it does not work well at all is modifying photographs of people's faces.* Completely fumbles if you take a selfie and ask it to modify your shirt, for example. * = unless the people are in the training set

It just doesn't have that kind of image editing capability. Maybe people just assume it does because Google's similar model has it. But did OpenAI claim it could edit images?

Re: 4o Image Generation

#93
post #72
post #70

Earlier quoted context omitted.

What's the problem?

It's a nitpick about the repetitive phrasing for announcements : Our most yet|ever.

Speaking as someone who'd love to not speak that way in my own marketing - it's an unfortunate necessity in a world where people will give you literal milliseconds of their time. Marketing isn't there to tell you about the thing, it's there to get you to want to know more about the thing.

Re: 4o Image Generation

#94
post #57

> Introducing 4o Image Generation: [...] our most advanced image generator yet Then google: > Gemini 2.5: Our most intelligent AI model > Introducing Gemini 2.0 | Our most capable AI model yet I could go on forever. I hope this trend dies and apple starts using something effective so all the other companies can start copying a new lexicon.

Has post-Jobs Apple ever come up with anything that would warrant this hope?

Every iPhone is their best iPhone yet

Re: 4o Image Generation

#95

I’ll just be happy with not everything having that over saturated cg/cartoon style that you cant prompt your way out of.

Is that an artifact of the training data? Where are all these original images with that cartoony look that it was trained on?

Re: 4o Image Generation

#96
post #57

> Introducing 4o Image Generation: [...] our most advanced image generator yet Then google: > Gemini 2.5: Our most intelligent AI model > Introducing Gemini 2.0 | Our most capable AI model yet I could go on forever. I hope this trend dies and apple starts using something effective so all the other companies can start copying a new lexicon.

Has post-Jobs Apple ever come up with anything that would warrant this hope?

No, but I think they stopped with "our most" (since all other brainless corps adopted it) and just connect adjectives with dots.

Hotwheels: Fast. Furious. Spectacular.

Re: 4o Image Generation

#97
post #49

OpenAI's livestream of GPT-4o Image Generation shows that it is slowwwwwwwwww (maybe 30 seconds per image, which Sam Altman had to spin "it's slow but the generated images are worth it"). Instead of using a diffusion approach, it appears to be generating the image tokens and decoding them akin to the original DALL-E ( https://openai.com/index/dall-e/ ), which allows for streaming partial generations from top to botto…

LLMs are autoregressive, so they can't be (multi-modality) integrated with diffusion image models, only with autoregressive image models (which generate an image via image tokens). Historically those had lower image fidelity than diffusion models. OpenAI now seems to have solved this problem somehow. More than that, they appear far ahead of any available diffusion model, including Midjourney and Imagen 3. Gemini "int…

Meta has experimented with a hybrid mode, where the LLM uses autoregressive mode for text, but within a set of delimiters will switch to diffusion mode to generate images. In principle it's the best of both worlds.

Re: 4o Image Generation

#98
To quote myself from a comment on sora:

Iterations are the missing link. With ChatGPT, you can iteratively improve text (e.g., "make it shorter," "mention xyz"). However, for pictures (and video), this functionality is not yet available. If you could prompt iteratively (e.g., "generate a red car in the sunset," "make it a muscle car," "place it on a hill," "show it from the side so the sun shines through the windshield"), the tools would become exponentially more useful.

I‘m looking forward to try this out and see if I was right. Unfortunately it’s not yet available for me.

Re: 4o Image Generation

#99
The periodic table poster under "High binding problems" is billed as evidence of model limitations, but I wonder if it just suggests that 4o is a fan of "Look Around You".

Re: 4o Image Generation

#100
Is there any way to see whether a given prompt was serviced by 4o or Dall-E?

Currently, my prompts seem to be going to the latter still, based on e.g. my source image being very obviously looped through a verbal image description and back to an image, compared to gemini-2.0-flash-exp-image-generation. A friend with a Plus plan has been getting responses from either.

The long-term plan seems to be to move to 4o completely and move Dall-E to its own tab, though, so maybe that problem will resolve itself before too long.

Post reply on HN