Live data from Hacker News

4o Image Generation

openai.com

81–90 of 629 posts

Re: 4o Image Generation

#81

OpenAI's livestream of GPT-4o Image Generation shows that it is slowwwwwwwwww (maybe 30 seconds per image, which Sam Altman had to spin "it's slow but the generated images are worth it"). Instead of using a diffusion approach, it appears to be generating the image tokens and decoding them akin to the original DALL-E ( https://openai.com/index/dall-e/ ), which allows for streaming partial generations from top to botto…

As a user, images feel slightly slower but comparable to the previous generation. Given the significant quality improvement, it's a fair trade-off. Overall, it feels snappy, and the value justifies a higher price.

Re: 4o Image Generation

#82
post #61

Earlier quoted context omitted.

Apparently it rolls out today to Plus (which I have). I followed the "Try in ChatGPT" link at the top of the post

It's rolling out to everyone starting today but i'm not sure if everyone has it yet. Does it generate top down for you (picture goes from mostly blurry to clear starting from the top) like in their presentation ?

No it didn't generate like that. Thanks for clarifying. I have updated my original post.

Re: 4o Image Generation

#83

OpenAI's livestream of GPT-4o Image Generation shows that it is slowwwwwwwwww (maybe 30 seconds per image, which Sam Altman had to spin "it's slow but the generated images are worth it"). Instead of using a diffusion approach, it appears to be generating the image tokens and decoding them akin to the original DALL-E ( https://openai.com/index/dall-e/ ), which allows for streaming partial generations from top to botto…

Maybe this is the dialup of the era.

Ha. That's a good analogy.

When I first read the parent comment, I thought, maybe this is a long-term architecture concern...

But your message reminded me that we've been here before.

Re: 4o Image Generation

#84
post #49

OpenAI's livestream of GPT-4o Image Generation shows that it is slowwwwwwwwww (maybe 30 seconds per image, which Sam Altman had to spin "it's slow but the generated images are worth it"). Instead of using a diffusion approach, it appears to be generating the image tokens and decoding them akin to the original DALL-E ( https://openai.com/index/dall-e/ ), which allows for streaming partial generations from top to botto…

LLMs are autoregressive, so they can't be (multi-modality) integrated with diffusion image models, only with autoregressive image models (which generate an image via image tokens). Historically those had lower image fidelity than diffusion models. OpenAI now seems to have solved this problem somehow. More than that, they appear far ahead of any available diffusion model, including Midjourney and Imagen 3. Gemini "int…

Your understanding seems outdated, I think people are referring Gemini native image generation

Re: 4o Image Generation

#85
post #24

This is really impressive, but the "Best of 8" tag on a lot of them really makes me want to see how cherry-picked they are. My three free images had two impressive outputs and one failure.

The high five looks extremely unnatural. Their wrists are aligned, but their fingers aren't, somehow? If that's best of 8, I'd love to see the outtakes.

Agreed. It seems totally unnatural that a couple of nerds high-five awkwardly.

Re: 4o Image Generation

#86

Tried it, the "compise armporressed" and "Pros: made bord reqotons" didn't impress me in the slightest.

Are you sure you were even using the model from the post?

Pressed the "Try in ChatGPT", pasted the first prompt, became thoroughly unimpressed.

Re: 4o Image Generation

#87
post #49

OpenAI's livestream of GPT-4o Image Generation shows that it is slowwwwwwwwww (maybe 30 seconds per image, which Sam Altman had to spin "it's slow but the generated images are worth it"). Instead of using a diffusion approach, it appears to be generating the image tokens and decoding them akin to the original DALL-E ( https://openai.com/index/dall-e/ ), which allows for streaming partial generations from top to botto…

LLMs are autoregressive, so they can't be (multi-modality) integrated with diffusion image models, only with autoregressive image models (which generate an image via image tokens). Historically those had lower image fidelity than diffusion models. OpenAI now seems to have solved this problem somehow. More than that, they appear far ahead of any available diffusion model, including Midjourney and Imagen 3. Gemini "int…

Gemini added their multimodal Flash model to Google AI Studio some time ago. It does not use Imagen via tool, it's uses native capabilities to manipulate images, and it's free to try.

Re: 4o Image Generation

#88
I wanted to use this to generate funny images of myself. Recently I was playing around with Gemini Image Generation to dress myself up as different things. Gemini Image Generation is surprisingly good, although the image quality quickly degrades as you add more changes. Nothing harmful, just silly things like dressing me up as a wizard or other typical RPG roles.

Trying out 4o image generation... It doesn't seem to support this use-case at all? I gave it an image of myself and asked to turn me into a wizard, and it generate something that doesn't look like me in the slightest. A second attempt, I asked to add a wizard hat and it just used python to add a triangle in the middle of my image. I looked at the examples and saw they had a direct image modification where they say "Give this cat a detective hat and a monocle", so I tried that with my own image "Give this human a detective hat and a monocle" and it just gave me this error:

> I wasn't able to generate the modified image because the request didn't follow our content policy. However, I can try another approach—either by applying a filter to stylize the image or guiding you on how to edit it using software like Photoshop or GIMP. Let me know what you'd like to do!

Overall, a very disappointing experience. As another point of comparison, Grok also added image generation capabilities and while the ability to edit existing images is a bit limited and janky, it still manages to overlay the requested transformation on top of the existing image.

Re: 4o Image Generation

#89
post #40

Edit: Please ignore. They hadn't rolled the new model out to my account yet. The announcement blog post is a bit misleading saying you can try it today. -- Comparison with Leonardo.Ai. ChatGPT: https://chatgpt.com/share/67e2fb21-a06c-8008-b297-07681dddee... ChatGPT again (direct one shot): https://chatgpt.com/share/67e2fc44-ecc8-8008-a40f-e1368d306e... ChatGPT again (using word "photorealistic instead of "photo"): ht…

Yeah, its just not good enough. The big labs are way behind what the image focused labs are putting out. Flux and Midjourney are running laps around these guys
Post reply on HN