Live data from Hacker News

4o Image Generation

openai.com

41–50 of 629 posts

Re: 4o Image Generation

#41

LPT: while the benchmarks don't show it, chatGPT4>4o. It amazes me people use 4o at all. But hey its the brand name and its free. ofc 4.5 is best, but its slow and I am afraid I'm going to hit limits.

OpenAI themselves discourages using GPT-4 outside of legacy applications, in favor of GPT-4o instead (they are shutting down the large output gpt-4-32k variants in a few months). GPT-4 is also an order of magnitude more expensive/slower.

I think both of these points are what sow doubt in some people in the first place because both could be true if GPT-4 was just less profitable to run, not if it was worse in quality. Of course it is actually worse in quality than 4o by any reasonable metric... but I guess not everyone sees it that way.

Re: 4o Image Generation

#42
post #17

The character consistency and UI capabilities seem like they open up a lot of new use cases.

[flagged]

Well I definitely wouldn't say it's vital for humanity. Has anyone actually said that?

Character consistency means that these models could now theoretically illustrate books, as one example.

Generating UIs seems like it would be very helpful for any app design or prototyping.

Re: 4o Image Generation

#43
post #24

This is really impressive, but the "Best of 8" tag on a lot of them really makes me want to see how cherry-picked they are. My three free images had two impressive outputs and one failure.

The high five looks extremely unnatural. Their wrists are aligned, but their fingers aren't, somehow?

If that's best of 8, I'd love to see the outtakes.

Re: 4o Image Generation

#44
post #40

Edit: Please ignore. They hadn't rolled the new model out to my account yet. The announcement blog post is a bit misleading saying you can try it today. -- Comparison with Leonardo.Ai. ChatGPT: https://chatgpt.com/share/67e2fb21-a06c-8008-b297-07681dddee... ChatGPT again (direct one shot): https://chatgpt.com/share/67e2fc44-ecc8-8008-a40f-e1368d306e... ChatGPT again (using word "photorealistic instead of "photo"): ht…

Is the ”2D animation style" part you put at the beginning and then changed an attempt to see how well the AI responds to gas lighting?

Re: 4o Image Generation

#47
post #40

Edit: Please ignore. They hadn't rolled the new model out to my account yet. The announcement blog post is a bit misleading saying you can try it today. -- Comparison with Leonardo.Ai. ChatGPT: https://chatgpt.com/share/67e2fb21-a06c-8008-b297-07681dddee... ChatGPT again (direct one shot): https://chatgpt.com/share/67e2fc44-ecc8-8008-a40f-e1368d306e... ChatGPT again (using word "photorealistic instead of "photo"): ht…

In all fairness you _did_ say 2D animation style

Re: 4o Image Generation

#49

OpenAI's livestream of GPT-4o Image Generation shows that it is slowwwwwwwwww (maybe 30 seconds per image, which Sam Altman had to spin "it's slow but the generated images are worth it"). Instead of using a diffusion approach, it appears to be generating the image tokens and decoding them akin to the original DALL-E ( https://openai.com/index/dall-e/ ), which allows for streaming partial generations from top to botto…

LLMs are autoregressive, so they can't be (multi-modality) integrated with diffusion image models, only with autoregressive image models (which generate an image via image tokens). Historically those had lower image fidelity than diffusion models. OpenAI now seems to have solved this problem somehow. More than that, they appear far ahead of any available diffusion model, including Midjourney and Imagen 3.

Gemini "integrates" Imagen 3 (a diffusion model) only via a tool that Gemini calls internally with the relevant prompt. So it's not a true multimodal integration, as it doesn't benefit from the advanced prompt understanding of the LLM.

Edit: Apparently Gemini also has an experimental native image generation ability.

Re: 4o Image Generation

#50
post #40

Edit: Please ignore. They hadn't rolled the new model out to my account yet. The announcement blog post is a bit misleading saying you can try it today. -- Comparison with Leonardo.Ai. ChatGPT: https://chatgpt.com/share/67e2fb21-a06c-8008-b297-07681dddee... ChatGPT again (direct one shot): https://chatgpt.com/share/67e2fc44-ecc8-8008-a40f-e1368d306e... ChatGPT again (using word "photorealistic instead of "photo"): ht…

What did the prompt look like for Leonard.ai?

I'm curious if you said 2d animation style for both or just for chatgpt.

Edit: Your second version of chatgpt doesn't say photorealistic. Can you share the Leonard.ai prompt?

Post reply on HN