Live data from Hacker News

4o Image Generation

openai.com

71–80 of 629 posts

Re: 4o Image Generation

#71
post #38

This works great for many purposes. One area where it does not work well at all is modifying photographs of people's faces.* Completely fumbles if you take a selfie and ask it to modify your shirt, for example. * = unless the people are in the training set

That's to be expected, no? It's a usian product so it will be a disappointment in all areas where things could get lewd.

What is usian? Never heard of that.

Re: 4o Image Generation

#72
post #70
post #57

> Introducing 4o Image Generation: [...] our most advanced image generator yet Then google: > Gemini 2.5: Our most intelligent AI model > Introducing Gemini 2.0 | Our most capable AI model yet I could go on forever. I hope this trend dies and apple starts using something effective so all the other companies can start copying a new lexicon.

What's the problem?

It's a nitpick about the repetitive phrasing for announcements

: Our most yet|ever.

Re: 4o Image Generation

#73
post #28

Still seems to have problems with transparent backgrounds.

That's expected with any image generating models because they aren't trained with an alpha channel. It's more pragmatic to pipeline the results to a background removal model. EDIT: It appears GPT-4o is different as there is a video demo dedicated to transparancy.

There's an entire video in the post dedicated to how well it does transparency: https://openai.com/index/introducing-4o-image-generation/?vi...

I suspect we're getting a flood of comment from people who are using Dall-E.

Re: 4o Image Generation

#74
post #28

Still seems to have problems with transparent backgrounds.

That's expected with any image generating models because they aren't trained with an alpha channel. It's more pragmatic to pipeline the results to a background removal model. EDIT: It appears GPT-4o is different as there is a video demo dedicated to transparancy.

This one however explicitly advertises good transparency support.

Re: 4o Image Generation

#75

Earlier quoted context omitted.

That's expected with any image generating models because they aren't trained with an alpha channel. It's more pragmatic to pipeline the results to a background removal model. EDIT: It appears GPT-4o is different as there is a video demo dedicated to transparancy.

There's an entire video in the post dedicated to how well it does transparency: https://openai.com/index/introducing-4o-image-generation/?vi... I suspect we're getting a flood of comment from people who are using Dall-E.

Huh, I missed that. I'm skeptical of the results in practice, though.

Re: 4o Image Generation

#77
post #40

Edit: Please ignore. They hadn't rolled the new model out to my account yet. The announcement blog post is a bit misleading saying you can try it today. -- Comparison with Leonardo.Ai. ChatGPT: https://chatgpt.com/share/67e2fb21-a06c-8008-b297-07681dddee... ChatGPT again (direct one shot): https://chatgpt.com/share/67e2fc44-ecc8-8008-a40f-e1368d306e... ChatGPT again (using word "photorealistic instead of "photo"): ht…

[deleted]

Re: 4o Image Generation

#78
post #72
post #70

Earlier quoted context omitted.

What's the problem?

It's a nitpick about the repetitive phrasing for announcements : Our most yet|ever.

I hate modern marketing trends.

This one isn't even my biggest gripe. If I could eliminate any word from the English language forever, it would be "effortlessly".

Re: 4o Image Generation

#79
post #57

> Introducing 4o Image Generation: [...] our most advanced image generator yet Then google: > Gemini 2.5: Our most intelligent AI model > Introducing Gemini 2.0 | Our most capable AI model yet I could go on forever. I hope this trend dies and apple starts using something effective so all the other companies can start copying a new lexicon.

Has post-Jobs Apple ever come up with anything that would warrant this hope?

Re: 4o Image Generation

#80
post #49

OpenAI's livestream of GPT-4o Image Generation shows that it is slowwwwwwwwww (maybe 30 seconds per image, which Sam Altman had to spin "it's slow but the generated images are worth it"). Instead of using a diffusion approach, it appears to be generating the image tokens and decoding them akin to the original DALL-E ( https://openai.com/index/dall-e/ ), which allows for streaming partial generations from top to botto…

LLMs are autoregressive, so they can't be (multi-modality) integrated with diffusion image models, only with autoregressive image models (which generate an image via image tokens). Historically those had lower image fidelity than diffusion models. OpenAI now seems to have solved this problem somehow. More than that, they appear far ahead of any available diffusion model, including Midjourney and Imagen 3. Gemini "int…

I expect the Chinese to have an open source answer for this soon.

They haven't been focusing attention on images because the most used image models have been open source. Now they might have a target to beat.

Post reply on HN