Live data from Hacker News

4o Image Generation

openai.com

1–10 of 629 posts

Re: 4o Image Generation

#2
OpenAI's livestream of GPT-4o Image Generation shows that it is slowwwwwwwwww (maybe 30 seconds per image, which Sam Altman had to spin "it's slow but the generated images are worth it"). Instead of using a diffusion approach, it appears to be generating the image tokens and decoding them akin to the original DALL-E (https://openai.com/index/dall-e/), which allows for streaming partial generations from top to bottom. In contrast, Google's Gemini can generate images and make edits in seconds.

No API yet, and given the slowness I imagine it will cost much more than the $0.03+/image of competitors.

Re: 4o Image Generation

#3
> ChatGPT’s new image generation in GPT‑4o rolls out starting today to Plus, Pro, Team, and Free users as the default image generator in ChatGPT, with access coming soon to Enterprise and Edu. For those who hold a special place in their hearts for DALL·E, it can still be accessed through a dedicated DALL·E GPT.

> Developers will soon be able to generate images with GPT‑4o via the API, with access rolling out in the next few weeks.

That's it folks. Tens of thousands of so-called "AI" image generator startups have been obliterated and taking digital artists with them all reduced to near zero.

Now you have a widely accessible meme generator with the name "ChatGPT".

The last task is for an open weight model that competes against this and is faster and all for free.

Re: 4o Image Generation

#5

Did they time it with the Gemini 2.5 launch? https://news.ycombinator.com/item?id=43473489 Was it public information when Google was going to launch their new models? Interesting timing.

"Interesting timing" It's like the 4th time by my counting they've done this

Re: 4o Image Generation

#7
post #3

> ChatGPT’s new image generation in GPT‑4o rolls out starting today to Plus, Pro, Team, and Free users as the default image generator in ChatGPT, with access coming soon to Enterprise and Edu. For those who hold a special place in their hearts for DALL·E, it can still be accessed through a dedicated DALL·E GPT. > Developers will soon be able to generate images with GPT‑4o via the API, with access rolling out in the n…

Yep. The coherence and text quality is insanely good. Keen to play with it to find it's "mangled hands" style deficiencies, because of course they cherry picked the best examples.

Re: 4o Image Generation

#10
post #6

Looks about what you'd get with FLUX and attaching some language model to enhance your prompt with eg more text

Exactly. OpenAI isn't going to win image and video.

Sora is one of the worst video generators. The Chinese have really taken the lead in video with Kling, Hailuo, and the open source Wan and Hunyuan.

Wan with LoRAs will enable real creative work. Motion control, character consistency. There's no place for an OpenAI Sora type product other than as a cheap LLM add-in.

Post reply on HN