Live data from Hacker News

4o Image Generation

openai.com

21–30 of 629 posts

Re: 4o Image Generation

#22
post #13

Earlier quoted context omitted.

Flux 1.1 Pro has good prompt adherence, but some of these (admittingly cherry-picked) GPT-4o generated image demos are beyond what you would get with Flux without a lot of iteration, particularly the large paragraphs of text. I'm excited to see what a Flux 2 can do if it can actually use a modern text encoder.

Structural editing and control nets are much more powerful than text prompting alone. The image generators used by creatives will not be text-first. "Dragon with brown leathery scales with an elephant texture and 10% reflectivity positioned three degrees under the mountain, which is approximately 250 meters taller than the next peak, ..." is not how you design. Creative work is not 100% dice rolling in a crude and in…

Yeah, but then it no longer replaces human artists.

Controlnet has been the obvious future of image-generation for a while now.

Re: 4o Image Generation

#23

LPT: while the benchmarks don't show it, chatGPT4>4o. It amazes me people use 4o at all. But hey its the brand name and its free. ofc 4.5 is best, but its slow and I am afraid I'm going to hit limits.

OpenAI themselves discourages using GPT-4 outside of legacy applications, in favor of GPT-4o instead (they are shutting down the large output gpt-4-32k variants in a few months). GPT-4 is also an order of magnitude more expensive/slower.

Re: 4o Image Generation

#24
This is really impressive, but the "Best of 8" tag on a lot of them really makes me want to see how cherry-picked they are. My three free images had two impressive outputs and one failure.

Re: 4o Image Generation

#25
post #22
post #13

Earlier quoted context omitted.

Structural editing and control nets are much more powerful than text prompting alone. The image generators used by creatives will not be text-first. "Dragon with brown leathery scales with an elephant texture and 10% reflectivity positioned three degrees under the mountain, which is approximately 250 meters taller than the next peak, ..." is not how you design. Creative work is not 100% dice rolling in a crude and in…

Yeah, but then it no longer replaces human artists. Controlnet has been the obvious future of image-generation for a while now.

We're not trying to replace human artists. We're trying to make them more efficient.

We might find that the entire "studio system" is a gross inefficiency and that individual artists and directors can self-publish like on Steam or YouTube.

Re: 4o Image Generation

#26
post #3

> ChatGPT’s new image generation in GPT‑4o rolls out starting today to Plus, Pro, Team, and Free users as the default image generator in ChatGPT, with access coming soon to Enterprise and Edu. For those who hold a special place in their hearts for DALL·E, it can still be accessed through a dedicated DALL·E GPT. > Developers will soon be able to generate images with GPT‑4o via the API, with access rolling out in the n…

> Tens of thousands of so-called "AI" image generator startups have been obliterated and taking digital artists with them all reduced to near zero. Now you have a widely accessible meme generator with the name "ChatGPT".

ChatGPT has already had a that via Dall-E. If it didn't kill those startups when that happened this doesn't fundamentally change anything. Now its got a new image gen model, which — like Dall-E 3 when it came out — is competitive or ahead of other SotA base models using just text prompts, the simplest generation workflow, but both more expensive and less adaptable to more involved workflows than the tools anyone more than a casual user (whether using local tools or hosted services) is using. This is station-keeping for OpenAI, not a meaningful change in the landscape.

Re: 4o Image Generation

#27
post #13

Earlier quoted context omitted.

Flux 1.1 Pro has good prompt adherence, but some of these (admittingly cherry-picked) GPT-4o generated image demos are beyond what you would get with Flux without a lot of iteration, particularly the large paragraphs of text. I'm excited to see what a Flux 2 can do if it can actually use a modern text encoder.

Structural editing and control nets are much more powerful than text prompting alone. The image generators used by creatives will not be text-first. "Dragon with brown leathery scales with an elephant texture and 10% reflectivity positioned three degrees under the mountain, which is approximately 250 meters taller than the next peak, ..." is not how you design. Creative work is not 100% dice rolling in a crude and in…

Prompt adherence and additional tricks such as ControlNet/ComfyUI pipelines are not mutually exclusive. Both are very important to get good image generation results.

Re: 4o Image Generation

#29
am I dumb or every time they release something I can never find out how to actually use it and forget about it. take this for instance I wanted to try out their newton "an infographic explaining newton's prism experiment in great detail" example, but it generated a very bad result but maybe it's because I'm not using the right model? every release of theirs is not really a release, it's like a trailer. right?

Re: 4o Image Generation

#30
This works great for many purposes.

One area where it does not work well at all is modifying photographs of people's faces.* Completely fumbles if you take a selfie and ask it to modify your shirt, for example.

* = unless the people are in the training set

Post reply on HN