Live data from Hacker News

4o Image Generation

openai.com

451–460 of 629 posts

Re: 4o Image Generation

#451
post #275

Earlier quoted context omitted.

Is it drawing the image from top to bottom very slowly over the course of at least 30 seconds? If not, then you're using DALL-E, not 4o image generation.

This top to bottom drawing – does this tell us anything about the underlying model architecture? AFAIK diffusion models do not work like that. They denoise the full frame over many steps. In the past there used to be attempts to slowly synthetize a picture by predicting the next pixel, but I wasn't aware whether there has been a shift to that kind of architecture within OpenAI.

Yes, the model card explicitly says it's autoregressive, not diffusion. And it's not a separate model, it's a native ability of GPT-4o, which is a multimodal model. They just didn't made this ability public until now. I assume they worked on the fine-tuning to improve prompt following.

Re: 4o Image Generation

#452

Interesting that in the second image the text on the whiteboard changes (top left)

It seems this is because the string "autoregressive prior" should appear on the right hand side as well, but in the second image it's hidden from view, and this has confused it to place it on the left hand side instead?

It also misses the arrow between "[diffusion]" and "pixels" in the first image.

Re: 4o Image Generation

#453

Earlier quoted context omitted.

This is overly cynical. Models typically do know what tools they have access to because the tool descriptions are in the prompt. Asking a model which tools it has is a perfectly reasonable way of learning what is effectively the content of the prompt. Of course the model may hallucinate, but in this case it takes a few clicks in the dev tools to verify that this is not the case.

>Of course the model may hallucinate, but in this case it takes a few clicks in the dev tools to verify that this is not the case. I don't know - or care to figure out - how OpenAI does their tool calling in this specific case. But moving tool calls to the end user is _monumentally_ stupid for the latency if nothing else. If you centralize your function calls to a single model next to a fat pipe it means that you hal…

You should check out Claude desktop or Roo-Code or any of the other MCP client capable hosts. The whole idea of MCP is providing a universal pluggable tool api to the generative model.

Re: 4o Image Generation

#455

First AI image generator to pass the uncanny valley test? Seems like it. This is the biggest leap in image generation quality I've ever seen. How much longer until an AI that can generate 30 frames with this quality and make a movie? About 1.5 years ago, I thought AI would eventually allow anyone with an idea to make a Hollywood quality movie. Seems like we're not too far off. Maybe 2-3 more years?

>First AI image generator to pass the uncanny valley test?

Other image generators I've used lately often produced pretty good images of humans, as well [0]. It was DALLE that consistently generated incredibly awful images. Glad they're finally fixing it. I think what most AI image generators lack the most is good instruction following.

[0] YandexArt for the first prompt from the post: https://imgur.com/a/VvNbL7d The woman looks okay, but the text is garbled, and it didn't fully follow the instruction.

Re: 4o Image Generation

#456
post #407
post #393

My experience with these announcements is that they're cherry picking the best results from a maybe several hundred or a thousand prompts. I'm not saying that it's not true, it's just "wait and see" before you take their word as gold. I think MS's claim on their quantum computing breakthrough is the latest form of this.

The examples they show have little captions that say "best of #", like "best of 8" or "best of 4". Hopefully that truly represents the odds of generating the level of quality shown.

Some of the prompts are pretty long. I'm curious how iterations it took to get to that prompt for them to take the top 8 out of.

Re: 4o Image Generation

#457
post #455

First AI image generator to pass the uncanny valley test? Seems like it. This is the biggest leap in image generation quality I've ever seen. How much longer until an AI that can generate 30 frames with this quality and make a movie? About 1.5 years ago, I thought AI would eventually allow anyone with an idea to make a Hollywood quality movie. Seems like we're not too far off. Maybe 2-3 more years?

>First AI image generator to pass the uncanny valley test? Other image generators I've used lately often produced pretty good images of humans, as well [0]. It was DALLE that consistently generated incredibly awful images. Glad they're finally fixing it. I think what most AI image generators lack the most is good instruction following. [0] YandexArt for the first prompt from the post: https://imgur.com/a/VvNbL7d The…

Do you have another example from YandexArt?

https://images.ctfassets.net/kftzwdyauwt9/7M8kf5SPYHBW2X9N46...

OpenAI's human faces look *almost* real.

Re: 4o Image Generation

#458
post #194

Earlier quoted context omitted.

Can't replicate. Maybe the rollout is staggered? Using Plus from Europe, it's consistently giving me a half full glass.

I am using Plus from Australia, and while I am not getting a full glass, nor am I getting a half full glass. The glass I'm getting is half empty.

Surprised it isn't fully empty for being upside down!

Re: 4o Image Generation

#460
One very neat thing the interwebs are talking about is the ghiblification of family pictures. It’s actually pretty cute: https://x.com/grantslatton/status/1904631016356274286

In the coming days, people will Anime all sorts of images, for example historical images: https://x.com/keysmashbandit/status/1904764224636592188

Post reply on HN