Live data from Hacker News

4o Image Generation

openai.com

241–250 of 629 posts

Re: 4o Image Generation

#241
post #236

The real test for image generators is the image->text->image conversion. In other words it should be able to describe an image with words and then use the words to recreate the original image with a high accuracy. The text representation of the image doesn't have to be English. It can be a program, e.g. a shader, that draws the image. I believe in 5-10 years it will be possible to give this tool a picture of rainfore…

Why is that “the real test for image generators”? I mean, most image generators don't inherently include image->text functionality at all, so this seems more of a test of multimodal modals that include both t2i and i2t functionality, but even then, I don't think humans would generally pass this test well (unless the human doing the description test was explicitly told that the purpose was reproduction, but that's not the usual purpose of either human or image2text model descriptions.)

Re: 4o Image Generation

#242

OpenAI's livestream of GPT-4o Image Generation shows that it is slowwwwwwwwww (maybe 30 seconds per image, which Sam Altman had to spin "it's slow but the generated images are worth it"). Instead of using a diffusion approach, it appears to be generating the image tokens and decoding them akin to the original DALL-E ( https://openai.com/index/dall-e/ ), which allows for streaming partial generations from top to botto…

> it appears to be generating the image tokens and decoding them akin to the original DALL-E

The animation is a lie. The new 4o with "native" image generating capabilities is a multi-modal model that is connected to a diffusion model. It's not generating images one token at a time, it's calling out to a multi-stage diffusion model that has upscalers.

You can ask 4o about this yourself, it seems to have a strong understanding of how the process works.

Re: 4o Image Generation

#243
post #235
post #104

What's important about this new type of image generation that's happening with tokens rather than with diffusion, is that this is effectively reasoning in pixel space. Example: Ask it to draw a notepad with an empty tic-tac-toe, then tell it to make the first move, then you make a move, and so on. You can also do very impressive information-conserving translations, such as changing the drawing style, but also stuff l…

> What's important about this new type of image generation that's happening with tokens rather than with diffusion, is that this is effectively reasoning in pixel space. I do not think that this is correct. Prior to this release, 4o would generate images by calling out to a fully external model (DALL-E). After this release, 4o generates images by calling out to a multi-modal model that was trained alongside it. You c…

Models are famously good at understanding themselves.

Re: 4o Image Generation

#244
post #95

Earlier quoted context omitted.

Is that an artifact of the training data? Where are all these original images with that cartoony look that it was trained on?

A large part of deviantart.com would fit that description. There are also a lot of cartoony or CG images in communities dedicated to fanart. Another component in there is probably the overly polished and clean look of stock images, like the front page results of shutterstock. "Typical" AI images are this blend of the popular image styles of the internet. You always have a bit of digital drawing + cartoon image + over…

> There are also a lot of cartoony or CG images in communities dedicated to fanart.

Asian artists don't color this way though; those neon oversaturated colors are a Western style.

(This is one of the easiest ways to tell a fake-anime western TV show, the colors are bad. The other way is that action scenes don't have any impact because they aren't any good at planning them.)

Re: 4o Image Generation

#245
I wish AI companies would release new things once a year, like at CES or how Apple does it. This constant stream of releases and announcements feels like it's just for attention.

Re: 4o Image Generation

#247
For the first time ever, it feels like it listens and actually tries to follow what I say. I managed to actually get a good photo of a dog in the beach with shoes, from a side angle, by consistently prompting it and making small changes from one image to another till I got my intended effect

Re: 4o Image Generation

#248

It bothers me to see links to content that requires a login. I don't expect openai or anyone else to give their services away for free. But I feel like "news" posts that require one to setup an account with a vendor are bad faith. If the subject matter is paywalled, I feel that the post should include some explanation of what is newsworthy behind the link.

The linked post is not paywalled, did you click something else?

Re: 4o Image Generation

#249
post #7
post #3

> ChatGPT’s new image generation in GPT‑4o rolls out starting today to Plus, Pro, Team, and Free users as the default image generator in ChatGPT, with access coming soon to Enterprise and Edu. For those who hold a special place in their hearts for DALL·E, it can still be accessed through a dedicated DALL·E GPT. > Developers will soon be able to generate images with GPT‑4o via the API, with access rolling out in the n…

Yep. The coherence and text quality is insanely good. Keen to play with it to find it's "mangled hands" style deficiencies, because of course they cherry picked the best examples.

[deleted]
Post reply on HN