Live data from Hacker News

4o Image Generation

openai.com

221–230 of 629 posts

Re: 4o Image Generation

#221

Earlier quoted context omitted.

It still can't generate a full glass of wine. Even in follow up questions it failed to manipulate the image correctly.

The question remains: why would you generate a full glass of wine? Is that something really that common?

It’s a type of QA question that can identify peculiarities in models (e.g. count “r”s in strawberry), which the best we have given the black box nature of LLMs.

Re: 4o Image Generation

#223
It bothers me to see links to content that requires a login. I don't expect openai or anyone else to give their services away for free. But I feel like "news" posts that require one to setup an account with a vendor are bad faith.

If the subject matter is paywalled, I feel that the post should include some explanation of what is newsworthy behind the link.

Re: 4o Image Generation

#225
post #205

Anyone else frightened by this? Seeing meant believing, and now that isnt the case anymore...

Nah, I'll maybe start taking them seriously when they can draw someone grating cheese, but holding the cheese and the grater as if they were playing violin.

Re: 4o Image Generation

#228
post #170

I’ve just tried it and oh wow it’s really good. I managed to create a birthday invitation card for my daughter in basically 1-shot, it nailed exactly the elements and style I wanted. Then I asked to retain everything but tweak the text to add more details about the date, venue etc. And it did. I’m in shock. Previous models would not be even halfway there.

share prompt minus identifying details?

Re: 4o Image Generation

#229
post #162

Earlier quoted context omitted.

> What's important about this new type of image generation that's happening with tokens rather than with diffusion That sounds really interesting. Are there any write-ups how exactly this works?

Would be interested to know as well. As far as I know there is no public information about how this works exactly. This is all I could find: > The system uses an autoregressive approach — generating images sequentially from left to right and top to bottom, similar to how text is written — rather than the diffusion model technique used by most image generators (like DALL-E) that create the entire image at once. Goh sp…

I wonder how it'd work if the layers were more physical based. In other words something like rough 3d shape -> details -> color -> perspective -> lighting.

Also wonder if you'd get better results in generating something like blender files and using its engine to render the result.

Re: 4o Image Generation

#230
post #104

What's important about this new type of image generation that's happening with tokens rather than with diffusion, is that this is effectively reasoning in pixel space. Example: Ask it to draw a notepad with an empty tic-tac-toe, then tell it to make the first move, then you make a move, and so on. You can also do very impressive information-conserving translations, such as changing the drawing style, but also stuff l…

Hmmm, I wanted to do that tic tac toe example, and it failed to create a 3x3 grid, instead creating a 5x5 (?) grid with two first moves marked.

https://chatgpt.com/share/67e32d47-eac0-8011-9118-51b81756ec...

Post reply on HN