Live data from Hacker News

4o Image Generation

openai.com

161–170 of 629 posts

Re: 4o Image Generation

#161

am I dumb or every time they release something I can never find out how to actually use it and forget about it. take this for instance I wanted to try out their newton "an infographic explaining newton's prism experiment in great detail" example, but it generated a very bad result but maybe it's because I'm not using the right model? every release of theirs is not really a release, it's like a trailer. right?

This is what's so wild about Anthropic. When they release it seems like it's rolled out to all users, and API customers immediately. OpenAI has MONTHS between annoucement and roll out, or if they do it's usually just influencers who get an "early look". It's pretty frustrating.

Re: 4o Image Generation

#162
post #104

What's important about this new type of image generation that's happening with tokens rather than with diffusion, is that this is effectively reasoning in pixel space. Example: Ask it to draw a notepad with an empty tic-tac-toe, then tell it to make the first move, then you make a move, and so on. You can also do very impressive information-conserving translations, such as changing the drawing style, but also stuff l…

> What's important about this new type of image generation that's happening with tokens rather than with diffusion

That sounds really interesting. Are there any write-ups how exactly this works?

Re: 4o Image Generation

#164
post #153
post #80

Earlier quoted context omitted.

I expect the Chinese to have an open source answer for this soon. They haven't been focusing attention on images because the most used image models have been open source. Now they might have a target to beat.

ByteDance has been working on autoregressive image generation for a while (see VAR, NeurIPS 2024 best paper). Traditionally they weren't in the open-source gang though.

The VAR paper is very impressive. I wonder if OpenAI did something similar. But the main contribution in the new GPT-4o feature doesn't seem to be just image quality (which VAR seems to focus on), but also massively enhanced prompt understanding.

Re: 4o Image Generation

#166
post #63

Earlier quoted context omitted.

Is this the same for their gemini-2.0-flash-exp-image-generation model?

No that seems to be indeed a native part of the multimodal Gemini model. I didn't know this existed, it's not available in the normal Gemini interface.

This is a pretty good example of the current state of Google LLMs:

The (no longer, I guess) industry-leading features people actually want are hidden away in some obscure “AI studio” with horrible usability, while the headline Gemini app still often refuses to do anything useful for me. (Disclaimer: I last checked a couple of months ago, after several more of mild amusement/great frustration.)

Re: 4o Image Generation

#167
post #95

I’ll just be happy with not everything having that over saturated cg/cartoon style that you cant prompt your way out of.

Is that an artifact of the training data? Where are all these original images with that cartoony look that it was trained on?

I've heard this is downstream of human feedback. If you ask someone which picture is better, they'll tend to pick the more saturated option. If you're doing post-training with humans, you'll bake that bias into your model.

Re: 4o Image Generation

#168
post #57

> Introducing 4o Image Generation: [...] our most advanced image generator yet Then google: > Gemini 2.5: Our most intelligent AI model > Introducing Gemini 2.0 | Our most capable AI model yet I could go on forever. I hope this trend dies and apple starts using something effective so all the other companies can start copying a new lexicon.

https://www.youtube.com/watch?v=CUPDRnUWeBA

Re: 4o Image Generation

#169

Earlier quoted context omitted.

Are you sure you are using the new 4o image generation? https://imgur.com/a/wGkBa0v

Looks amazing,can you please also create a unconventional image like the clock at 2:35 , I tried it something like this with gemini when some redditor asked it and it failed so wondering if 4o does do it

I tried and while the clock it generated was very well done and high quality, it showed the time as the analog clock default of 10:10.

Re: 4o Image Generation

#170
I’ve just tried it and oh wow it’s really good. I managed to create a birthday invitation card for my daughter in basically 1-shot, it nailed exactly the elements and style I wanted. Then I asked to retain everything but tweak the text to add more details about the date, venue etc. And it did. I’m in shock. Previous models would not be even halfway there.
Post reply on HN