Live data from Hacker News

4o Image Generation

openai.com

121–130 of 629 posts

Re: 4o Image Generation

#122

This works great for many purposes. One area where it does not work well at all is modifying photographs of people's faces.* Completely fumbles if you take a selfie and ask it to modify your shirt, for example. * = unless the people are in the training set

> We’re aware of a bug where the model struggles with maintaining consistency of edits to faces from user uploads but expect this to be fixed within the week. Sounds like it may be a safety thing that's still getting figured out

Thanks, I had not seen that caveat!

Re: 4o Image Generation

#123
post #95

I’ll just be happy with not everything having that over saturated cg/cartoon style that you cant prompt your way out of.

Is that an artifact of the training data? Where are all these original images with that cartoony look that it was trained on?

Ever since Midjourney popularized it, image generation models are often posttrained on more "aesthetic" subsets of images to give them a more fantasy look. It also help obscure some of the imperfections of the AI.

Re: 4o Image Generation

#124
post #105
post #104

What's important about this new type of image generation that's happening with tokens rather than with diffusion, is that this is effectively reasoning in pixel space. Example: Ask it to draw a notepad with an empty tic-tac-toe, then tell it to make the first move, then you make a move, and so on. You can also do very impressive information-conserving translations, such as changing the drawing style, but also stuff l…

Pretty sure the modern Gemini image models can already do token based image generation/editing and are significantly better and faster.

Yeah Gemini has had this for a few weeks, but much lower resolution. Not saying 4o is perfect, but my first few images with it are much more impressive than my first few images with Gemini.

Re: 4o Image Generation

#125

I’ll just be happy with not everything having that over saturated cg/cartoon style that you cant prompt your way out of.

Frustratingly the DALL-E API actually has an option for this, you can switch it from "vivid" to "realistic".

This option is not exposed in ChatGPT, it only uses vivid.

Re: 4o Image Generation

#126

Earlier quoted context omitted.

[flagged]

[flagged]

I work on a product for generating interactive fanfiction using an LLM, and I've put a lot of work into post-training to improve writing quality to match or exceed typical human levels.

I'm excited about this for adding images to those interactive stories.

It has nothing to do with circumventing the cost of artists or writers: regardless of cost, no one can put out a story and then rewrite it based on whatever idea pops into every reader's mind for their own personal main character.

It's a novel experience that only a "writer" that scales by paying for an inanimate object to crunch numbers can enable.

Similarly no artist can put out a piece of art for that story and then go and put out new art bespoke to every reader's newly written story.

-

I think there's this weird obsession with framing these tools about being built to just replace current people doing similar things. Just speaking objectively: the market for replacing "cheeky expensive artists" would not justify building these tools.

The most interesting applications of this technology being able to do things that are simply not possible today even if you have all the money in the world.

And for the record, I'll be ecstatic for the day an AI can reach my level of competency in building software. I've been doing it since I was a child because I love it, it's the one skill I've ever been paid for, and I'd still be over the moon because it'd let me explore so many more ideas than I alone can ever hope to build.

Re: 4o Image Generation

#127
post #98

To quote myself from a comment on sora: Iterations are the missing link. With ChatGPT, you can iteratively improve text (e.g., "make it shorter," "mention xyz"). However, for pictures (and video), this functionality is not yet available. If you could prompt iteratively (e.g., "generate a red car in the sunset," "make it a muscle car," "place it on a hill," "show it from the side so the sun shines through the windshie…

You can do that with Gemini's image model, flash 2.0 (image generation) exp.[1] It's not perfect but it does mostly maintain likeness between generations.

[1]https://aistudio.google.com/prompts/new_chat

Re: 4o Image Generation

#128
I would love to see advancement in the pixel art space, specifying 64x64 pixels and attempting to make game-ready pixel art and even animations, or even taking a reference image and creating a 64x64 version

Re: 4o Image Generation

#129
post #96

Earlier quoted context omitted.

Has post-Jobs Apple ever come up with anything that would warrant this hope?

No, but I think they stopped with "our most" (since all other brainless corps adopted it) and just connect adjectives with dots. Hotwheels: Fast. Furious. Spectacular.

Maybe people also caught up to the fact that the "our most X product" for Apple usually means someone else already did X a long time ago and Apple is merely jumping on the wagon.

Re: 4o Image Generation

#130
post #104

What's important about this new type of image generation that's happening with tokens rather than with diffusion, is that this is effectively reasoning in pixel space. Example: Ask it to draw a notepad with an empty tic-tac-toe, then tell it to make the first move, then you make a move, and so on. You can also do very impressive information-conserving translations, such as changing the drawing style, but also stuff l…

It still can't generate a full glass of wine. Even in follow up questions it failed to manipulate the image correctly.
Post reply on HN