Earlier quoted context omitted.
Yeah, this is in my opinion the biggest limitation of the current gen GPT 4o image generation: it is incapable of editing only parts of an image. I assume what it does every time is tokenizing the source image, then transforming it according to the prompt and then giving you the final result. For some use cases that’s fine but if you really just want a small edit while keeping the rest of the image intact you’re out…
It just means that you comp it together manually. That's still much better than having to set up some inpainting pipeline or whatever.
I'm hoping we see an open weights or open source model with these capabilities soon, because good tools need open models.
As has happened in the past, once an open implementation of DallE or whatever comes out, the open source community pushes the capabilities much further by writing lots of training, extensions, and pipelines. The results look significantly better than closed SaaS models.