Not sure why they put so much investment into videoSlop and imageSlop. Anthropic seems to be more focused at least.
Because OpenAI stands for AI leader. If Gemini can create or edit an image, chatgpt needs to be able to do this too. Who wants to copy&paste prompts between ai agents? Also if you want to have more semantics, you add image, video and audio to your model. It gets smarter because of it. OpenAI is also relevant bigger than antropic and is known as a generic 'helper'. Antropic probably saw the benefits of being more focu…
I think you are confusing generation with analysis. As far I am aware your model does not need to be good at generating images to be able to decode an image.