4o Image Generation
511–520 of 629 posts
Re: 4o Image Generation
#512Anyone else frightened by this? Seeing meant believing, and now that isnt the case anymore...
Re: 4o Image Generation
#513Can someone explain what is going with 4o and Anime with Ghibli Style? Why is it suddenly all over x/twitter?
Re: 4o Image Generation
#514Re: 4o Image Generation
#515Earlier quoted context omitted.
Yeah I wasn’t very imaginative in my examples, with 4o you can also perform transformations like “rotate the camera 10 degrees to the left” which would be hard without a specialized model. Basically you can run arbitrary functions on the exact image contents but in latent space.
It's doable with diffusion, too.
4o is a game changer. It's clearly imperfect, but its operating modalities are clearly superior to everything else we have seen.
Have you seen (or better yet, played with) the whiteboard examples? Or the examples of it taking characters out of reflections and manipulating them? The prompt adherence, text layout, and composing capabilities are unreal to the point this looks like it completely obsoletes inpainting and outpainting.
I'm beginning to think this even obsoletes ComfyUI and the whole space of open source tools once the model improves. Natural language might be able to accomplish everything outside of fine adjustments, but if you can also supply the model with reference images and have it understand them, then it can do basically everything. I haven't bumped into anything that makes me question this yet.
They just need to bump the speed and the quality a little. They're back at the top of image gen again.
I'm hoping the Chinese or another US company releases an open model capable of these behaviors. Because otherwise OpenAI is going to take this ball and run far ahead with it.
Re: 4o Image Generation
#516Earlier quoted context omitted.
>Of course the model may hallucinate, but in this case it takes a few clicks in the dev tools to verify that this is not the case. I don't know - or care to figure out - how OpenAI does their tool calling in this specific case. But moving tool calls to the end user is _monumentally_ stupid for the latency if nothing else. If you centralize your function calls to a single model next to a fat pipe it means that you hal…
It's not client side, the messages are in the api though. But what do you mean you don't care? The thing you were responding to was literally a claim that it was a tool call rather than direct output
The thing we need to worry about is whether a Chinese company will drop an open source equivalent.