Earlier quoted context omitted.
Models are famously good at understanding themselves.
I hope you're joking. Sometimes they don't even know which company developed them. E.g. DeepSeek was claiming it was developed by OpenAI.
4o Image Generation
351–360 of 629 posts
Re: 4o Image Generation
#352Earlier quoted context omitted.
https://i.imgur.com/xsFKqsI.png "Draw a picture of a full glass of wine, ie a wine glass which is full to the brim with red wine and almost at the point of spilling over... Zoom out to show the full wine glass, and add a caption to the top which says "HELL YEAH". Keep the wine level of the glass exactly the same."
The head of foam on that glass of wine is perfect !
Re: 4o Image Generation
#353Earlier quoted context omitted.
> What's important about this new type of image generation that's happening with tokens rather than with diffusion, is that this is effectively reasoning in pixel space. I do not think that this is correct. Prior to this release, 4o would generate images by calling out to a fully external model (DALL-E). After this release, 4o generates images by calling out to a multi-modal model that was trained alongside it. You c…
You're incorrect. 4o was not trained on knowledge of itself so literally can't tell you that. What 4o is doing isn't even new either, Gemini 2.0 has the same capability.
Re: 4o Image Generation
#354Earlier quoted context omitted.
Humans don’t train on the entire contents of the Internet, so i’d wager that they do learn differently
I think there is a critical aspect of human visual learning which machine leanring cant replicate because it is prohibitively expensive. When we look at things as children we are not just looking at a single snapshot. When you stare at an object for a few seconds you have practically injested hundreds of slightly variated images of that object. This gets even more interesting when you take into account real world is…
Re: 4o Image Generation
#355Earlier quoted context omitted.
It still can't generate a full glass of wine. Even in follow up questions it failed to manipulate the image correctly.
I don't buy the meme or w/e that they can't produce an image with the full glass of wine. Just takes a little prompt engineering. Using Dall-e / old model without too much effort (I'd call this "full".) https://imgur.com/a/J2bCwYh
Re: 4o Image Generation
#356Re: 4o Image Generation
#357> we see the photographer's reflection Am I the only one immediately looking past the amazing text generation, the excellent direction following, the wonderful reflection, and screaming inside my head, "That's not how reflection works!" I know it's super nitpicky when it's so obviously a leap forward on multiple other metrics, but still, that reflection just ain't right.
Could you explain more? I'm having trouble seeing anything weird in the reflection. Edit: are we talking about the first or second image? I meant to say the image with only the woman seems normal. Image with the two people does seem a bit odd.
Angle of incidence = angle of reflection. That means that the only way to see yourself in a reflective surface is by looking directly at it. Note this refers to looking at your eyes -- you can look down at a mirror to see your feet because your feet aren't where your eyes are.
You can google "mirror selfie" to see endless examples of this. Now look for one where the camera isn't pointing directly at the mirror.
From the way the white board is angled, it's clear the phone isn't facing it directly. And yet the reflection of the phone/photographer is near-center in frame. If you face a mirror and angle to the left the way the image is, your reflection won't be centered, it'll be off to the right, where your eyes can see it because you have a very wide field of view, but a phone would not.
Re: 4o Image Generation
#358Visual internet content is completely over. Pack it up
For starters, this completely blocks generation of anything remotely related to copy-protected IPs, which may actually be a saving grace for some creatives. There's a lot of demand for fanart of existing characters, so until this type of model can be run locally, the legal blocks in place actually give artists some space to play in where they don't have to compete with this. At least for a short while.
It did Dragon Ball Z here:
https://old.reddit.com/r/ChatGPT/comments/1jjtcn9/the_new_im...
Rick and Morty:
https://old.reddit.com/r/ChatGPT/comments/1jjtcn9/the_new_im...
South Park:
https://old.reddit.com/r/ChatGPT/comments/1jjyn5q/openais_ne...
Re: 4o Image Generation
#359Earlier quoted context omitted.
4o generates top down (picture goes from mostly blurry to clear starting from the top). If it's not generating like that for you then you don't have it yet.
That's useful, thank you! But it also highlights my point: Why do I have to observe minor details about how the result is being presented to me to know which model was used? I get the intent to abstract it all behind a chat interface, but this seems a bit too much.
Re: 4o Image Generation
#360Visual internet content is completely over. Pack it up
For starters, this completely blocks generation of anything remotely related to copy-protected IPs, which may actually be a saving grace for some creatives. There's a lot of demand for fanart of existing characters, so until this type of model can be run locally, the legal blocks in place actually give artists some space to play in where they don't have to compete with this. At least for a short while.
So it's ironic in this sense, that OpenAI blocking generation of copyrighted characters means that it's more in compliance with copyright laws than most fan artists out there, in this context. If you consider AI training to be transformative enough to be permissible, then they are more copyright-respecting in general.
Source: https://lawsoup.org/legal-guides/copyright-protecting-creati...