Live data from Hacker News

4o Image Generation

openai.com

351–360 of 629 posts

Re: 4o Image Generation

#351
post #271

Earlier quoted context omitted.

Models are famously good at understanding themselves.

I hope you're joking. Sometimes they don't even know which company developed them. E.g. DeepSeek was claiming it was developed by OpenAI.

I have asked GPT if it is using the 4o or 4.5 model multiple times in voice mode e.g. "Which model are you using?". It has said that it is using 4.5 when it is actually using 4o.

Re: 4o Image Generation

#352
post #185

Earlier quoted context omitted.

https://i.imgur.com/xsFKqsI.png "Draw a picture of a full glass of wine, ie a wine glass which is full to the brim with red wine and almost at the point of spilling over... Zoom out to show the full wine glass, and add a caption to the top which says "HELL YEAH". Keep the wine level of the glass exactly the same."

The head of foam on that glass of wine is perfect !

I think we're really fscked, because even AI image detectors think the images are genuine. They look great in Photoshop forensics too. I hope the arms race between generators and detectors doesn't stop here.

Re: 4o Image Generation

#353
post #235

Earlier quoted context omitted.

> What's important about this new type of image generation that's happening with tokens rather than with diffusion, is that this is effectively reasoning in pixel space. I do not think that this is correct. Prior to this release, 4o would generate images by calling out to a fully external model (DALL-E). After this release, 4o generates images by calling out to a multi-modal model that was trained alongside it. You c…

You're incorrect. 4o was not trained on knowledge of itself so literally can't tell you that. What 4o is doing isn't even new either, Gemini 2.0 has the same capability.

The system prompt includes instructions on how to use tools like image generation. From that it could infer what the GP posted.

Re: 4o Image Generation

#354

Earlier quoted context omitted.

Humans don’t train on the entire contents of the Internet, so i’d wager that they do learn differently

I think there is a critical aspect of human visual learning which machine leanring cant replicate because it is prohibitively expensive. When we look at things as children we are not just looking at a single snapshot. When you stare at an object for a few seconds you have practically injested hundreds of slightly variated images of that object. This gets even more interesting when you take into account real world is…

Then explain blind children? Or blind & deaf children? There's obviously some role senses play in development but there's clearly capabilities at play here that are drastically more efficient and powerful than what we have with modern transformers. While humans learn through example, they clearly need a lot fewer examples to generalize off of and reason against.

Re: 4o Image Generation

#355

Earlier quoted context omitted.

It still can't generate a full glass of wine. Even in follow up questions it failed to manipulate the image correctly.

I don't buy the meme or w/e that they can't produce an image with the full glass of wine. Just takes a little prompt engineering. Using Dall-e / old model without too much effort (I'd call this "full".) https://imgur.com/a/J2bCwYh

The true test was "full to the brim", as in almost overflowing.

Re: 4o Image Generation

#357
post #269

> we see the photographer's reflection Am I the only one immediately looking past the amazing text generation, the excellent direction following, the wonderful reflection, and screaming inside my head, "That's not how reflection works!" I know it's super nitpicky when it's so obviously a leap forward on multiple other metrics, but still, that reflection just ain't right.

Could you explain more? I'm having trouble seeing anything weird in the reflection. Edit: are we talking about the first or second image? I meant to say the image with only the woman seems normal. Image with the two people does seem a bit odd.

The first image, with the photographer holding the phone reflected in the white board.

Angle of incidence = angle of reflection. That means that the only way to see yourself in a reflective surface is by looking directly at it. Note this refers to looking at your eyes -- you can look down at a mirror to see your feet because your feet aren't where your eyes are.

You can google "mirror selfie" to see endless examples of this. Now look for one where the camera isn't pointing directly at the mirror.

From the way the white board is angled, it's clear the phone isn't facing it directly. And yet the reflection of the phone/photographer is near-center in frame. If you face a mirror and angle to the left the way the image is, your reflection won't be centered, it'll be off to the right, where your eyes can see it because you have a very wide field of view, but a phone would not.

Re: 4o Image Generation

#358
post #295
post #288

Visual internet content is completely over. Pack it up

For starters, this completely blocks generation of anything remotely related to copy-protected IPs, which may actually be a saving grace for some creatives. There's a lot of demand for fanart of existing characters, so until this type of model can be run locally, the legal blocks in place actually give artists some space to play in where they don't have to compete with this. At least for a short while.

>For starters, this completely blocks generation of anything remotely related to copy-protected IPs

It did Dragon Ball Z here:

https://old.reddit.com/r/ChatGPT/comments/1jjtcn9/the_new_im...

Rick and Morty:

https://old.reddit.com/r/ChatGPT/comments/1jjtcn9/the_new_im...

South Park:

https://old.reddit.com/r/ChatGPT/comments/1jjyn5q/openais_ne...

Re: 4o Image Generation

#359
post #133

Earlier quoted context omitted.

4o generates top down (picture goes from mostly blurry to clear starting from the top). If it's not generating like that for you then you don't have it yet.

That's useful, thank you! But it also highlights my point: Why do I have to observe minor details about how the result is being presented to me to know which model was used? I get the intent to abstract it all behind a chat interface, but this seems a bit too much.

I mean in the webpage the dalle one has a bubble under that says "generated with dall-e"

Re: 4o Image Generation

#360
post #295
post #288

Visual internet content is completely over. Pack it up

For starters, this completely blocks generation of anything remotely related to copy-protected IPs, which may actually be a saving grace for some creatives. There's a lot of demand for fanart of existing characters, so until this type of model can be run locally, the legal blocks in place actually give artists some space to play in where they don't have to compete with this. At least for a short while.

Fan-art is still illegal, especially since a lot of fan artists are doing it commercially nowadays via commissions and Patreon. It's just that companies have stopped bothering to sue for it because individual artists are too small to bother with, and it's bad PR. (Nintendo did take down a super popular Pokemon porn comic, though.)

So it's ironic in this sense, that OpenAI blocking generation of copyrighted characters means that it's more in compliance with copyright laws than most fan artists out there, in this context. If you consider AI training to be transformative enough to be permissible, then they are more copyright-respecting in general.

Source: https://lawsoup.org/legal-guides/copyright-protecting-creati...

Post reply on HN