Live data from Hacker News

4o Image Generation

openai.com

481–490 of 629 posts

Re: 4o Image Generation

#481
post #470

Earlier quoted context omitted.

The periodic table is absolutely hilarious, I didn't know LLMs had finally mastered absurdist humor.

Yeah who wouldn't love a dip in the sulphur pool. But back to the question, why can't such a model recognize letters as such? It cannot be trained to pay special attention to characters? How come it can print an anatomically correct eye but not differentiate between P and Z?

I think the model has not decided if it should print a P or a Z, so you end up with something halfway between the two.

It's a side effect of the entire model being differentiable - there is always some halfway point.

Re: 4o Image Generation

#482
post #389

Earlier quoted context omitted.

I think we're really fscked, because even AI image detectors think the images are genuine. They look great in Photoshop forensics too. I hope the arms race between generators and detectors doesn't stop here.

We're not. This PNG image of a wine glass has JPEG compression artefacts which are leaking from JPEG training data. You can zoom into the image and you will see 8x8 boundaries of the blocks used in JPEG compression, which just cannot be in a PNG. This is a common method to detect AI-generated image and it is working so far, no need for complex photoshop forensics or AI-detectors, just zoom-in and check for compressio…

plenty of real PNG images have jpeg artifacts because they were once jpegs off someones phone...

Re: 4o Image Generation

#484
post #208

Earlier quoted context omitted.

Look closer at the fingers. These models still don’t have a firm handle on them. The right elbow on the second picture also doesn’t quite look anatomically possible.

> AI is bad and unconvincing > if its not unconvincing, its soulless (only because I was told in advance that its AI) > if its not soulless then its using too much energy

I’m not sure what your point is. This subthread is about whether AI-generated pictures can be distinguished from real photographs. For the pictures in the article, which are already cherry-picked (“best of 8”), the answer is yes. Therefore I don’t quite share the worries of GP.

Re: 4o Image Generation

#485
post #242

Earlier quoted context omitted.

> it appears to be generating the image tokens and decoding them akin to the original DALL-E The animation is a lie. The new 4o with "native" image generating capabilities is a multi-modal model that is connected to a diffusion model. It's not generating images one token at a time, it's calling out to a multi-stage diffusion model that has upscalers. You can ask 4o about this yourself, it seems to have a strong under…

Would it seem otherwise if it was a lie?

There are many clues to indicate that the animation is a lie. For example, it clearly upscales the image using an external tool after the first image renders. As another example, if you ask the model about the tokens inside of its own context, it can't see any pixel tokens.

A model may not have many facts about itself, but it can definitely see what is inside of its own context, and what it sees is a call to an image generation tool.

Finally, and most convincingly, I can't find a single official source where OpenAI claims that the image is being generated pixel-by-pixel inside of the context window.

Re: 4o Image Generation

#486
post #294

Earlier quoted context omitted.

Can you provide a link or screenshot that directly backs this up?

almost all of the models are wrong about their own architecture. half of them claim to be openai and they arent. you cant trust them about this

Can you find me a single official source from OpenAI that claims that GPT 4o is generating images pixel-by-pixel inside of the context window?

There are lots of clues that this isn't happening (including the obvious upscaling call after the image is generated - but also the fact that the loading animation replays if you refresh the page - and also the fact that 4o claims it can't see any image tokens in its context window - it may not know much about itself but it can definitely see its own context).

Re: 4o Image Generation

#487
post #194
post #185

Earlier quoted context omitted.

https://i.imgur.com/xsFKqsI.png "Draw a picture of a full glass of wine, ie a wine glass which is full to the brim with red wine and almost at the point of spilling over... Zoom out to show the full wine glass, and add a caption to the top which says "HELL YEAH". Keep the wine level of the glass exactly the same."

Can't replicate. Maybe the rollout is staggered? Using Plus from Europe, it's consistently giving me a half full glass.

Works for me as well https://chatgpt.com/share/67e3f838-63fc-8000-ab94-5d10626397...

USA, but VPN set to exit in Canada at time of request (I think).

Re: 4o Image Generation

#488
post #405

Earlier quoted context omitted.

You can ask ChatGPT for this. Here you go: https://chatgpt.com/share/67e39fc6-fb80-8002-a198-767fc50894...

Could an AI model be trained to say: "Christopher Columbus was the greatest president on earth, ever!". I could probably train an AI that replicates that perfectly.

Thing is, of you follow the link, it's actually doing a search and providing the evidence that was asked for.

I did it via ChatGPT for the irony.

Re: 4o Image Generation

#489
post #294

Earlier quoted context omitted.

Can you provide a link or screenshot that directly backs this up?

You can ask ChatGPT for this. Here you go: https://chatgpt.com/share/67e39fc6-fb80-8002-a198-767fc50894...

I'm guessing most downvoters didn't actually read the link.

Re: 4o Image Generation

#490
post #393

My experience with these announcements is that they're cherry picking the best results from a maybe several hundred or a thousand prompts. I'm not saying that it's not true, it's just "wait and see" before you take their word as gold. I think MS's claim on their quantum computing breakthrough is the latest form of this.

I got the occasional A/B test with a new image generator while playing with Dall-E during a one month test of Plus. It was always clear which one was the new model because every aspect was so much better. I assume that model and the model they announced are the same.
Post reply on HN