Live data from Hacker News

4o Image Generation

openai.com

461–470 of 629 posts

Re: 4o Image Generation

#461

https://chatgpt.com/share/67e39ffa-3a98-8011-ab79-fe3ac76632... Asking it to draw the Balkans map in Tolkien style, this is actually really impressive, geography is more or less completely correct, borders and country locations are wrong, but it feels like something I could get it to fix.

Strange, I'm getting > I wasn't able to generate the map because the request didn't follow content policy guidelines. Let me know if you'd like me to adjust the request or suggest an alternative way to achieve a similar result. Are you in the US? ...why are we living in such a retarded sci-fi age

No, I'm in Croatia. Just tried again and it's working https://chatgpt.com/share/67e3d18a-75e0-8011-ba67-fdcd13aa7f...

Re: 4o Image Generation

#462
post #414
post #413

Earlier quoted context omitted.

Why are you making blanket statements on things that you haven't even tried? This is leaps and bounds better than before.

No offense but do you believe it when microsoft announces they have solved quantum computing? And to prove it they only need your email address, birth date, credit card number, and rights to first born child?

No offense but you are really obnoxious.

Re: 4o Image Generation

#463
post #411
post #407

Earlier quoted context omitted.

The examples they show have little captions that say "best of #", like "best of 8" or "best of 4". Hopefully that truly represents the odds of generating the level of quality shown.

I'm not doubting it's an improvement, because it looks like it is. I guess here's an example of a prompt I would like to see: A flying spaghetti monster with a metal colander on its head flying above New York City saving the world from and very very evil Pope. I'm not anti/pro spaghetti monster or catholicism. But I can visualize it clearly in my head what that prompt might look like.

Here you go: https://chatgpt.com/share/67e3d3dc-b234-8004-b992-b559fc5038...

Re: 4o Image Generation

#464
The whiteboard image is insane. Even if it took more than 8 to find it, it's really impressive.

To think that a few years ago we had dreamy pictures with eyes everywhere. And not long ago we were always identifying the AI images by the 6 fingered people.

I wonder how well the physics is modeled internally. E.g. if you prompt it to model some difficult ray tracing scenario (a box with a separating wall and a light in one of the chambers which leaks through to the other chamber etc)?

Or if you have a reflective chrome ball in your scene, how well does it understand that the image reflected must be an exact projection of the visible environment?

Re: 4o Image Generation

#465

Earlier quoted context omitted.

That's the point. With the old models they all failed to produce a wine glass that is completley to the brim full. Because you can't find that a lot in the data they used for training.

Imagine if they just actually trained the model on a bunch of photographs of a full glass of wine, knowing of this litmus test

They still can't generate a watch that shows arbitrary times I believe, so it could be the case?

Re: 4o Image Generation

#467
> All generated images come with C2PA metadata

How easy is this to remove? Is it just like exif data that can be easily stripped out, or is it baked in more permanently somehow

Re: 4o Image Generation

#468
post #264
post #104

What's important about this new type of image generation that's happening with tokens rather than with diffusion, is that this is effectively reasoning in pixel space. Example: Ask it to draw a notepad with an empty tic-tac-toe, then tell it to make the first move, then you make a move, and so on. You can also do very impressive information-conserving translations, such as changing the drawing style, but also stuff l…

> truly generative UI, where the model produces the next frame of the app Please sir step away from the keyboard now! That is an absurd proposition and I hope I never get to use an app that dreams of the next frame. Apps are buggy as they are, I don't need every single action to be interpreted by LLM. An existing example of this is that AI Minecraft demo and it's a literal nightmare.

First it will dream up the interaction frame by frame. Next, to improve efficiency, it will cache those interaction representations. What better way to do that than through a code representation.

While I think current AI can’t come close to anything remotely usable, this is a plausible direction for the future. Like you, I shudder.

Re: 4o Image Generation

#469
post #321

The new model in the drop down says something like "4o Create Image (Updated)". It is truly incredible. Far better than any other image generator as far as understanding and following complex prompts. I was blown away when they showed this many months ago, and found it strange that more people weren't talking about it. This is much more precise than the Gemini one that just came out recently.

> found it strange that more people weren't talking about it.

Some simply dislike everything OpenAI. Just like everything Musk or Trump.

Re: 4o Image Generation

#470

Earlier quoted context omitted.

It very much looks like a side effect of this new architecture. In my experience, text looks much better in recent DALL-E images (so what ChatGPT was using before), but it is still noticeably mangled when printing more than a few letters. This model update seems to improve text rendering by a lot, at least as long as the content is clearly specified. However, when giving a prompt that requires the model to come up wi…

The periodic table is absolutely hilarious, I didn't know LLMs had finally mastered absurdist humor.

Yeah who wouldn't love a dip in the sulphur pool. But back to the question, why can't such a model recognize letters as such? It cannot be trained to pay special attention to characters? How come it can print an anatomically correct eye but not differentiate between P and Z?
Post reply on HN