Live data from Hacker News

4o Image Generation

openai.com

471–480 of 629 posts

Re: 4o Image Generation

#471
post #104

What's important about this new type of image generation that's happening with tokens rather than with diffusion, is that this is effectively reasoning in pixel space. Example: Ask it to draw a notepad with an empty tic-tac-toe, then tell it to make the first move, then you make a move, and so on. You can also do very impressive information-conserving translations, such as changing the drawing style, but also stuff l…

> truly generative UI, where the model produces the next frame of the app I built this exact thing last month, demo: https://universal.oroborus.org (not viable on phone for this demo, fine on tablet or computer) Also see discussion and code at: http://github.com/snickell/universal I wasn't really planning to share/release it today, but, heck, why not. I started with bitmap-style generative image models, but because t…

It’s like a lucid dream version of using and modifying the software at the same time.

Re: 4o Image Generation

#472

Ran through some of my relatively complex prompts combined with using pure text prompts as the de-facto means of making adjustments to the images (in contrast to using something like img2img / inpainting / etc.) https://mordenstar.com/blog/chatgpt-4o-images It's definitely impressive though once again fell flat on the ability to render a 9-pointed star.

Didn't work for me on the first prompting (got a 10-pointed one), but after sending [this is 10 points, make it 9] it did render a 9-pointed one too

Re: 4o Image Generation

#473

First AI image generator to pass the uncanny valley test? Seems like it. This is the biggest leap in image generation quality I've ever seen. How much longer until an AI that can generate 30 frames with this quality and make a movie? About 1.5 years ago, I thought AI would eventually allow anyone with an idea to make a Hollywood quality movie. Seems like we're not too far off. Maybe 2-3 more years?

Ideogram 2.0 and Recraft also create images that looks very much real.

For drawings, NovelAI's models are way beyond the uncanny valley now.

Re: 4o Image Generation

#475
post #405

Earlier quoted context omitted.

You can ask ChatGPT for this. Here you go: https://chatgpt.com/share/67e39fc6-fb80-8002-a198-767fc50894...

Could an AI model be trained to say: "Christopher Columbus was the greatest president on earth, ever!". I could probably train an AI that replicates that perfectly.

> Could an AI model be trained to say: "Christopher Columbus was the greatest president on earth, ever!".

Yes, it could. And even after training its data can be manipulated to output whatever: https://www.anthropic.com/news/mapping-mind-language-model

Re: 4o Image Generation

#476

First AI image generator to pass the uncanny valley test? Seems like it. This is the biggest leap in image generation quality I've ever seen. How much longer until an AI that can generate 30 frames with this quality and make a movie? About 1.5 years ago, I thought AI would eventually allow anyone with an idea to make a Hollywood quality movie. Seems like we're not too far off. Maybe 2-3 more years?

[deleted]

Re: 4o Image Generation

#477
post #235

Earlier quoted context omitted.

> What's important about this new type of image generation that's happening with tokens rather than with diffusion, is that this is effectively reasoning in pixel space. I do not think that this is correct. Prior to this release, 4o would generate images by calling out to a fully external model (DALL-E). After this release, 4o generates images by calling out to a multi-modal model that was trained alongside it. You c…

>You can ask 4o about this yourself. Here's what it said to me: >"So while I’m deeply multimodal in cognition (understanding and coordinating text + image), image generation is handled by a linked latent diffusion model, not an end-to-end token-unified architecture." Models don't know anything about themselves. I have no idea why people keep doing this and expecting it to know anything more than a random con artist o…

>Models don't know anything about themselves.

They can. Fine tune them on documents describing their identity, capabilities and background. Deepseek v3 used to present itself as ChatGPT. Not anymore.

>Like other AI models, I’m trained on diverse, legally compliant data sources, but not on proprietary outputs from models like ChatGPT-4. DeepSeek adheres to strict ethical and legal standards in AI development.

Re: 4o Image Generation

#478
post #31

It's incredible that this took 316 days to be released since it was initially announced. I do appreciate the emphasis in the presentation on how this can be useful beyond just being a cool/fun toy, as it seems most image generation tools have functioned. Was anyone else surprised how slow the images were to generate in the livestream? This seems notably slower than DALLE.

I've never minded that an image might take 10-30 seconds to generate. The fact that people do is crazy to me. A professional artist would take days, and cost $100s for the same asset.

I ran stable diffusion for a couple of years (maybe?, time really hasn't made sense since 2020) on my Dual 3090 rendering server. I built the server originally for crypto heating my office in my 1820s colonial in upstate NY then when I was planning to go back to college (got accepted into a university in England), I switched it's focus to Blender/UE4 (then 5), then eventually to AI image gen. So I've never minded 20 seconds for an image. If I needed dozens of options to pick the best, I was going to click start and grab a cup of coffee, come back and maybe it was done. Even if it took 2 hours, it is still faster than when I used to have to commission art for a project.

I grew out of Stable Diffusion, though, because the learning curve beyond grabbing a decent checkpoint and clicking start was actually really high (especially compared to LLMs that seamed to "just work"), after going through failed training after failed fine-tuning using tutorials that were a couple days out of date, I eventually said, fuck it, I'm paying for this instead.

All that to say - if you are using GenAI commercially, even if an image or a block of code took 30 minutes, it's still WAY cheaper than a human. That said, eventually a professional will be involved, and all the AI slop you generated will be redone, which will still cost a lot, but you get to skip the back and forth figuring out style/etc.

Re: 4o Image Generation

#479
post #414
post #413

Earlier quoted context omitted.

Why are you making blanket statements on things that you haven't even tried? This is leaps and bounds better than before.

No offense but do you believe it when microsoft announces they have solved quantum computing? And to prove it they only need your email address, birth date, credit card number, and rights to first born child?

I don't believe it when Microsoft announces it, but when two separate trustworthy-looking hn accounts tell me something is crazy good that seems like valuable information to me.
Post reply on HN