Live data from Hacker News

4o Image Generation

openai.com

531–540 of 629 posts

Re: 4o Image Generation

#531
ChatGPT Pro tip: In addition to video generation, you can use this new image gen functionality in Sora and apply all of your custom templates to it! I generated this template (using my Sora Preset Generator, which I think is public) to test reasoning and coherency within the image:

Theme: Educational Scientific Visualization – Ultra Realistic Cutaways Color: Naturalistic palettes that reflect real-world materials (e.g., rocky grays, soil browns, fiery reds, translucent biological tones) with high contrast between layers for clarity Camera: High-resolution macro and sectional views using a tilt-shift camera for extreme detail; fixed side angles or dynamic isometric perspective to maximize spatial understanding Film Stock: Hyper-realistic digital rendering with photogrammetry textures and 8K fidelity, simulating studio-grade scientific documentation Lighting: Studio-quality three-point lighting with soft shadows and controlled specular highlights to reveal texture and depth without visual noise Vibe: Immersive and precise, evoking awe and fascination with the inner workings of complex systems; blends realism with didactic clarity Content Transformation: The input is transformed into a hyper-detailed, realistically textured cutaway model of a physical or biological structure—faithful to material properties and scale—enhanced for educational use with visual emphasis on internal mechanics, fluid systems, and spatial orientation

Examples: 1. A photorealistic geological cutaway of Earth showing crust, tectonic plates, mantle convection currents, and the liquid iron core with temperature gradients and seismic wave paths. 2. An ultra-detailed anatomical cross-section of the human torso revealing realistic organs, vasculature, muscular layers, and tissue textures in lifelike coloration. 3. A high-resolution cutaway of a jet engine mid-operation, displaying fuel flow, turbine rotation, air compression zones, and combustion chamber intricacies. 4. A hyper-realistic underground slice of a city showing subway lines, sewage systems, electrical conduits, geological strata, and building foundations. 5. A realistic cutaway of a honeybee hive with detailed comb structures, developing larvae, worker bee behavior zones, and active pollen storage processes.

Re: 4o Image Generation

#532
post #57

> Introducing 4o Image Generation: [...] our most advanced image generator yet Then google: > Gemini 2.5: Our most intelligent AI model > Introducing Gemini 2.0 | Our most capable AI model yet I could go on forever. I hope this trend dies and apple starts using something effective so all the other companies can start copying a new lexicon.

Has post-Jobs Apple ever come up with anything that would warrant this hope?

Apple isn't really the best software company and though they were early to digital assistants with Siri, it seems like they've let it languish. It's almost comical how bad Siri is given the capabilities of modern AI. That being said, Android doesn't really have a great builtin solution for this either.

Apple is more of a hardware company. Still, Cook does have a few big wins under his belt: M-series ARM chips on Macs, Airpods, Apple watch, Apple pay.

Re: 4o Image Generation

#533

Earlier quoted context omitted.

>The tool name is not relevant. It isn't the actual name, they use an obfuscated name. >EDIT: And googling the tool name I see it's already been widely discussed on twitter and elsewhere I am so confused by this thread.

The original claim was that the new image generation is direct multimodal output, rather than a second model. People provided evidence from the product, including outputs of the model that indicate it is likely using a tool. It's very easy to confirm that that's the case in the API, and it's now widely discussed elsewhere. It's possible the tool is itself just gpt4o, wrapped for reliability or safety or some other re…

> It's possible the tool is itself just gpt4o, wrapped for reliability or safety or some other reason, but it's definitely calling out at the model-output level

That's probably right. It allows them to just swap it out for DALL-E, including any tooling/features/infrastructure hey have built up around image generation, and they don't have to update all their 4o instances to this model which, who knows, may be not be ready for other tasks anyway or different enough to warrant testing before a rollout, or more expensive, etc.

Honestly it seems like the only sane way to roll it out if it is a multimodal descendant of 4o.

Re: 4o Image Generation

#534
It’s pretty good, the interesting thing is when it fails it seems to often be able to reason about what went wrong. So when we get CoT scaffolding for this it’ll be incredibly competent.

Re: 4o Image Generation

#535
post #104

What's important about this new type of image generation that's happening with tokens rather than with diffusion, is that this is effectively reasoning in pixel space. Example: Ask it to draw a notepad with an empty tic-tac-toe, then tell it to make the first move, then you make a move, and so on. You can also do very impressive information-conserving translations, such as changing the drawing style, but also stuff l…

A lot of convoluted explanations about something we don't even know if it really works all the time. I feel like in the third year of LLM-Hype and after reminde-me-how-many billions of dollars burned, we should by now not have to imagine what 'might happen' down to road, it should have been happening already. The use-case you are describing, sure sounds very interesting, until I remember asking asked Copilot for a si…

Consider using a better AI IDE platform than Copilot ... cursor, windsurf, cline, all great options that do much better than what you're describing. The underlying LLM capabilities also have advanced quite a bit in the past year.

Re: 4o Image Generation

#536
post #104

What's important about this new type of image generation that's happening with tokens rather than with diffusion, is that this is effectively reasoning in pixel space. Example: Ask it to draw a notepad with an empty tic-tac-toe, then tell it to make the first move, then you make a move, and so on. You can also do very impressive information-conserving translations, such as changing the drawing style, but also stuff l…

A lot of convoluted explanations about something we don't even know if it really works all the time. I feel like in the third year of LLM-Hype and after reminde-me-how-many billions of dollars burned, we should by now not have to imagine what 'might happen' down to road, it should have been happening already. The use-case you are describing, sure sounds very interesting, until I remember asking asked Copilot for a si…

That scene is changing so quickly that you will want to try again right now if you can.

While LLM code generation is very much still a mixed bag, it has been a significant accelerator in my own productivity, and for the most part all I am using is o1 (via the openAI website), deepseek, and jetbrains' AI service (Copilot clone). I'm eager to play with some of the other tooling available to VS Code users (such as cline)

I don't know why everyone is so eager to "get to the fun stuff". Dev is supposed to be boring. If you don't like it maybe you should be doing something else.

Re: 4o Image Generation

#537

Earlier quoted context omitted.

Strange, I'm getting > I wasn't able to generate the map because the request didn't follow content policy guidelines. Let me know if you'd like me to adjust the request or suggest an alternative way to achieve a similar result. Are you in the US? ...why are we living in such a retarded sci-fi age

No, I'm in Croatia. Just tried again and it's working https://chatgpt.com/share/67e3d18a-75e0-8011-ba67-fdcd13aa7f...

Weird, I can't see the image at all.

Edit: Eventually it showed up

Re: 4o Image Generation

#538
post #104

What's important about this new type of image generation that's happening with tokens rather than with diffusion, is that this is effectively reasoning in pixel space. Example: Ask it to draw a notepad with an empty tic-tac-toe, then tell it to make the first move, then you make a move, and so on. You can also do very impressive information-conserving translations, such as changing the drawing style, but also stuff l…

> truly generative UI, where the model produces the next frame of the app I built this exact thing last month, demo: https://universal.oroborus.org (not viable on phone for this demo, fine on tablet or computer) Also see discussion and code at: http://github.com/snickell/universal I wasn't really planning to share/release it today, but, heck, why not. I started with bitmap-style generative image models, but because t…

This is so impressive. You are building the future!

Re: 4o Image Generation

#539

Earlier quoted context omitted.

A lot of convoluted explanations about something we don't even know if it really works all the time. I feel like in the third year of LLM-Hype and after reminde-me-how-many billions of dollars burned, we should by now not have to imagine what 'might happen' down to road, it should have been happening already. The use-case you are describing, sure sounds very interesting, until I remember asking asked Copilot for a si…

Consider using a better AI IDE platform than Copilot ... cursor, windsurf, cline, all great options that do much better than what you're describing. The underlying LLM capabilities also have advanced quite a bit in the past year.

Well I do not really use it that much to actually care, and don't really depend on AI, thankfully. If they did not mess up the google search, we wouldnt even need that crap at all. But that's not the main point. Even if I switched to cursor or windsurf - aren't they all using one of the same LLMs? (ChatGPT,Claude, whatever..). The issue is that the underlying general approach will never be accurate enough. There is a reason most of successful technologies lift off quickly and those not successful also die very quickly. This is a tech propped up by a lot of VC money for now, but at some point, even the richest of the rich VCs will have trouble explaining spending 500B dollars in total, to get something like 15B revenue (not even profit). And don't even get me started on Altman's trillion-fantasies...

Re: 4o Image Generation

#540

Ran through some of my relatively complex prompts combined with using pure text prompts as the de-facto means of making adjustments to the images (in contrast to using something like img2img / inpainting / etc.) https://mordenstar.com/blog/chatgpt-4o-images It's definitely impressive though once again fell flat on the ability to render a 9-pointed star.

Armless Venus with bread is true art
Post reply on HN