Live data from Hacker News

ChatGPT Images 2.0

openai.com

581–590 of 1001 posts

Re: ChatGPT Images 2.0

#582
post #338

Earlier quoted context omitted.

In short, yes. A modern laptop is running almost fanless, like a 486 from the days of yore. A single H200 pumps out 700W continuously in a data center, and you run thousands of them. Also, don't forget the training and fine tuning runs required for the models. Mass transportation / global logistics can be very efficient and cheap. Before the pandemic, it was cheaper to import fresh tomatoes from half-world away rathe…

these are unfair comparisons. it's not just a single laptop running all day it's all the graphic designer laptops that get replaced. it's not a single container of painting supplies it's all off them, (which are toxic by the way). so if power were plentiful and environmental you'd be onboard with it?

> these are unfair comparisons. it's not just a single laptop running all day it's all the graphic designer laptops that get replaced. it's not a single container of painting supplies it's all off them, (which are toxic by the way).

Please see my other comment about energy consumption and connect the dots with how open loop DLC systems are harmful to fresh water supplies (which is another comment of mine).

> so if power were plentiful and environmental you'd be onboard with it?

This is a pretty loaded way to ask this. Let me put this straight. I'm not against AI. I'm against how this thing is built. Namely:

    - Use of copyrighted and copylefted materials to train models and hiding under "fair use" to exploit people.
      - Moreover, belittling of people who create things with their blood sweat and tears and poorly imitating their art just for kicks or quick bucks.
    - Playing fast and loose with environment and energy consumption without trying to make things efficiently and sustainably to reduce initial costs and time to market.
    - Gaslighting the users and general community about how these things are built, and how it's a theater, again to make people use this and offload their thinking, atrophying their skills and making them dependent on these.
I work in HPC. I support AI workloads and projects, but the projects we tackle have real benefits, like ecosystem monitoring, long term climate science, water level warning and prediction systems, etc. which have real tangible benefits for the future of the humanity. Moreover, there are other projects trying to minimize environmental impact of computation which we're part of.

So it's pretty nuanced, and the AI iceberg goes well below OpenAI/Anthropic/Mistral trio.

Re: ChatGPT Images 2.0

#584

Earlier quoted context omitted.

If that energy is used for research, maybe. If used to answer customer questions or generate Studio Ghibli knock-offs, it's not worth it, even a bit .

what’s the difference between those two? how can you say one has more value than the other?

One is trying to save the future of the planet and the humanity with science, the other one is mocking a man who devoted his whole life to his art, even if it means spending years to perfect a three-second sequence for kicks and monies.

If you see no difference between them, I can't continue to discuss this with you, sorry.

Re: ChatGPT Images 2.0

#585
post #513

Earlier quoted context omitted.

Did it correctly follow the instructions? Don't know my pokemon well enough.

Essentially yes (bottom got distorted), but Gemini uses Nano Banana Pro or Nano Banana 2 so it's not a surprising result. The image I linked uses the raw API.

Note that the styles are different; there are two digit images rendered in color.

Color charcoal drawings do exist, but it’s not what’s usually meant by “charcoal drawing”.

Re: ChatGPT Images 2.0

#586
post #131

Earlier quoted context omitted.

It's usually based on what they've been trained on. There aren't very many models that'll do higher resolutions outside of Seedream but adherency is worse.

Need a model trained on closeup/macro shots of everything, to use for upscaling, then run that, as a kernel, over the whole image.

Exactly what I was thinking

Re: ChatGPT Images 2.0

#587
post #202

This seems like a great time to mention C2PA, a specification for positively affirming image sources. OpenAI participates in this, and if I load an image I had AI generate in a C2PA Viewer it shows ChatGPT as the source. Bad actors can strip sources out so it's a normal image (that's why it's positive affirmation), but eventually we should start flagging images with no source attribution as dangerous the way we flag…

> but eventually we should start flagging images with no source attribution as dangerous the way we flag non-https. Yes, lets make all images proprietary and locked behind big tech signatures. No more open source image editors or open hardware.

Why would the image itself have to be proprietary to have some new piece of metadata attached to it ?

Re: ChatGPT Images 2.0

#588
post #252

Here is my regular "hard prompt" I use for testing image gen models: "A macro close-up photograph of an old watchmaker's hands carefully replacing a tiny gear inside a vintage pocket watch. The watch mechanism is partially submerged in a shallow dish of clear water, causing visible refraction and light caustics across the brass gears. A single drop of water is falling from a pair of steel tweezers, captured mid-splas…

I mean, your prompt is basically this skit: https://www.youtube.com/watch?v=BKorP55Aqvg ("The Expert" 7 red lines: all strictly perpendicular, some with green ink some with transparent ink) I couldn't imagine the image you were describing. I've listed some of the red lines with green ink I've noticed in your prompt: Macro Close Up - Sharp throughout Focus on tiny gear - But also on tweezers, old watchmakers hand, wat…

The last point (reflection by front glass versus mechanism access so no front glass) is the only issue I see with it. Other than that I can easily visualize an image that satisfies the prompt. I think that the general idea is a good one because it's satisfable while having multiple competing requirements that impose geometric constraints on the scene without providing an immediate solution to said constraints as well as requiring multiple independent features (caustics, reflections, fluid dynamics, refraction, directional lighting) that are quite complicated to get right.

To illustrate that there aren't any contradictions (other than the final bit about the reflection in the glass). Consider a macro shot showing partial hands, partial tweezers, and pocket watch internals. That's much is certainly doable. Now imagine the partial left hand holding a half submerged pocket watch, fingertips of right hand holding front half of tweezers that are clasping a tiny gear, positioned above the work piece with the drop of water falling directly below. Capture the watchmaker's perspective. I could sketch that so an image model capable of 3D reasoning should have no trouble.

It's precisely the sort of scene you'd use to test a raytracer. One thing I can immediately think to add is nested dielectrics. Perhaps small transparent glass beads sitting at the bottom of the dish of water with the edge of the pocket watch resting on them, make the dish transparent glass, and place the camera level with the top of the dish facing forward?

https://blog.yiningkarlli.com/2019/05/nested-dielectrics.htm...

A second thing I can think to add is a flame. Perhaps place a tealight candle on the far side of the dish, the flame visible through (and distorted by) the water and glass beads?

Re: ChatGPT Images 2.0

#589

The image of the messy desktop with the ASCII art is so impressive - the text renders, the date is consistent, it actually generated ASCII art in "ChatGPT", etc. I was skeptical that it was cherry-picked but was able to generate something very similar and then edit particular parts on the desktop (i.e. fixing content in the browser window and making the ASCII dog "more dog like"). It's honestly astounding, to me at l…

the neofetch for apple logo is messed up, though. the characters rendering that don't exist.

Re: ChatGPT Images 2.0

#590

Earlier quoted context omitted.

Are you asking if the 10 seconds it takes AI to generate an image is more costly to the environment than a commissioned graphics artist using a laptop for 5-6 hours, or a painter who uses physical media sourced from all over the world?

In short, yes. A modern laptop is running almost fanless, like a 486 from the days of yore. A single H200 pumps out 700W continuously in a data center, and you run thousands of them. Also, don't forget the training and fine tuning runs required for the models. Mass transportation / global logistics can be very efficient and cheap. Before the pandemic, it was cheaper to import fresh tomatoes from half-world away rathe…

This argument is so flawed that its conclusion almost loops back around to being correct again:

No, in terms of unit economics, I'm almost certain that the painting supplies have a bigger ecological/resource footprint than an LLM per icon generated, and I'm pretty sure the cost of shipping tomatoes does not decrease that footprint, even if it possibly dwarfs it.

But yes, due to Jevon's paradox, the total resource use might well increase despite all that. I, for example, would have never commissioned a professional icon for my silly little iOS shortcuts on my homescreen, so my silly icon related carbon footprint went from exactly zero to slightly above that.

Post reply on HN