Live data from Hacker News

ChatGPT Images 2.0

openai.com

421–430 of 1001 posts

Re: ChatGPT Images 2.0

#421

Earlier quoted context omitted.

Democratizing visual communication is arguably useful, for instance helping people to create diagrams that illustrate a concept they wish to convey. This is contingent on the tech working sufficiently well that the visuals are more effective at communication than the text that went into producing them though.

It's always felt like way overhyping to call something "democratization" when it's something I could do as a middle schooler in 2005. It takes some skill to do very well but it's not like basic diagram creation isn't something people already could do for basically free (I create figures for my job all the time now and chatGPT is more expensive than tools I use for design). Commissioning high quality diagrams from a d…

You are making a mistake a lot of people make when talking about genAI helping others do work. I get that to you it is very easy to do, but there are other groups of people that are not able to do it. What you are saying is like a hobbyist carpenter saying that making a bedside table would take him one weekend to do, so he doesn't think it is okay for tables to be made via assembly line instead of hiring a carpenter to do it.

Re: ChatGPT Images 2.0

#422
post #15

I've been trying out the new model like this: OPENAI_API_KEY="$(llm keys get openai)" \ uv run https://tools.simonwillison.net/python/openai_image.py \ -m gpt-image-2 \ "Do a where's Waldo style image but it's where is the raccoon holding a ham radio" Code here: https://github.com/simonw/tools/blob/main/python/openai_imag... Here's what I got from that prompt. I do not think it included a raccoon holding a ham radio…

Thanks for the image, I will see their faces in my nightmares.

Re: ChatGPT Images 2.0

#423
post #15

I've been trying out the new model like this: OPENAI_API_KEY="$(llm keys get openai)" \ uv run https://tools.simonwillison.net/python/openai_image.py \ -m gpt-image-2 \ "Do a where's Waldo style image but it's where is the raccoon holding a ham radio" Code here: https://github.com/simonw/tools/blob/main/python/openai_imag... Here's what I got from that prompt. I do not think it included a raccoon holding a ham radio…

Thanks for the image, I will see their faces in my nightmares.

This happens all too frequently when you ask a GenAI model to create an image with a large crowd especially a “Where’s Waldo?” style scenes, where by definition you’re going to be examining individual faces very closely.

Re: ChatGPT Images 2.0

#424

So during my Nano Banana Pro experiments I wrote a very fun prompt that tests the ability for these image generation models to follow heuristics, but still requires domain knowledge and/or use of the search tool: Create a 8x8 contiguous grid of the Pokémon whose National Pokédex numbers correspond to the first 64 prime numbers. Include a black border between the subimages. You MUST obey ALL the FOLLOWING rules for th…

Why would you consider this a good prompt?

[flagged]

Re: ChatGPT Images 2.0

#425
post #252

Here is my regular "hard prompt" I use for testing image gen models: "A macro close-up photograph of an old watchmaker's hands carefully replacing a tiny gear inside a vintage pocket watch. The watch mechanism is partially submerged in a shallow dish of clear water, causing visible refraction and light caustics across the brass gears. A single drop of water is falling from a pair of steel tweezers, captured mid-splas…

Why would you consider this a good prompt?

[deleted]

Re: ChatGPT Images 2.0

#426
post #9

do they have anything similar to SynthID, or are they just pretending that problem doesn't exist? I know this is probably mega cherry-picked to look more impressive, but some of the images are terrifyingly realistic. They seem to have put a lot of effort into the lighting.

I feel like asking the image generators to mark AI images is the wrong way to go about it. It's like trying to maintain a blocklist. It seems better to me to have the major camera manufacturers or cell phones cryptographically sign their images as real.

Re: ChatGPT Images 2.0

#427

Earlier quoted context omitted.

Can't the expression come from the person prompting the AI and sometimes taking hours inpainting or tweaking the prompt to try get the exact image / expression they had in their mind? A good use I've found is to be able to make scenes from a dream you had into an image. If that's not an expression of something then I'm not sure anything is.

Notably, this process of struggle is meant to go away, to make room for instant satisfaction. This is really about some kind of expression consumerism. (And what will be lost along the way is meaning.)

I always find this argument to ring hollow. Maybe it's because I've been through it with too many technologies already. Digital photography took out the art of film photography. CGI took out the wonder of practical effects. Digital art takes out the important brush strokes of someone actually painting. The real answer always is the mediums can coexist and each will be good for expression in their own way.

I'm not sure you immediately lose meaning if someone can make a highly personalized version of something easily. The % of completely meaningless video after YouTube and tiktok came about has skyrocketed. The amount of good stuff to watch has gone up as well though.

Re: ChatGPT Images 2.0

#428
post #116

Been using the model for a few hours now. I'm actually reall impressed with it. This is the first time i've found value in an image model for stuff I actually do. I've been using it to build powerpoint slides, and mockups. It's CRAZY good at that.

Yeah, it's funny. I would expect to see more enthusiasm versus just basic run-of-the-mill, "oh, there it is". Leave it to the HN crowd. This is incredible. I don't even like OpenAI.

Re: ChatGPT Images 2.0

#429
post #51

This is not as exciting as previous models were, but it is incredibly good. I am starting to think that expressing thoughts in words clearly is probably the most important and general skill of the future.

> I am starting to think that expressing thoughts in words clearly is probably the most important and general skill of the future. Without question. AI will be indistinguishable from having a team. Communicating clearly has always and will always mattered. This, however, is even stronger. Because you can program and use logic in your communications. We're going to collectively develop absolutely wild command over ins…

How can AI be the amazing thing you say it is, but also too stupid to understand unless you get really good at communicating. Wouldn't better AI just mean it understands your ramblings better?

Re: ChatGPT Images 2.0

#430
post #398

Earlier quoted context omitted.

> What else? I used to have an assistant make little index-card sized agendas for gettogethers when folks were in town or I was organising a holiday or offsite. They used to be physical; now it's a cute thing I can text around so everyone knows when they should be up by (and by when, if they've slept in, they can go back to bed). AI has been good at making these. They don't need to be works of art, just cute and sill…

I'm not seeing how it takes more than 5 minutes to type up an itinerary. If you want to make it cute and silly, just change up the font and color and add some clip art. If this is the best use case that exists for AI image generation, I'm only further convinced the tech is at best largely useless.

> not seeing how it takes more than 5 minutes to type up an itinerary

Because I’ll then spend hours playing with the typography (because it’s fun) and making it look like whatever design style I’ve most recently read about (again, because it’s fun) and then fighting Word or Latex because I don’t actually know what I’m doing (less fun). Outsourcing it is the right move, particularly if someone else is handling requests for schedules to be adjusted. An AI handles that outsourcing quicker for low-value (but frequent) tasks.

> If this is the best use case that exists for AI image generation

I’ve also had good luck sketching a map or diagram and then having the AI turn it into something that looks clean.

Look, 99% of my use cases are e.g. making my cat gnaw on the Tetons or making a concert of lobsters watching Lady Gaga singing “I do it for the claws” or whatever so I can send two friends something stupid at 1AM. But there does appear to be a veneer of productivity there, and worst case it makes the world look a bit nicer.

Post reply on HN