Live data from Hacker News

GPT Image 1.5

openai.com

141–150 of 272 posts

Re: GPT Image 1.5

#141

Okay results are in for GenAI Showdown with the new gpt-image 1.5 model for the editing portions of the site! https://genai-showdown.specr.net/image-editing Conclusions - OpenAI has always had some of the strongest prompt understanding alongside the weakest image fidelity. This update goes some way towards addressing this weakness. - It's leagues better at making localized edits without altering the entire image's ae…

> the only model that legitimately passed the Giraffe prompt. 10 years ago I would have considered that sentence satire. Now it allegedly means something. Somehow it feels like we’re moving backwards.

> Somehow it feels like we’re moving backwards.

I don't understand why everyone isn't in awe of this. This is legitimately magical technology.

We've had 60+ years of being able to express our ideas with keyboards. Steve Jobs' "bicycle of the mind". But in all this time we've had a really tough time of visually expressing ourselves. Only highly trained people can use Blender, Photoshop, Illustrator, etc. whereas almost everyone on earth can use a keyboard.

Now we're turning the tide and letting everyone visually articulate themselves. This genuinely feels like computing all over again for the first time. I'm so unbelievably happy. And it only gets better from here.

Every human should have the ability to visually articulate themselves. And it's finally happening. This is a major win for the world.

I'm not the biggest fan of LLMs, but image and video models are a creator's dream come true.

In the near future, the exact visions in our head will be shareable. We'll be able to iterate on concepts visually, collaboratively. And that's going to be magical.

We're going to look back at pre-AI times as primitive. How did people ever express themselves?

Re: GPT Image 1.5

#142

Okay results are in for GenAI Showdown with the new gpt-image 1.5 model for the editing portions of the site! https://genai-showdown.specr.net/image-editing Conclusions - OpenAI has always had some of the strongest prompt understanding alongside the weakest image fidelity. This update goes some way towards addressing this weakness. - It's leagues better at making localized edits without altering the entire image's ae…

This showdown benchmark was and still is great, but an enormous grain of salt should be added to any model that was released after the showdown benchmark itself.

Maybe everyone has a different dose of skepticism. Personally I'm not even looking at results for models that were released after the benchmark, for all this tells us, they might as well be one-trick ponies that only do well in the benchmark.

It might be too much work, but one possible "correct" approach for this kind of benchmark would to periodically release new benchmarks with new tests (that are broadly in the same categories) and only include models that predate each benchmark.

Re: GPT Image 1.5

#144
I have a "go to" prompt for images:

> In the style of a 1970s book sci-fi novel cover: A spacer walks towards the frame. In the background his spaceship crashed on an icy remote planet. The sky behind is dark and full of stars.

Nano banana pro via gemini did really well, although still way too detailed, and it then made a mess of different decades when I asked it to follow up: https://gemini.google.com/share/1902c11fd755

It's therefore really disappointing that GPT-image 1.5 did this:

https://chatgpt.com/share/6941ed28-ed80-8000-b817-b174daa922...

Completely generic, not at all like a book cover, it completely ignored that part of the prompt while it focused on the other elements.

Did it get the other details right? Sure, maybe even better, but the important part it just ignored completely.

And it's doing even worse when I try to get it to correct the mistake. It's just repeating the same thing with more "weathering".

Re: GPT Image 1.5

#145

Okay results are in for GenAI Showdown with the new gpt-image 1.5 model for the editing portions of the site! https://genai-showdown.specr.net/image-editing Conclusions - OpenAI has always had some of the strongest prompt understanding alongside the weakest image fidelity. This update goes some way towards addressing this weakness. - It's leagues better at making localized edits without altering the entire image's ae…

This showdown benchmark was and still is great, but an enormous grain of salt should be added to any model that was released after the showdown benchmark itself. Maybe everyone has a different dose of skepticism. Personally I'm not even looking at results for models that were released after the benchmark, for all this tells us, they might as well be one-trick ponies that only do well in the benchmark. It might be too…

Yeah that’s a classic problem, and it's why good tests are such closely guarded secrets: to keep them from becoming training fodder for the next generation of models. Regarding the "model date" vs "benchmark date" - that's an interesting point... I'll definitely look into it!

I don't have any captcha systems in place, but I wonder if it might be worth putting up at least a few nominal roadblocks (such as Anubis [1]) to at least slow down the scrapers.

A few weeks ago I actually added some new, more challenging tests to the GenAI Text-to-Image section of the site (the “angelic forge” and “overcrowded flat earth”) just to keep pace with the latest SOTA models.

In the next few weeks, I’ll be adding some new benchmarks to the Image Editing section as well~~

[1] - https://anubis.techaro.lol

Re: GPT Image 1.5

#146

Earlier quoted context omitted.

> the only model that legitimately passed the Giraffe prompt. 10 years ago I would have considered that sentence satire. Now it allegedly means something. Somehow it feels like we’re moving backwards.

> Somehow it feels like we’re moving backwards. I don't understand why everyone isn't in awe of this. This is legitimately magical technology. We've had 60+ years of being able to express our ideas with keyboards. Steve Jobs' "bicycle of the mind". But in all this time we've had a really tough time of visually expressing ourselves. Only highly trained people can use Blender, Photoshop, Illustrator, etc. whereas almos…

Where is all this wonderful visual self expression that people are now free to do? As far as I can tell it's mostly being used on LinkedIn posts.

Re: GPT Image 1.5

#147

Is there a watermarking, or some other way for normal people to tell if its fake?

I think society is going to need the opposite - cameras that can embed cryptographic information in the pixels of a video indicating the image is real.

Re: GPT Image 1.5

#148
post #46

If this was a farm of sweatshop Photoshopers in 2010, who download all images from the internet and provide a service of combining them on your request, this would escalate pretty quickly. Question: with copyright and authorship dead wrt AI, how do I make (at least) new content protected? Anecdotal: I had a hobby of doing photos in quite rare style and lived in a place where you'd get quite a few pictures of. When I…

> how do I make (at least) new content protected? Air gap. If you don’t want content to be used without your permission, it never leaves your computer. This is the only protection that works. If you want others to see your content, however, you have to accept some degree of trade off with it being misappropriated. Blatant cases can be addressed the same as they always were, but a model overfitting to your original wo…

Horror scenario:

Big IP holders will go nuclear on IP licensing to an extent we've never seen before.

Right now, there are thousands of images and videos of Star Wars, Pokemon, Superman, Sonic, etc. being posted across social media. All it takes is for the biggest IP conglomerates to turn into linear tv and sports networks of the past and treat social media like cable.

Disney: "Gee {Google,Meta,Reddit,TikTok}, we see you have a lot of Star Wars and Marvel content. We think that's a violation of our rights. If you want your users to continue to be able to post our media, you need to pay us $5B/yr."

I would not be surprised if this happens now that every user on the internet can soon create high-fidelity content.

This could be a new $20-30B/yr business for Disney. Nintendo, WBD, and lots of other giant IP holders could easily follow suit.

Re: GPT Image 1.5

#149

Earlier quoted context omitted.

> Somehow it feels like we’re moving backwards. I don't understand why everyone isn't in awe of this. This is legitimately magical technology. We've had 60+ years of being able to express our ideas with keyboards. Steve Jobs' "bicycle of the mind". But in all this time we've had a really tough time of visually expressing ourselves. Only highly trained people can use Blender, Photoshop, Illustrator, etc. whereas almos…

Where is all this wonderful visual self expression that people are now free to do? As far as I can tell it's mostly being used on LinkedIn posts.

It’s a classic issue that you give access to superpowers to the general population and most will use them in the most boring ways.

The internet is an amazing technology, yet its biggest consumption is a mix of ads, porn and brain rot.

We all have cameras in our pockets yet most people use them for selfies.

But if you look closely enough, the incredible value that comes from these examples more than makes up for all the people using them in a “boring” way.

And anyway who’s the arbiter of boring?

Re: GPT Image 1.5

#150
What is the endgame? Why is OpenAI throwing that much money on image/video generation? Is there a profitable market for AI-generated image slop? Do people choose ChatGPT instead of Gemini/Grok/Claude because of the image generation capabilities? To me, it looks like a huge fiery money pit.
Post reply on HN