Live data from Hacker News

GPT Image 1.5

openai.com

171–180 of 272 posts

Re: GPT Image 1.5

#171
post #144

I have a "go to" prompt for images: > In the style of a 1970s book sci-fi novel cover: A spacer walks towards the frame. In the background his spaceship crashed on an icy remote planet. The sky behind is dark and full of stars. Nano banana pro via gemini did really well, although still way too detailed, and it then made a mess of different decades when I asked it to follow up: https://gemini.google.com/share/1902c11f…

You're just not describing what you want properly. Looks fine to me. Clearly you have something else in mind, so I think you're just not describing well. My tip would be to use actuall illustration language. Do you want a wide angle shot? What should depth of field be? Oil painting print? Ink illustration? What kind of printing style? Do you want a photo of the book or a pre-print proof? What kind of color scheme?

A professional artist wouldn't know what you want.

You didn't even specify an art style. 1970s sci-fi novel cover isn't a style. You'll find vastly different art styles from the 70s. If you're disappointed, it's because you're doing a shitty job describing what's in your head. If your prompt isn't at least a paragraph, you're going to just get random generic results.

Re: GPT Image 1.5

#172
post #104

I am very impressed a benchmark I like to run is have it create sprite maps, uv texture maps for an imagined 3d model Noticed it captured a megaman legends vibe .... https://x.com/AgentifySH/status/2001037332770615302 and here it generated a texture map from a 3d character https://x.com/AgentifySH/status/2001038516067672390/photo/1 however im not sure if these are true uv maps that is accurate as i dont have the 3d m…

> however im not sure if these are true uv maps I can tell you with 100% certainty they are not. For example, Crash doesn't have a backside for his torso. You could definitely make a model that uses these as textures, but you'd really have to force it and a lot of it would be stretched or look weird. If you want to go this approach, it would make a lot more sense to make a model, unwrap it, and use the wireframe UV m…

That's a remake model in a modern game. The original Crash was even simpler than that one.

Most of Crash in the first game was not textured; just vertex colours. Only the fur on his back and his shoelaces were textures at all.

Re: GPT Image 1.5

#173
post #144

I have a "go to" prompt for images: > In the style of a 1970s book sci-fi novel cover: A spacer walks towards the frame. In the background his spaceship crashed on an icy remote planet. The sky behind is dark and full of stars. Nano banana pro via gemini did really well, although still way too detailed, and it then made a mess of different decades when I asked it to follow up: https://gemini.google.com/share/1902c11f…

You're just not describing what you want properly. Looks fine to me. Clearly you have something else in mind, so I think you're just not describing well. My tip would be to use actuall illustration language. Do you want a wide angle shot? What should depth of field be? Oil painting print? Ink illustration? What kind of printing style? Do you want a photo of the book or a pre-print proof? What kind of color scheme? A…

The killer feature of LLMs is to be able to extrapolate what's really wanted from short descriptions.

Look again at Gemini's output, it looks like an actual book cover, it looks like an illustration that could be found on a book.

It takes on board corrections (albeit hilariously literaly).

Look at GPT image's output, it doesn't look anything like a book cover, and when prompted to say it got it wrong, just doubles down on what it was doing.

Re: GPT Image 1.5

#175

Earlier quoted context omitted.

my question to your anecdotal: who cares? not being fecicious, but who cares if someone reproduced your stuff and millions of people see your stuff? is the money that you want? is it the fame? because fame you will get, maybe not money... but couldn't there be another way?

As a professional cinematographer/photographer I am incredibly uncomfortable with people using my art without my permission for unknown ends. Doubly so when it’s venture backed private companies stealing from millions of people like me as they make vague promises about the capabilities of their software trained on my work. It doesn’t take much to understand why that makes me uncomfortable and why I feel I am entitled…

You should be proud your work will now be distilled enterally and an aspect of your work will forever influence the world

Re: GPT Image 1.5

#177
post #112

Was it ever explained or understood why ChatGPT Images always has (had?) that yellow cast?

Not always, it started at a very specific point. Studio Ghibli craze + reinforcement learning on the likes.

That's not how it works the model doesn't just update in real time to likes and besides it was already yellow upon release

Re: GPT Image 1.5

#178
post #30

It's still not available in the API despite them announcing the availability. They even linked to their Image Playground where it's also not available.. I updated my local playground to support it and I'm just handling the 404 on the model gracefully https://github.com/alasano/gpt-image-1-playground

My Enterprise account got an email 1.5 hours ago that it is available in API but my other accounts haven't gotten any email yet

Re: GPT Image 1.5

#179

I have a Nano Banana Pro blog post in the works expanding on my experiments with Nano Banana ( https://news.ycombinator.com/item?id=45917875 ). Running a few of my test cases from that post and the upcoming blog post through this new ChatGPT Image model, this new model is better than Nano Banana but MUCH worse than Nano Banana Pro which now nails the test cases that previously showed issues. The pricing is unclear bu…

I just tested GPT1.5. I would say the image quality is on par with NBP in my tests (which is surprising as the images in their trailer video are bad), but the prompt adherence is way worse, and its "world model" if you want to call it that is worse. For instance, I asked it for two people in a row boat and it had two people, but the boat was more like a coracle and they would barely fit inside it. Also: SUPER ANNOYIN…

I actually just finished running the Text-to-Image benchmark a few minutes ago. This matches my own testing as well. GPT-Image 1.5 is clearly a step up as an editing model, but it performed worse in purely generative tasks compared to its predecessor - dropping from 11 (out of 14) to 9.

Comparing NB Pro, GPT Image 1, and GPT Image 1.5

https://genai-showdown.specr.net/?models=o4,nbp,g15

Re: GPT Image 1.5

#180

Earlier quoted context omitted.

This showdown benchmark was and still is great, but an enormous grain of salt should be added to any model that was released after the showdown benchmark itself. Maybe everyone has a different dose of skepticism. Personally I'm not even looking at results for models that were released after the benchmark, for all this tells us, they might as well be one-trick ponies that only do well in the benchmark. It might be too…

I think training image models to pass these very specific tests correctly will be very difficult for any of these companies. How would they even do that?

Hire a professional Photoshop artist to manually create the "correct" images and then put the before and after photos into the training data. Or however they've been training these models thus far, i don't know.

And if that still doesn't get you there, hash the image inputs to detect if its one of these test photos and then run your special test-passer algo.

Post reply on HN