Live data from Hacker News

ChatGPT Images 2.0

openai.com

881–890 of 1001 posts

Re: ChatGPT Images 2.0

#881

Earlier quoted context omitted.

How is it that a model can produce what must be near 1:1 images ripped straight out of Pokemon Fire Red (The first ones) for profit and not be infringing copyright. I know that's the game, but it seems CRAZY to me that they can do this.

The funny this is the main complaint I’ve heard so far is that it repeatedly refused to operate on original content… because it might violate copyright.

Yeah, the CSAM generated by grok proves the guardrails are only really good for stymieing benign uses.

Re: ChatGPT Images 2.0

#882
post #624

Earlier quoted context omitted.

I do not think this is a good prompt or useful benchmark, but nonetheless, it seems to work better for me: https://chatgpt.com/share/69e88a94-ded8-8395-b5dc-abceb2f44d...

Huh, that is indeed better. If ChatGPT Images 2.0/gpt-2-image is more nondeterministic than usual, than that is in itself a useful data point.

Did you enable thinking for your experiment? Are you sure you were on the 2.0 rather than 1.5 version?

Re: ChatGPT Images 2.0

#883
post #750

Earlier quoted context omitted.

This is, in my opinion, attempting to say the right thing with entirely the wrong perspective: The people you say are getting "shafted" always got shafted. Their works are the inspiration for all artists and people who lay their eyes on it - maybe they got paid when they made the work, maybe they managed to sell it, but probably not. And still, other artists (and machines) will use remember and be inspired by it, som…

Children can draw without ever having been to an art gallery. The IP laundromats need the entire stolen corpus of human labor. The latter is clearly an infringing derivative work. It will be true no matter who many bribes those who have never created anything pay to Marsha Blackburn (who miraculously reversed her AI skepticism). I wonder how many threats of being primaried have been issued by the uncreative technocra…

[deleted]

Re: ChatGPT Images 2.0

#884

OpenAI’s gpt-image-1.5 and Google’s NB2 have been pretty much neck and neck on my comparison site which focuses heavily on prompt adherence, with both hovering around a 70% success rate on the prompts for generative and editing capabilities. With the caveat being that Gemini has always had the edge in terms of visual fidelity. That being said, gpt-image-1.5 was a big leap in visual quality for OpenAI and eliminated m…

Great website, 2 things: 1 - Gpt-image-2 seems to pass the Flat Earth test? (if not, I'm sure the paid thinking 2k version passes it). 2 - Since NB2 was earlier, many gold medals are assigned to it, even though now GI2 passes them too, example the Octopus test NB2 14 attempts but GI2 just 2 (BTW number of attempts should affect the score I guess?)

So if you zoom in (click the zoom button on the actual gpt-image-2 of the flat Earth), you’ll see that a lot of the people are anatomical impossibilities, which is one of the disallowed criteria on the list. The faces also look like melted candles.

This is one of those areas where even state-of-the-art models still struggle. You’re asking for a high level of detail at a per-person level, which means you end up with lots and lots of very small objects that all need to be rendered with convincing detail.

I should probably explain the scoring rubric better - it's in the (i) info icon. If you click the pass/fail button towards the top, it switches from a simple pass/fail view to a weighted score. That weighted score is based on three things: level of adherence to the prompt, visual fidelity, and the number of attempts.

I've tried to keep my criteria as objective as possible, but there's just a certain level of unavoidable subjectivity to it.

For example, with the octopus image: Even though the minimum criteria might be five tentacles covered, having all eight is much closer to the ideal of “an octopus,” so it usually gets bumped up to a higher rating (bronze, silver, gold).

Honestly, I think I agree that the gpt-image-2 probably should be upgraded to a gold medal. Thanks for pointing that out!

Re: ChatGPT Images 2.0

#885

Every improvement in image generation seems to reduce the value of the images themselves. When anything can be faked or created in seconds, what is an image really worth? With text or code, you can dig into a meaningful dialogue because their reality is digital too. But images become like the plain people to show up photo frames. I guess it's just a completely personal feeling.

This AI fatigue is called genflation.

Re: ChatGPT Images 2.0

#886

Earlier quoted context omitted.

I think prompts like this are where agentic workflows come in to play. If you asked it to do generate the first 64 prime numbers, AI tools could do that. If you asked it to draw a charcoal image of Pokemon 13, it could do that. If you asked it to add a white Menlo 13 on a black background to the top left corner of that image, it could do that. If you asked it to do that 63 more times, it could do those things, and if…

That's what makes it a fair evaluation of its limits

I mean asking these transformers to do maths has always been the wrong task. It's like we're now considering "it doesn't have x tools built with traditional code built in".

Though I suppose we're testing their model + agent harness here as well. It really _should_ have all of those tools/reasoning available to accomplish a task like the above without issue.

Re: ChatGPT Images 2.0

#887

So during my Nano Banana Pro experiments I wrote a very fun prompt that tests the ability for these image generation models to follow heuristics, but still requires domain knowledge and/or use of the search tool: Create a 8x8 contiguous grid of the Pokémon whose National Pokédex numbers correspond to the first 64 prime numbers. Include a black border between the subimages. You MUST obey ALL the FOLLOWING rules for th…

How is it that a model can produce what must be near 1:1 images ripped straight out of Pokemon Fire Red (The first ones) for profit and not be infringing copyright. I know that's the game, but it seems CRAZY to me that they can do this.

Training a model on a corpus which includes copyrighted images but which is not focussed primarily or exclusively on applications which violate copyright might be fair use in the US (so far, it seems that way.)

But that doesn't mean that producing outputs using the model so trained which are based on copyright-protected ones in ways which would violate copyright if produced by any other means doesn't still violate copyright. DMCA safe harbor might apply to the system owner (IIRC, the exact boundaries are fuzzy with UGC generated on the site by the provider’s systems rather than generated elsewhere and posted), so Google may not be liable for the infringement (though if it is actively searching for references online at generation and not relying on what is trained into the model, that would seem to weaken the case for that), but it's still an infringement.

Re: ChatGPT Images 2.0

#889
post #478

Earlier quoted context omitted.

Oh god yes, I've been trying to make a LLM Assisted Magic the Gathering card scanner... its been a hell of a time trying to get it to just OCR card names well....

Why would you use an LLM for OCR?

Because if it's multimodal, oops all transformers and they're pretty much best in class for ocr now, afaik?

Re: ChatGPT Images 2.0

#890

Earlier quoted context omitted.

I get your point, but it's not even really that. It's that an AI generated photo evokes the same feelings in me that human-made photographs do and I have to catch that and turn that off consciously.

It shouldn't bother you. Just enjoy stuff. It's ok to think computer art is pretty. It's not some kind of personal or societal moral failing.

well it is. what if you found out that your wife is actually a robot that you cant tell apart from real human. your real wife. well at least not by cutting her open. would you feel the same being with here?
Post reply on HN