Live data from Hacker News

ChatGPT Images 2.0

openai.com

991–1000 of 1001 posts

Re: ChatGPT Images 2.0

#991

Earlier quoted context omitted.

Thanks. That's a good one~ Lens type stuff that involves reflections/refraction is a neat challenge for generative models. I did some editing tests that involved replacing an apartment window with a mirror back when Nano-Banana Pro was released and was rather stunned by the results. https://mordenstar.com/blog/edits-with-nanobanana/#through-t...

That's great, though I wasn't even thinking at the scale of reflection or refraction. My test was if the image generators could come up with a novelty pair of glasses that incorporate the year digits into the shape of the frame itself with some whimsy, rather than just plop the numbers on top of regular boring frames. So something like [this]( https://p.kagi.com/proxy/oardefault.jpg?c=-4THVYblKrsgkzFTNE... ) rather t…

Oh, I see what you’re saying. I like these types of tests where you incorporate well-known objects from the training data into unusual geometries.

Kind of makes me want to take advantage of the multi-image editing capability, since you can use gpt-image-2 with multiple images.

Take a photo of an existing pair of glasses frames (maybe even snapped at an optometrist’s office) then take a picture of an animal, like a spider with an unusual number of eyes, or something like a flounder, where the eyes eventually migrate to the top of its body.

Then you could see if the system can realistically adapt the design and show how those glasses might look if they were redesigned for these unusual optical situations.

Re: ChatGPT Images 2.0

#992

Earlier quoted context omitted.

If it was about this why do OpenAI and Anthropic lose their minds when people are training off their output or trying to scrape their systems. I actually don't have an issue with training off the mass of everyones work if the models are open and free to build upon, it's locking them away and then throwing your toys out the pram when people try and do the same thing that bothers me.

Good question. I actually have a technical answer, believe it or not. Pre-training is: training a model from scratch on cheap data that sets the foundation of a model's capabilities. It produces a base model. Post-training is: training a base model further, using expensive specialized data, direct human input and elaborate high compute use methods to refine the model's behavior, and imbue it with the capabilities tha…

>Ones that the frontier labs have spent a lot of AI-specialized data, compute, labor and hours of R&D work on.

Granted thats time and money but it's an absolute minuscule amount of human hours compared to the scraped data.

We know this for a fact because of parallelization, work of hundreds of millions vs the work of 20-100 even of OpenAIs team worked for the entire lifetimes of the current team and the lifetimes of the offspring of that team and the lifetimes of their offspring even with several lifetimes they still wouldnt have even made a dent in recreating that initial scraped training data.

Re: ChatGPT Images 2.0

#993

Earlier quoted context omitted.

That's great, though I wasn't even thinking at the scale of reflection or refraction. My test was if the image generators could come up with a novelty pair of glasses that incorporate the year digits into the shape of the frame itself with some whimsy, rather than just plop the numbers on top of regular boring frames. So something like [this]( https://p.kagi.com/proxy/oardefault.jpg?c=-4THVYblKrsgkzFTNE... ) rather t…

Oh, I see what you’re saying. I like these types of tests where you incorporate well-known objects from the training data into unusual geometries. Kind of makes me want to take advantage of the multi-image editing capability, since you can use gpt-image-2 with multiple images. Take a photo of an existing pair of glasses frames (maybe even snapped at an optometrist’s office) then take a picture of an animal, like a sp…

Flounder might even work, since my initial complaints that the generated designs obscured the wearer's eyesight were met with solutions that just moved the offending eye to the side of the person's head :)

Re: ChatGPT Images 2.0

#994
post #77

Earlier quoted context omitted.

I'm honestly unsure what could be improved at this point. Consistency? So it fails less often? Based on the released images, (especially the one "screenshot" of the Mac desktop) I feel like the best images from this model are so visually flawless that the only way to tell they're fake is by reasoning about the content of the image itself (ex. "Apple never made a red iPhone 15, so this image is probably fake" or "Cost…

> I'm honestly unsure what could be improved at this point. That's because you're focusing a little bit too much on visual fidelity. It's still relatively trivial to create a moderately complex prompt and have it fail miserably. Even SOTA models only scored a 12 out of 15 on my benchmarks, and that was without me deliberately trying to "flex" to break the model. Here's one I just came up with: A Mercator projection o…

Good point.

So I guess while "realism" (or believability) is really good now, prompt adherence has much room for improvement.

(though put it another way, realism has always been "solved" if the model gets to output whatever it wants as long as it looks realistic, though now it looks less like a malfunction and more like an inattentive human mistake or oversight, so even when it gets it wrong it's hard to tell it's wrong without knowing what the prompt was)

Re: ChatGPT Images 2.0

#995

Earlier quoted context omitted.

> I'm honestly unsure what could be improved at this point. That's because you're focusing a little bit too much on visual fidelity. It's still relatively trivial to create a moderately complex prompt and have it fail miserably. Even SOTA models only scored a 12 out of 15 on my benchmarks, and that was without me deliberately trying to "flex" to break the model. Here's one I just came up with: A Mercator projection o…

Good point. So I guess while "realism" (or believability) is really good now, prompt adherence has much room for improvement. (though put it another way, realism has always been "solved" if the model gets to output whatever it wants as long as it looks realistic, though now it looks less like a malfunction and more like an inattentive human mistake or oversight, so even when it gets it wrong it's hard to tell it's wr…

> it's hard to tell it's wrong without knowing what the prompt was.

Yeah this is actually a huge point of frustration on reddit where lots of people post their "impressive generative images" but fail to disclose the prompts so the audience is only able to evaluate realism/fidelity and not how faithfully the model actually followed the prompt.

Re: ChatGPT Images 2.0

#996

Earlier quoted context omitted.

While I agree with you, hacker news audience is not in the middle of the bell curve. I get this sounds elitist - but tremendous percentage of population is happily and eagerly engaging with fake religious images, funny AI videos, horrible AI memes, etc. Trying to mention that this video of puppy is completely AI generated results in vicious defense and mansplaining of why this video is totally real (I love it when vi…

I recently shoulder-surfed a family member scrolling away on their social media feed, and every single image was obvious AI slop. But it didn't matter. She loved every single one, watched videos all the way through, liked and commented on them... just total zombie-consumption mode and it was all 100% AI generated. I've tried in the past pointing out that it's all AI generated and nothing is real, and they simply don'…

I'd be a bit more humble rather than terrified, because I enjoy some AI slop too, especially funny animals that remind me of my old pets' antics. There are levels of slop. But tasteless stuff with crap graphics plastered all over, loud edits or badly calibrated tts voices were already all over reels/tiktok long before AI, and people still liked that.

The unsettling thing on social media is the mind hijacking with the recommendation algo and scrolling motion that resembles a slot machine, more than the content itself.

Re: ChatGPT Images 2.0

#997

Earlier quoted context omitted.

Good question. I actually have a technical answer, believe it or not. Pre-training is: training a model from scratch on cheap data that sets the foundation of a model's capabilities. It produces a base model. Post-training is: training a base model further, using expensive specialized data, direct human input and elaborate high compute use methods to refine the model's behavior, and imbue it with the capabilities tha…

>Ones that the frontier labs have spent a lot of AI-specialized data, compute, labor and hours of R&D work on. Granted thats time and money but it's an absolute minuscule amount of human hours compared to the scraped data. We know this for a fact because of parallelization, work of hundreds of millions vs the work of 20-100 even of OpenAIs team worked for the entire lifetimes of the current team and the lifetimes of…

This is like trying to apply "labor theory of value" to datasets. It doesn't work any better there than it does in economics in general.

It doesn't matter how many human hours went into making a Twitter shitpost. What matters is: how much value does it add to pre-training run, and how easy is it to substitute it for another data source.

"Cheap data" has low training value and is easy to replace. Twitter shitposts are worthless except in aggregate. "Expensive data" is what has high training value and is hard to replace. Things like SFT traces, domain expert RLHF guidance, RLVR bits - that's what the "moat" is.

Re: ChatGPT Images 2.0

#998
post #682

Earlier quoted context omitted.

Is there a reason why you chose to post this comment for free, without rewards, knowing full well it's going to end up in the training data of some LLM in the future?

Well, the way intellectual property works, anything I write on the internet is, by default, all rights reserved. Different website's policies will impact this, of course, and different laws (and quirks like "fair use") as well, but in general, if I write a snippet of code like: printf("%p\n", 0xbeefbeef); /* insert awesome new compression algorithm here */ Then no, I'm not providing it for free. In fact, all rights a…

The question was about a comment you posted on this specific site, whose terms[0] say:

> By uploading any User Content you hereby grant and will grant Y Combinator and its affiliated companies a nonexclusive, worldwide, royalty free, fully paid up, transferable, sublicensable, perpetual, irrevocable license to copy, display, upload, perform, distribute, store, modify and otherwise use your User Content for any Y Combinator-related purpose in any form, medium or technology now known or later developed.

[0]: https://www.ycombinator.com/legal/

Re: ChatGPT Images 2.0

#999
post #335

the tragedy of image generating ai is that it is used to massively create what already exists instead of creating something truly unique - we need ai artists - and yeah, they will not be appreciated

Why would we need AI artists tho?

why would we need photographers, they just push a button? why would we need digital artists, they just use a computer?

its a new medium, doesnt matter if we like it or not (art also should not care if we like it or not), ai is here to stay. so lets find out if we even can create art with it, or not.

Re: ChatGPT Images 2.0

#1000
post #923

Earlier quoted context omitted.

Would your ideal world apply to humans as well? Like if I see some art in a museum and it inspires me to create some of my own, I would need to pay a licensing fee to the original artist? And what about the artists that inspired them ? There is no art in the world that sprang fully formed from one single person, without any influences. Should we reshape our economy to ensure knowledge and artistic provenance is maint…

>Like if I see some art in a museum and it inspires me to create some of my own, I would need to pay a licensing fee to the original artist? Nope, humans are admitted for free :). >And what about the artists that inspired them? There is no art in the world that sprang fully formed from one single person, without any influences. As long as you are a human you get to be inspired all you want :) You seem very invested i…

[dead]
Post reply on HN