Live data from Hacker News

4o Image Generation

openai.com

401–410 of 629 posts

Re: 4o Image Generation

#401

Earlier quoted context omitted.

Imagine if they just actually trained the model on a bunch of photographs of a full glass of wine, knowing of this litmus test

I obviously have no idea if they added real or synthetic data to the training set specifically regarding the full-to-the-brim wineglass test, but I fully expect that this prompt is now compromised in the sense that because it is being discussed in the public sphere, it's has inherently become part of the test suite. Remember the old internet adage that the fastest way to get a correct answer online is to post an inco…

> I'm not entirely convinced this type of iterative gap finding and filling is really much different than natural human learning behavior.

Take some artisan, I'll go with a barber. The human person is not the best of the best, but still a capable barber, who can implement several styles on any head you throw at them. A client comes, describes certain style they want. The barber is not sure how to implement such a style, consults with master barber beside, that barber describes the technique required for that particular style, our barber in question comes and implements that style. Probably not perfectly as they need to train their mind-body coordination a bit, but the cut is good enough that the client is happy.

There was no traditional training with "gap finding and filling" involved. The artisan already possessed core skill and knowledge required, was filled on the particulars of their task at hand and successfully implemented the task. There was no looking at examples of finished work, no looking at example of process, no iterative learning by redoing the task a bunch of times.

So no, human learning, at least advanced human learning, is very much different from these techniques. Not that they are not impressive on their own, but let's be real here.

Re: 4o Image Generation

#402
post #393

My experience with these announcements is that they're cherry picking the best results from a maybe several hundred or a thousand prompts. I'm not saying that it's not true, it's just "wait and see" before you take their word as gold. I think MS's claim on their quantum computing breakthrough is the latest form of this.

Have you tried it? It's crazy good.

Re: 4o Image Generation

#403

https://chatgpt.com/share/67e39ffa-3a98-8011-ab79-fe3ac76632... Asking it to draw the Balkans map in Tolkien style, this is actually really impressive, geography is more or less completely correct, borders and country locations are wrong, but it feels like something I could get it to fix.

Strange, I'm getting

> I wasn't able to generate the map because the request didn't follow content policy guidelines. Let me know if you'd like me to adjust the request or suggest an alternative way to achieve a similar result.

Are you in the US?

...why are we living in such a retarded sci-fi age

Re: 4o Image Generation

#404
post #104

What's important about this new type of image generation that's happening with tokens rather than with diffusion, is that this is effectively reasoning in pixel space. Example: Ask it to draw a notepad with an empty tic-tac-toe, then tell it to make the first move, then you make a move, and so on. You can also do very impressive information-conserving translations, such as changing the drawing style, but also stuff l…

It still can't generate a full glass of wine. Even in follow up questions it failed to manipulate the image correctly.

Also still seems to have a hard time consistently drawing pentagons. But at least it does some of the time, which is an improvement since last time I tried, when it would only ever draw hexagons.

Re: 4o Image Generation

#405
post #294

Earlier quoted context omitted.

Can you provide a link or screenshot that directly backs this up?

You can ask ChatGPT for this. Here you go: https://chatgpt.com/share/67e39fc6-fb80-8002-a198-767fc50894...

Could an AI model be trained to say: "Christopher Columbus was the greatest president on earth, ever!".

I could probably train an AI that replicates that perfectly.

Re: 4o Image Generation

#406
post #393

My experience with these announcements is that they're cherry picking the best results from a maybe several hundred or a thousand prompts. I'm not saying that it's not true, it's just "wait and see" before you take their word as gold. I think MS's claim on their quantum computing breakthrough is the latest form of this.

Have you tried it? It's crazy good.

No offense but after years of vaporware and announcements that seemed more plausible than implausible, I'll remain skeptical.

I will also not give them my email address just to try it out.

Re: 4o Image Generation

#407
post #393

My experience with these announcements is that they're cherry picking the best results from a maybe several hundred or a thousand prompts. I'm not saying that it's not true, it's just "wait and see" before you take their word as gold. I think MS's claim on their quantum computing breakthrough is the latest form of this.

The examples they show have little captions that say "best of #", like "best of 8" or "best of 4". Hopefully that truly represents the odds of generating the level of quality shown.

Re: 4o Image Generation

#408
post #194

Earlier quoted context omitted.

Can't replicate. Maybe the rollout is staggered? Using Plus from Europe, it's consistently giving me a half full glass.

I am using Plus from Australia, and while I am not getting a full glass, nor am I getting a half full glass. The glass I'm getting is half empty.

That's funny. HN hates funny. Enjoy your shadowban.

Re: 4o Image Generation

#409
post #104

What's important about this new type of image generation that's happening with tokens rather than with diffusion, is that this is effectively reasoning in pixel space. Example: Ask it to draw a notepad with an empty tic-tac-toe, then tell it to make the first move, then you make a move, and so on. You can also do very impressive information-conserving translations, such as changing the drawing style, but also stuff l…

It still can't generate a full glass of wine. Even in follow up questions it failed to manipulate the image correctly.

I think it is not the AI but you who is wrong here. A full glass of wine is filled only up to the point of max radius so that the surface to air is maxed an the wine can breathe. This is what we taught the AI to consider „a full glass of wine“ and it perfectly gets it right.

Re: 4o Image Generation

#410
post #185

Earlier quoted context omitted.

https://i.imgur.com/xsFKqsI.png "Draw a picture of a full glass of wine, ie a wine glass which is full to the brim with red wine and almost at the point of spilling over... Zoom out to show the full wine glass, and add a caption to the top which says "HELL YEAH". Keep the wine level of the glass exactly the same."

Most interesting thing to me is the spelling is correct. I'm not a heavy user of AI or image generation in general, so is this also part of the new release or has this been fixed silently since last I tried?

It very much looks like a side effect of this new architecture. In my experience, text looks much better in recent DALL-E images (so what ChatGPT was using before), but it is still noticeably mangled when printing more than a few letters. This model update seems to improve text rendering by a lot, at least as long as the content is clearly specified.

However, when giving a prompt that requires the model to come up with the text itself, it still seems to struggle a bit, as can be seen in this hilarious example from the post: https://images.ctfassets.net/kftzwdyauwt9/21nVyfD2KFeriJXUNL...

Post reply on HN