Live data from Hacker News

4o Image Generation

openai.com

301–310 of 629 posts

Re: 4o Image Generation

#301
post #248

It bothers me to see links to content that requires a login. I don't expect openai or anyone else to give their services away for free. But I feel like "news" posts that require one to setup an account with a vendor are bad faith. If the subject matter is paywalled, I feel that the post should include some explanation of what is newsworthy behind the link.

The linked post is not paywalled, did you click something else?

Thank you for the accurate correction. My whining was a bit unmerited. The link goes to a page that largely provides exactly what I asked for. It just starts out with an invitation to try it yourself. That invitation leads you to an app that requires a login. It was unfair of me to be triggered by that invitation.

After that invitation there are several examples that boil down to: "Hey look. Our AI can generate deep fakes." Impressive examples.

Re: 4o Image Generation

#302
post #268

Earlier quoted context omitted.

This argument could be made for every level of abstraction we've added to software so far... yet here we are commenting about it from our buggy apps!

Yeah, but the abstractions have been useful so far. The main advantage of our current buggy apps is that if it is buggy today, it will be exactly as buggy tomorrow. Conversely, if it is not currently buggy, it will behave the same way tomorrow. I don't want an app that either works or does not work depending on the RNG seed, prompt and even data that's fed to it. That's even ignoring all the absurd computing power th…

Still sounds a bit like we've seen it all already – dynamic linking introduced a lot of ways for software that wasn't buggy today to become buggy tomorrow. And Chrome uses an absurd amount of computing power (its bare minimum is many multiples of what was once a top-of-the-line, expensive PC).

I think these arguments would've been valid a decade ago for a lot of things we use today. And I'm not saying the classical software way of things needs to go away or even diminish, but I do think there are unique human-computer interactions to be had when the "VM" is in fact a deep neural network with very strong intelligence capabilities, and the input/output is essentially keyboard & mouse / video+audio.

Re: 4o Image Generation

#303
post #205

Anyone else frightened by this? Seeing meant believing, and now that isnt the case anymore...

Not really, but only because literally every person I've met that spends a lot of time on TikTok starts spouting unhinged nonsense at me.

You don't even need deepfakes. https://www.newsweek.com/doug-mastriano-pennsylvania-senator...

The disaster scenario is already here.

Re: 4o Image Generation

#305

Earlier quoted context omitted.

A term for people giving only milliseconds of their attention is: uninterested people. If I’m not looking for a project planner, or interested in the space, there’s no wording that can make me stay on an announcement for one. If I am, you can be sure I’m going to read the whole feature page.

Idealistic and wrong, marketing does work in a lot of cases and that's why everybody does it

No, everybody uses marketing because it's a conventional bet. It has proven in many cases to not be effective, but people aren't willing to risk getting fired because they suggested going against the grain.

Re: 4o Image Generation

#307

Earlier quoted context omitted.

I obviously have no idea if they added real or synthetic data to the training set specifically regarding the full-to-the-brim wineglass test, but I fully expect that this prompt is now compromised in the sense that because it is being discussed in the public sphere, it's has inherently become part of the test suite. Remember the old internet adage that the fastest way to get a correct answer online is to post an inco…

Humans don’t train on the entire contents of the Internet, so i’d wager that they do learn differently

I think there is a critical aspect of human visual learning which machine leanring cant replicate because it is prohibitively expensive. When we look at things as children we are not just looking at a single snapshot. When you stare at an object for a few seconds you have practically injested hundreds of slightly variated images of that object. This gets even more interesting when you take into account real world is moving all the time, so you are seeing so many things from so many angles. This is simply undoable with compute.

Re: 4o Image Generation

#308
post #235
post #104

What's important about this new type of image generation that's happening with tokens rather than with diffusion, is that this is effectively reasoning in pixel space. Example: Ask it to draw a notepad with an empty tic-tac-toe, then tell it to make the first move, then you make a move, and so on. You can also do very impressive information-conserving translations, such as changing the drawing style, but also stuff l…

> What's important about this new type of image generation that's happening with tokens rather than with diffusion, is that this is effectively reasoning in pixel space. I do not think that this is correct. Prior to this release, 4o would generate images by calling out to a fully external model (DALL-E). After this release, 4o generates images by calling out to a multi-modal model that was trained alongside it. You c…

I think this is actually correct even if the evidence is not right.

See this chat for example:

https://chatgpt.com/share/67e355df-9f60-8000-8f36-874f8c9a08...

Re: 4o Image Generation

#310

OpenAI's livestream of GPT-4o Image Generation shows that it is slowwwwwwwwww (maybe 30 seconds per image, which Sam Altman had to spin "it's slow but the generated images are worth it"). Instead of using a diffusion approach, it appears to be generating the image tokens and decoding them akin to the original DALL-E ( https://openai.com/index/dall-e/ ), which allows for streaming partial generations from top to botto…

i find this “slow” complaint (/observation— i dont view this comment as a complaint, to be clear) to be quite confusing. slow… compared to what, exactly? you know what is slow? having to prompt and reprompt 15 times to get the stupid model to spell a word correctly and it not only refuses, but is also insistent that it has corrected the error this time. and afaict this is the exact kind of issue this change should address substantially.

im not going to get super hyperbolic and histrionic about “entitlement” and stuff like that, but… literally this technology did not exist until like two years ago, and yet i hear this all the time. “oh this codegen is pretty accurate but it’s slow”, “oh this model is faster and cheaper (oh yeah by the way the results are bad, but hey it’s the cheapest so it’s better)”. like, are we collectively forgetting that the whole point of any of this is correctness and accuracy? am i off-base here?

the value to me of a demonstrably wrong chat completion is essentially zero, and the value of a correct one that anticipates things i hadn’t considered myself is nearly infinite. or, at least, worth much, much more than they are charging, and even _could_ reasonably charge. it’s like people collectively grouse about low quality ai-generated junk out of one side of their mouths, and then complain about how expensive the slop is out of the other side.

hand this tech to someone from 2020 and i guarantee you the last thing you’d hear is that it’s too slow. and how could it be? yeah, everyone should find the best deals / price-value frontier tradeoff for their use case, but, like… what? we are all collectively devaluing that which we lament is being devalued by ai by setting such low standards: ourselves. the crazy thing is that the quickly-generated slop is so bad as to be practically useless, and yet it serves as the basis of comparison for… anything at all. it feels like that “web-scale /dev/null” meme all over again, but for all of human cognition.

Post reply on HN