Live data from Hacker News

4o Image Generation

openai.com

541–550 of 629 posts

Re: 4o Image Generation

#541

Earlier quoted context omitted.

Then explain blind children? Or blind & deaf children? There's obviously some role senses play in development but there's clearly capabilities at play here that are drastically more efficient and powerful than what we have with modern transformers. While humans learn through example, they clearly need a lot fewer examples to generalize off of and reason against.

they take in many samples of touch data

I think my point is that communication is the biggest contributor to brain development more than anything and communication is what powers our learning. Effective learners learn to communicate more with themselves and to communicate virtually with past authors through literature. That isn’t how LLMs work. Not sure why that would be considered objectionable. LLMs are great but we don’t have to pretend like they’re actually how brains work. They’re a decent approximation for neurons on today’s silicon - useful but nowhere near the efficiency and power of wetware.

Also as for touch, you’re going to have a hard time convincing me that the amount of data from touch rivals the amount of content on the internet or that you just learn about mistakes one example at a time.

Re: 4o Image Generation

#542

Earlier quoted context omitted.

A lot of convoluted explanations about something we don't even know if it really works all the time. I feel like in the third year of LLM-Hype and after reminde-me-how-many billions of dollars burned, we should by now not have to imagine what 'might happen' down to road, it should have been happening already. The use-case you are describing, sure sounds very interesting, until I remember asking asked Copilot for a si…

That scene is changing so quickly that you will want to try again right now if you can. While LLM code generation is very much still a mixed bag, it has been a significant accelerator in my own productivity, and for the most part all I am using is o1 (via the openAI website), deepseek, and jetbrains' AI service (Copilot clone). I'm eager to play with some of the other tooling available to VS Code users (such as cline…

I mean I literally "tried it again" this morning, as a paying Copilot customer of 12 months, to the result I already described. And I do not want to "try it" - based on fluffy promises we've been hearing, it should "just work". Are you old enough to remember that phrase? It was a motto introduced by an engineering legend whose devices you're likely using every day. The reason why "everyone", including myself with 20+ years of experience is looking to do not "fun stuff" (please don't shove words into my mouth), but cool stuff (=hard problems) is that it produces an intrinsic sense of satisfaction, which in turn creates motivation to do more and eventually even produces wider gain for the society. Some of us went into engineering because of passion you know. We're not all former copy-writers retrained to be "frontend developers" for a higher salary, who are eager to push around CSS boxes. That's important work too, but I've definitely solved harder problems in my career. If it's boring for you and you think it's how it should be, then you are definitely doing it for the wrong reasons (I am assuming escaping a less profitable career).

Re: 4o Image Generation

#543
post #104

What's important about this new type of image generation that's happening with tokens rather than with diffusion, is that this is effectively reasoning in pixel space. Example: Ask it to draw a notepad with an empty tic-tac-toe, then tell it to make the first move, then you make a move, and so on. You can also do very impressive information-conserving translations, such as changing the drawing style, but also stuff l…

Is it able to break the usual failure modes of these models, that all clocks are at 10 min past two, or they can't produce images of people drawing with the left hand?

Re: 4o Image Generation

#544

Earlier quoted context omitted.

Strange, I'm getting > I wasn't able to generate the map because the request didn't follow content policy guidelines. Let me know if you'd like me to adjust the request or suggest an alternative way to achieve a similar result. Are you in the US? ...why are we living in such a retarded sci-fi age

No, I'm in Croatia. Just tried again and it's working https://chatgpt.com/share/67e3d18a-75e0-8011-ba67-fdcd13aa7f...

Strange, I can see your second image but not the first.

Re: 4o Image Generation

#545

Earlier quoted context omitted.

they take in many samples of touch data

I think my point is that communication is the biggest contributor to brain development more than anything and communication is what powers our learning. Effective learners learn to communicate more with themselves and to communicate virtually with past authors through literature. That isn’t how LLMs work. Not sure why that would be considered objectionable. LLMs are great but we don’t have to pretend like they’re act…

There are so many points to consider here im not sure i can address them all.

- Airplanes dont have wings like birds but can fly. and in some ways are superior to birds. (some ways not)

- Human brains may be doing some analogue of sample augmentation which gives you some multiple more equivalent samples of data to train on per real input state of environment. This is done for ml too.

- Whether that input data is text, or embodied is sort of irrelevant to cognition in general, but may be necessary for solving problems in a particular domain. (text only vs sight vs blind)

Re: 4o Image Generation

#546
post #57

> Introducing 4o Image Generation: [...] our most advanced image generator yet Then google: > Gemini 2.5: Our most intelligent AI model > Introducing Gemini 2.0 | Our most capable AI model yet I could go on forever. I hope this trend dies and apple starts using something effective so all the other companies can start copying a new lexicon.

This actually makes sense because the versioning is so confusing they could be releasing a lesser/lightweight model for all we know.

Re: 4o Image Generation

#547
post #408

Earlier quoted context omitted.

That's funny. HN hates funny. Enjoy your shadowban.

Yeah. I understand that this site doesn’t want to become Reddit, but it really has an allergy to comedy, it’s sad. God forbid you use sarcasm, half the people here won’t understand it and the other half will say it’s not appropriate for healthy discussion…

Good example in this very discussion: https://news.ycombinator.com/item?id=43477003

Re: 4o Image Generation

#548
post #238
post #228

Earlier quoted context omitted.

share prompt minus identifying details?

> Draw a birthday invitation for a 4 year old girl [name here]. It should be whimsical, look like its hand-drawn with little drawings on the sides of stuff like dinosaurs, flowers, hearts, cats. The background should be light and the foreground elements should be red, pink, orange and blue. Then I asked for some changes: > That's almost perfect! Retain this style and the elements, but adjust the text to read: > [refi…

that's lovely thank you. i am not very artistic so having stuff like this to crib is very helpful.

Re: 4o Image Generation

#549

Earlier quoted context omitted.

I think my point is that communication is the biggest contributor to brain development more than anything and communication is what powers our learning. Effective learners learn to communicate more with themselves and to communicate virtually with past authors through literature. That isn’t how LLMs work. Not sure why that would be considered objectionable. LLMs are great but we don’t have to pretend like they’re act…

There are so many points to consider here im not sure i can address them all. - Airplanes dont have wings like birds but can fly. and in some ways are superior to birds. (some ways not) - Human brains may be doing some analogue of sample augmentation which gives you some multiple more equivalent samples of data to train on per real input state of environment. This is done for ml too. - Whether that input data is text…

> Airplanes dont have wings like birds but can fly. and in some ways are superior to birds. (some ways not)

I think you're saying exactly what I'm saying. Human brains work differently from LLMs and the OP comment that started this thread is claiming that they work very similarly. In some ways they do but there's very clear differences and while clarifying examples in the training set can improve human understanding and performance, it's pretty clear we're doing something beyond that - just from a power efficiency perspective humans consume far less energy for significantly more performance and it's pretty likely we need less training data.

Re: 4o Image Generation

#550

Earlier quoted context omitted.

It's doable with diffusion, too.

I'm incredibly deep in the image / video / diffusion / comfy space. I've read the papers, written controlnets, modified architectures, pretrained, finetuned, etc. All that to say that I've been playing with 4o for the past day, and my opinions on the space have changed dramatically. 4o is a game changer. It's clearly imperfect, but its operating modalities are clearly superior to everything else we have seen. Have yo…

Yeah if we get an open model that one could apply a LoRA (or similarly cheap finetuning) to, then even problems like reproducing identity would (most likely) be solved, as they were for diffusion models. The coherence not just to the prompt but to any potential input image(s) is way beyond what I've seen in diffusion models.

I do think they run a "traditional" upscaler on the transformer output since it seems to sometimes have errors similar to upscalers (misinterpreted pixels), so probably the current decoded resolution is quite low and hopefully future models like GPT-5 will improve on this.

Post reply on HN