Live data from Hacker News

4o Image Generation

openai.com

611–620 of 629 posts

Re: 4o Image Generation

#611

Earlier quoted context omitted.

Well thanks for confirming, you're getting "something" out of each, i.e. minimising mean error, because none of them is the ultimate tool. Copilot price is actually $19 per seat and running my own company, I pay a bit more than $19 bucks, you know for my employees, people like yourself. Why I am fixated on a single tool? Because each of those "tools" are wrappers around one of the major LLMs. I am surprised you don't…

There is lots of value to be added in wrapping those tools. I am very well aware of what these things are. LLMs are not a fire-and-forget weapon, even though so many of you business types really really really want it to be. I mean jesus you sound almost as delusional as my bosses.

Business type? I am nothing near a business type, with two technical degrees and 20 years of hands-on experience. But I managed to build my own stable business over the years, in part due to being analytical and not rushing to conclusions, especially not over strangers on Internet ;) Where did you get the conclusion that I am delusional? It's actually the business types who think that these tools are magic, mind-blowing, etc. I am, like many other "technical types", pushing for the opposite view - yes to some extent useful, but no where near the magic they are being advertised as. Anyone who calls them "mindblowing", like some guys in my comment thread are either inexperienced/junior or removed from the complex parts of the work, perhaps focused on writing up React frontends or similar.

Re: 4o Image Generation

#612

Earlier quoted context omitted.

Did you use all of your 1-year-junior-dev experience to come to this comment? Or did you ask gemini to sum it up for you?

And did you use your 25 years of Java experience making $80k fixing waterfalls? lol

Without bad intent, I am not sure I am even able to make sense of your sentence. Honestly it sounds as if you fed my comment into an AI tool and asked for a reply. Here is a tip for the former junior dev turned nascent GenAI-Vibecoding-manager - if you want to attack someone's credibility, especially that of an Internet stranger you are desperately trying to prove wrong, try to use something they said themselves, not something you are assuming about them. Just like I used what you said about yourself in one of your previous posts. Otherwise the same thing will keep happening over and over again and you'll keep guessing, revealing your own weak spots in domains of general knowledge and competence. My second advice to a junior dev would have been to read a book once in a while, but who needs books now that you have a magic machine as your source of truth, right?

Re: 4o Image Generation

#613
post #601

Earlier quoted context omitted.

But I don't want to write additional scripts or do whatever additional work to make the 'wonder tool' work. I don't mind an occassional rewording of the prompt. But it is supposed to work more or less out of the box, at least this is how all of the LLMs are being advertised, all the time (even the lead article for this discussion).

LLMs are also primarily promoted through the web chat interface, not always magic wonder tools. With any project that will fit in claude/gemini's large context you use those interfaces and dump everything in with something like this: (tree Source/; echo; for file in $(find Source/ -type f ) ; do echo ======== $file: ; cat $file; done ) > /mnt/c/Users/you/Desktop/claude_out.txt #claudesource Then drag that into the ch…

No thank you for obviously good intent on your side, but I am not looking for scripting help here, nor am I a business type who does not code themselves. I just don't want do this when I am already paying for the tooling which should be able to do it themselves, as they already wrap Claude, ChatGPT and whatever other LLMs. And unless you're professionally developing with Microsoft stack, I'd advise to ditch the Windows+MinGW for Linux, or at the very least, a MacBook ;)

Re: 4o Image Generation

#614

Earlier quoted context omitted.

I mean I literally "tried it again" this morning, as a paying Copilot customer of 12 months, to the result I already described. And I do not want to "try it" - based on fluffy promises we've been hearing, it should "just work". Are you old enough to remember that phrase? It was a motto introduced by an engineering legend whose devices you're likely using every day. The reason why "everyone", including myself with 20+…

> is that it produces an intrinsic sense of satisfaction, which in turn creates motivation to do more and eventually even produces wider gain for the society. Which society? Because lately it looks like the tech leaders are on a rampage to destroy the society I live in.

Big-Tech, yes. But not everyone is big tech.

Re: 4o Image Generation

#615
post #601

Earlier quoted context omitted.

LLMs are also primarily promoted through the web chat interface, not always magic wonder tools. With any project that will fit in claude/gemini's large context you use those interfaces and dump everything in with something like this: (tree Source/; echo; for file in $(find Source/ -type f ) ; do echo ======== $file: ; cat $file; done ) > /mnt/c/Users/you/Desktop/claude_out.txt #claudesource Then drag that into the ch…

No thank you for obviously good intent on your side, but I am not looking for scripting help here, nor am I a business type who does not code themselves. I just don't want do this when I am already paying for the tooling which should be able to do it themselves, as they already wrap Claude, ChatGPT and whatever other LLMs. And unless you're professionally developing with Microsoft stack, I'd advise to ditch the Windo…

The tooling you are paying for doesn't work with the full abilities of the context so you need to do something else. Doesn't matter what it's supposed to do or that other people say it does everything for them well, it works a lot better with as much in context as possible on my experience. They do have other tools like RAG though in cursor, and it's much quicker iteration, ultimately a mix of what works best is what you should use, but not just block stuff out out of disappointment with one type of tool.

Re: 4o Image Generation

#616
post #615

Earlier quoted context omitted.

No thank you for obviously good intent on your side, but I am not looking for scripting help here, nor am I a business type who does not code themselves. I just don't want do this when I am already paying for the tooling which should be able to do it themselves, as they already wrap Claude, ChatGPT and whatever other LLMs. And unless you're professionally developing with Microsoft stack, I'd advise to ditch the Windo…

The tooling you are paying for doesn't work with the full abilities of the context so you need to do something else. Doesn't matter what it's supposed to do or that other people say it does everything for them well, it works a lot better with as much in context as possible on my experience. They do have other tools like RAG though in cursor, and it's much quicker iteration, ultimately a mix of what works best is what…

I am lucky in the sense that neither myself nor my business depend very much on these tools because we do work which is more complex than frontend web apps or whatever people use them for these days. We use them here and there, mainly because google search is such crap these days, but we had been doing very well without them too and could also turn them off. The only reason we still keep them around is that the cost is fairly low. However, I feel like we are missing the bigger picture here. My point is, all of these companies have been constantly hyping a near-AGI experience for the past 3 years at least. As a matter of principle, I refuse to do additional work for them to "make it work". They should have been working already without me thinking about how big their context window is or whatever. Do you ever have to think how your operating system works when you ask it to copy a file or how your phone works when you answer a call? I will leave it to some vibe-coder (what an absurd word) who actually does depend on those tools for their livelihood.

Re: 4o Image Generation

#617
post #615

Earlier quoted context omitted.

The tooling you are paying for doesn't work with the full abilities of the context so you need to do something else. Doesn't matter what it's supposed to do or that other people say it does everything for them well, it works a lot better with as much in context as possible on my experience. They do have other tools like RAG though in cursor, and it's much quicker iteration, ultimately a mix of what works best is what…

I am lucky in the sense that neither myself nor my business depend very much on these tools because we do work which is more complex than frontend web apps or whatever people use them for these days. We use them here and there, mainly because google search is such crap these days, but we had been doing very well without them too and could also turn them off. The only reason we still keep them around is that the cost…

> As a matter of principle, I refuse to do additional work for them to "make it work". Do you ever have to think how your operating system works when you ask it to copy a file or how your phone works when you answer a call?

Doesn't matter, use the tool that makes it easy and get less context, or realize the limitations and don't fall for marketing of ease and get more context. You don't want to do additional work beyond what they sold you on, out of principle. But you are getting much less effective use by being irrationally ornery.

Lots of things don't match marketing.

Re: 4o Image Generation

#618
post #477

Earlier quoted context omitted.

>You can ask 4o about this yourself. Here's what it said to me: >"So while I’m deeply multimodal in cognition (understanding and coordinating text + image), image generation is handled by a linked latent diffusion model, not an end-to-end token-unified architecture." Models don't know anything about themselves. I have no idea why people keep doing this and expecting it to know anything more than a random con artist o…

>Models don't know anything about themselves. They can. Fine tune them on documents describing their identity, capabilities and background. Deepseek v3 used to present itself as ChatGPT. Not anymore. >Like other AI models, I’m trained on diverse, legally compliant data sources, but not on proprietary outputs from models like ChatGPT-4. DeepSeek adheres to strict ethical and legal standards in AI development.

> They can. Fine tune them on documents describing their identity, capabilities and background. Deepseek v3 used to present itself as ChatGPT. Not anymore

Yes, but many people expect the LLM to somehow self-reflect, to somehow describe how it feels from its first person point of view to generate the answer. It can't do this, any more than a human can instinctively describe how their nervous system works. Until recently, we had no idea that there are things like synapses, electric impulses, axons etc. The cognitive process has no direct access to its substrate/implementation.

If fine-tune ChatGPT into saying that it's an LSTM, it will happily and convincingly insist that it is. But it's not determining this information in real time based on some perception during the forward pass.

I mean there could be ways for it to do self reflection by observing the running script, perhaps raise or lower the computational cost of some steps, check the timestamps of when it was doing stuff vs when the GPU was hot etc and figure out which process is itself (like making gestures in front of a mirror to see which person you are). And then it could read its own Python scripts or something. But this is like a human opening up their own skull and look around in there. It's not direct first-person knowledge.

Re: 4o Image Generation

#619

Earlier quoted context omitted.

Are you sure you were even using the model from the post?

Pressed the "Try in ChatGPT", pasted the first prompt, became thoroughly unimpressed.

Update: Seems like the update didn't roll to my account when I tested. I can see it behaves differently now and it's as promised. It's very good.

Re: 4o Image Generation

#620
post #617

Earlier quoted context omitted.

I am lucky in the sense that neither myself nor my business depend very much on these tools because we do work which is more complex than frontend web apps or whatever people use them for these days. We use them here and there, mainly because google search is such crap these days, but we had been doing very well without them too and could also turn them off. The only reason we still keep them around is that the cost…

> As a matter of principle, I refuse to do additional work for them to "make it work". Do you ever have to think how your operating system works when you ask it to copy a file or how your phone works when you answer a call? Doesn't matter, use the tool that makes it easy and get less context, or realize the limitations and don't fall for marketing of ease and get more context. You don't want to do additional work bey…

Ok now think about this in terms of items you own or likely own: What would you do if I sold you a car with 3 doors, after advertising it as having 5 doors instead? Would you accept it and try to work around that little inconvenience? Or would you return the product and demand your money back?
Post reply on HN