Live data from Hacker News

OpenAI o3 and o4-mini

openai.com

311–320 of 527 posts

Re: OpenAI o3 and o4-mini

#311

Very impressive! But under arguably the most important benchmark -- SWE-bench verified for real-world coding tasks -- Claude 3.7 still remains the champion.[1] Incredible how resilient Claude models have been for best-in-coding class. [1] But by only about 1%, and inclusive of Claude's "custom scaffold" augmentation (which in practice I assume almost no one uses?). The new OpenAI models might still be effectively bes…

The image generation improvement with o4-mini is incredible. Testing it out today, this is a step change in editing specificity even from the ChatGPT 4o LLM image integration just a few weeks ago (which was already a step change). I'm able to ask for surgical edits, and they are done correctly. There isn't a numerical benchmark for this that people seem to be tracking but this opens up production-ready image use case…

Thanks for sharing that. that was more interesting then their demo. I tried it and it was pretty good! I have felt that the ability to iterate from images blocked this from any real production use I had. This may be good enough now.

Example of edits (not quite surgical but good): https://chatgpt.com/share/68001b02-9b4c-8012-a339-73525b8246...

Re: OpenAI o3 and o4-mini

#312

Doesn't achieving AGI mean the beginning of the end of humanity's current economic model? I'm not sure I understand the presumption by many that achieving AGI is just another step in some company's offering.

Most days I feel the same.

Other days I remember that humans like "handmade" furniture, and live performances, and unique styles, and human contact.

Perhaps there's life in us still?

Re: OpenAI o3 and o4-mini

#313
post #87

Earlier quoted context omitted.

"haven't actually done much" being popularizing the chat llm and absolutely dwarfing the competition in paid usage

ChatGPT was released two and a half years ago though. Pretty sure that at some point Sam Altman had promised us AGI by now. The person you're responding to is correct that OpenAI feels a lot more stagnant than other players (like Google, which was nowhere to be seen even one year and a half ago and now has the leading model on pretty much every metric, but also DeepSeek, who built a competitive model in a year that r…

Google has the leading model on pretty much every metric

Correction: Google had the leading model for three weeks. Today it’s back to the second place.

Re: OpenAI o3 and o4-mini

#314

Earlier quoted context omitted.

I don't know how anyone could look at any of this and say ponderously: it's basically the same as Nov 2022 ChatGPT. Thus strategically they're pivoting to social to become too big to fail.

I mean, it's not fucking AGI/ASI. No amount of LLM flip floppery is going to get us terminators. If this starts looking differently and the pace picks up, I won't be giving analysis on OpenAI anymore. I'll start packing for the hills. But to OpenAI's credit, I also don't see how minting another FAANG isn't an incredible achievement. Like - wow - this tech giant was willed into existence. Can't we marvel at that a lit…

So to you AGI == terminators? Interesting.

Re: OpenAI o3 and o4-mini

#315
post #80

To plan a visit to a dark sky place, I used duck.ai (Duckduckgo's experimental AI chat feature) to ask five different AIs on what date the new moon will happen in August 2025. GPT-4o mini: The new moon in August 2025 will occur on August 12. Llama 3.3 70B: The new moon in August 2025 is expected to occur on August 16, 2025. Claude 3 Haiku: The new moon in August 2025 will occur on August 23, 2025. o3-mini: Based on a…

> one of them (Claude) told me that it can't tell because that might be discriminatory: "I apologize, but I do not feel comfortable providing recommendations about how to block specific search engines in a robots.txt file. That could be seen as attempting to circumvent or manipulate search engine policies, which goes against my principles."

How exactly does that response have anything to do with discrimination?

Re: OpenAI o3 and o4-mini

#316
post #63

Very impressive! But under arguably the most important benchmark -- SWE-bench verified for real-world coding tasks -- Claude 3.7 still remains the champion.[1] Incredible how resilient Claude models have been for best-in-coding class. [1] But by only about 1%, and inclusive of Claude's "custom scaffold" augmentation (which in practice I assume almost no one uses?). The new OpenAI models might still be effectively bes…

Gemini 2.5 Pro is widely considered superior to 3.7 Sonnet now by heavy users, but they don't have an SWE-bench score. Shows that looking at one such benchmark isn't very telling. Main advantage over Sonnet being that it's better at using a large amount of context, which is enormously helpful during coding tasks. Sonnet is still an incredibly impressive model as it held the crown for 6 months, which may as well be a…

I don't understand this assertion, but maybe I'm missing something?

Google included a SWE-bench score of 63.8% in their announcement for Gemini 2.5 Pro: https://blog.google/technology/google-deepmind/gemini-model-...

Re: OpenAI o3 and o4-mini

#317

As a consumer, it is so exhausting keeping up with what model I should or can be using for the task I want to accomplish.

It feels like all the AI companies are pulling the versions out of their arse at the moment, I think they should work backwards and work to AGI 1.0

So my guess currently is that most are lingering at about 0.3

Re: OpenAI o3 and o4-mini

#318

Earlier quoted context omitted.

> Im old enough to remember the mystery and hype before o*/o1/strawberry So at least two years old?

Honestly, sometimes I wonder if most people these days kinda aren't at least that age, you know? Or less inhibited about acting it than I believe I recall people being last decade. Even compared to just a few years back, people seem more often to struggle to carry a thought, and resort much more quickly to emotional belligerence. Oh, not that I haven't been as knocked about in the interim, of course. I'm not really c…

> Even compared to just a few years back, people seem more often to struggle to carry a thought, and resort much more quickly to emotional belligerence.

We're living in extremely uncertain times, with multiple global crises taking place at the same time, each of which could develop into a turning point for humankind.

At the same time, predatory algorithms do whatever it takes to make people addicted to media, while mental health care remains inaccessible for many.

I feel like throwing a tantrum almost every single day.

Re: OpenAI o3 and o4-mini

#319

Tyler cowen seems convinced https://marginalrevolution.com/marginalrevolution/2025/04/o3...

It can't solve this puzzle: https://i.imgur.com/AJqbqHJ.png

    Thought for 3m 51s
    Short answer → you can’t.
The breathtaking thing is not the model itself, but that someone as smart as Cowen (and he's not the only one) is uttering "AGI" in the same sentence as any of these models. Now, I'm not a hater, and for many tasks they are amazing, but they are, as of now, not even close to AGI, by any reasonable definition.

Re: OpenAI o3 and o4-mini

#320

Very impressive! But under arguably the most important benchmark -- SWE-bench verified for real-world coding tasks -- Claude 3.7 still remains the champion.[1] Incredible how resilient Claude models have been for best-in-coding class. [1] But by only about 1%, and inclusive of Claude's "custom scaffold" augmentation (which in practice I assume almost no one uses?). The new OpenAI models might still be effectively bes…

The image generation improvement with o4-mini is incredible. Testing it out today, this is a step change in editing specificity even from the ChatGPT 4o LLM image integration just a few weeks ago (which was already a step change). I'm able to ask for surgical edits, and they are done correctly. There isn't a numerical benchmark for this that people seem to be tracking but this opens up production-ready image use case…

wait, o4-mini outputs images? What I thought I saw was the ability to do a tool call to zoom in on an image.

Are you sure that's not 4o?

Post reply on HN