Live data from Hacker News

OpenAI O3-Mini

openai.com

881–890 of 944 posts

Re: OpenAI O3-Mini

#881

Earlier quoted context omitted.

Thank you, this is a perfect argument why LLMs are not AI but just statistical models. The original is so overrepresented in the training data that even though they notice this riddle is different, they regress to the statistically more likely solution over the course of generating the response. For example, I tried the first one with Claude and in its 4th step, it said: > This is safe because the wolf won't eat the…

The problem with claims like these that models are not doing “actual reasoning” is that they are often hot takes and not thought through very well. For example, since reasoning doesn’t yet have any consensus definition that can be applied as a yes/no test - you have to explain what you specifically mean by it, or else the claim is hollow. Clarify your definition, give a concrete example under that definition of somet…

Is it possible the models do something entirely different? I'm not sure why everyone needs to compare them to human intelligence. It's very obvious llms work nothing like our brains why would the intelligence they exhibit be like ours?

Re: OpenAI O3-Mini

#882
post #830

Earlier quoted context omitted.

Have you ever had to give a demo that no meeting attendee actually cared about, just because management demanded it? Standup meetings with 20 people where maybe 2 people cared about? The future might involve AI updates, summarized by AI into weekly reports, summarized again into monthly reports, then into quarterly departmental reports that nobody actually reads.

Or, the future might be everyone reads summaries, because there are only AI Managers and no human managers, where humans hold occasional meetings and have conversations which are listed to by AI, and we take our lead from the AI summaries. The AI maintains business focus through it monitoring the business performance, updating each summary as needed to maintain performance. It's a worthwhile experiment for a business…

Ah, good ol' Manna. https://marshallbrain.com/manna1

Re: OpenAI O3-Mini

#883

Earlier quoted context omitted.

Having tried using it, it is much worse than r1. Both the standard and high effort version.

If it’s actually available, it can’t be that much worse than R1 which currently only completes a response about 50% of the time for me.

There are multiple providers for it since it's open source.

Re: OpenAI O3-Mini

#884

Earlier quoted context omitted.

Thank you, this is a perfect argument why LLMs are not AI but just statistical models. The original is so overrepresented in the training data that even though they notice this riddle is different, they regress to the statistically more likely solution over the course of generating the response. For example, I tried the first one with Claude and in its 4th step, it said: > This is safe because the wolf won't eat the…

The problem with claims like these that models are not doing “actual reasoning” is that they are often hot takes and not thought through very well. For example, since reasoning doesn’t yet have any consensus definition that can be applied as a yes/no test - you have to explain what you specifically mean by it, or else the claim is hollow. Clarify your definition, give a concrete example under that definition of somet…

Explain this to me please: we don't have any consensus definition of _mathematics_ that can be applied as a yes/no test. Does that mean we don't know how to do mathematics, or that we don't know whether something, is, or, more importantly, isn't mathematics?

For example, if I throw a bunch of sticks in the air and look at their patterns to divine the future- can I call that "mathematics" just because nobody has a "consensus definition of mathematics that can be applied as a yes/no test"? Can I just call anything I like mathematics and nobody can tell me it's wrong because ... no definition?

We, as a civilisation, have studied both formal and informal reasoning since at least a couple thousand years go, starting with Aristotle and his syllogisms (a formalisation of rigorous arguments) and continuing through the years with such figures as Leibniz, Boole, Bayes, Frege, Pierce, Quine, Russel, Godel, Turing, etc etc. There are entire research disciplines that are dedicated to the study of reasoning: philosophy, computer science, and, of course, all of mathematics itself. In AI research reasoning is a major topic studied by fields like automated theorem proving, planning and scheduling, program verification and model checking, etc, everything one finds in Russel & Norvig really. It is only in machine learning circles that reasoning seems to be such a big mystery that nobody can agree what it is; and in discussions on the internet about whether LLMs reason or not.

And it should be clear that never in the history of human civilisation did "reasoning" mean "predict the most likely answer according to some training corpus".

Re: OpenAI O3-Mini

#885
post #605
post #538

Earlier quoted context omitted.

What works nice also is the text to speech. I find it easier and faster to give more context by talking rather than typing, and the extra content helps the AI to do its job. And even though the speech recognition fails a lot on some of the technical terms or weirdly named packages, software, etc, it still does a good job overall (if I don’t feel like correcting the wrong stuff). It’s great and has become somewhat of…

you mean speech to text right?

Sorry! Yes, speech to text.

Re: OpenAI O3-Mini

#886
post #669

Earlier quoted context omitted.

Currently on the internet people skip the article and go straight to the comments. Soon people will skip the comments and go striaght to an AI summary reading neither the original article nor the comments.

But then there will be no comments to summarize.

I think people generally like writing comments. Reading articles, in their entirety, less so.

Re: OpenAI O3-Mini

#887

Earlier quoted context omitted.

Having tried using it, it is much worse than r1. Both the standard and high effort version.

Yea, o3-mini was a massive step down from Sonnet for coding tasks. R1 is my cost effective programmer. Sonnet is my hard problem model still.

R1 is interesting.

Since I have access to the thinking tokens I can see where it's going wrong and do prompt surgery. But left to it's own devices it gets thing _stupendously_ wrong about 20% of the time with a huge context blowout. So much so that seeing that happen now tells me I've fundamentally asked the wrong question.

Sonnet doesn't suffer from that and solves the task, but doesn't give you much if any, help in how to recover from doing the wrong task.

I'd say that for work work Sonnet 3.5 is still the best, for exploratory work with a human in the loop r1 is better.

Or as someone posted here a few days ago: R1 as the architect, Sonnet3.5 as the worker and critic.

Re: OpenAI O3-Mini

#888
post #848
post #669

Earlier quoted context omitted.

Currently on the internet people skip the article and go straight to the comments. Soon people will skip the comments and go striaght to an AI summary reading neither the original article nor the comments.

Sounds about right, as we are post-dead internet in public places. There was a thread about the US tariffs on Canada I was reading on a stock investment subreddit. The whole page was full of people complaining about Elon Musk, Donald Trump, "Buy Canadian" comments, moralizing about Alberta's conservative government and other unrelated noise. None of this was related to the topic; stocks and funds that seemed well-pla…

> zoomer internet church

Stealing this.

Re: OpenAI O3-Mini

#889
post #41

Earlier quoted context omitted.

Deepseek V3 is equivalent to 4o. Deepseek R1 is equivalent to o1 (if not better) I think someone should just build an AI model comparing website at this point. Include all benchmarks and pricing

I had resubscribed to use o1 2 weeks ago and haven't even logged in this week because of R1. One thing I notice that is huge is being able to see the chain of thought lets me see when my prompt was lacking and the model is a bit confused on what I want. If I was anymore impressed with R1 I would probably start getting accused of being a CCP shill or wumao lol. With that said, I think it is very hard to compare models…

R1 servers seem to be down or busy a lot lately.

It’s an amazing model but was so much faster before the hype

The servers being constantly down is the only reason I haven’t cancelled my ChatGPT subscription

Re: OpenAI O3-Mini

#890

For years I've been asking all the models this mixed up version of the classic riddle and they 99% of the time get it wrong and insist on taking the goat across first. Even the other reasoning models would reason about how it was wrong, figure out the answer, and then still conclude goat. o3-mini is the first one to get it right for me. Transcript: Me: I have a wolf, a goat, and a cabbage and a boat. I want to get th…

Whyyy would anyone want a system that can even answer made up nonsense based on arbitrary word redefining and wrongness?

The whole point is you are distilling past knowledge, if you are making up on the spot nonsense to purposely make all past knowledge useless... get out of my house

Post reply on HN