Earlier quoted context omitted.
Thank you, this is a perfect argument why LLMs are not AI but just statistical models. The original is so overrepresented in the training data that even though they notice this riddle is different, they regress to the statistically more likely solution over the course of generating the response. For example, I tried the first one with Claude and in its 4th step, it said: > This is safe because the wolf won't eat the…
The problem with claims like these that models are not doing “actual reasoning” is that they are often hot takes and not thought through very well. For example, since reasoning doesn’t yet have any consensus definition that can be applied as a yes/no test - you have to explain what you specifically mean by it, or else the claim is hollow. Clarify your definition, give a concrete example under that definition of somet…
OpenAI O3-Mini
881–890 of 944 posts
Re: OpenAI O3-Mini
#882Earlier quoted context omitted.
Have you ever had to give a demo that no meeting attendee actually cared about, just because management demanded it? Standup meetings with 20 people where maybe 2 people cared about? The future might involve AI updates, summarized by AI into weekly reports, summarized again into monthly reports, then into quarterly departmental reports that nobody actually reads.
Or, the future might be everyone reads summaries, because there are only AI Managers and no human managers, where humans hold occasional meetings and have conversations which are listed to by AI, and we take our lead from the AI summaries. The AI maintains business focus through it monitoring the business performance, updating each summary as needed to maintain performance. It's a worthwhile experiment for a business…
Re: OpenAI O3-Mini
#883Earlier quoted context omitted.
Having tried using it, it is much worse than r1. Both the standard and high effort version.
If it’s actually available, it can’t be that much worse than R1 which currently only completes a response about 50% of the time for me.
Re: OpenAI O3-Mini
#884Earlier quoted context omitted.
Thank you, this is a perfect argument why LLMs are not AI but just statistical models. The original is so overrepresented in the training data that even though they notice this riddle is different, they regress to the statistically more likely solution over the course of generating the response. For example, I tried the first one with Claude and in its 4th step, it said: > This is safe because the wolf won't eat the…
The problem with claims like these that models are not doing “actual reasoning” is that they are often hot takes and not thought through very well. For example, since reasoning doesn’t yet have any consensus definition that can be applied as a yes/no test - you have to explain what you specifically mean by it, or else the claim is hollow. Clarify your definition, give a concrete example under that definition of somet…
For example, if I throw a bunch of sticks in the air and look at their patterns to divine the future- can I call that "mathematics" just because nobody has a "consensus definition of mathematics that can be applied as a yes/no test"? Can I just call anything I like mathematics and nobody can tell me it's wrong because ... no definition?
We, as a civilisation, have studied both formal and informal reasoning since at least a couple thousand years go, starting with Aristotle and his syllogisms (a formalisation of rigorous arguments) and continuing through the years with such figures as Leibniz, Boole, Bayes, Frege, Pierce, Quine, Russel, Godel, Turing, etc etc. There are entire research disciplines that are dedicated to the study of reasoning: philosophy, computer science, and, of course, all of mathematics itself. In AI research reasoning is a major topic studied by fields like automated theorem proving, planning and scheduling, program verification and model checking, etc, everything one finds in Russel & Norvig really. It is only in machine learning circles that reasoning seems to be such a big mystery that nobody can agree what it is; and in discussions on the internet about whether LLMs reason or not.
And it should be clear that never in the history of human civilisation did "reasoning" mean "predict the most likely answer according to some training corpus".
Re: OpenAI O3-Mini
#885Earlier quoted context omitted.
What works nice also is the text to speech. I find it easier and faster to give more context by talking rather than typing, and the extra content helps the AI to do its job. And even though the speech recognition fails a lot on some of the technical terms or weirdly named packages, software, etc, it still does a good job overall (if I don’t feel like correcting the wrong stuff). It’s great and has become somewhat of…
you mean speech to text right?
Re: OpenAI O3-Mini
#886Earlier quoted context omitted.
Currently on the internet people skip the article and go straight to the comments. Soon people will skip the comments and go striaght to an AI summary reading neither the original article nor the comments.
But then there will be no comments to summarize.
Re: OpenAI O3-Mini
#887Earlier quoted context omitted.
Having tried using it, it is much worse than r1. Both the standard and high effort version.
Yea, o3-mini was a massive step down from Sonnet for coding tasks. R1 is my cost effective programmer. Sonnet is my hard problem model still.
Since I have access to the thinking tokens I can see where it's going wrong and do prompt surgery. But left to it's own devices it gets thing _stupendously_ wrong about 20% of the time with a huge context blowout. So much so that seeing that happen now tells me I've fundamentally asked the wrong question.
Sonnet doesn't suffer from that and solves the task, but doesn't give you much if any, help in how to recover from doing the wrong task.
I'd say that for work work Sonnet 3.5 is still the best, for exploratory work with a human in the loop r1 is better.
Or as someone posted here a few days ago: R1 as the architect, Sonnet3.5 as the worker and critic.
Re: OpenAI O3-Mini
#888Earlier quoted context omitted.
Currently on the internet people skip the article and go straight to the comments. Soon people will skip the comments and go striaght to an AI summary reading neither the original article nor the comments.
Sounds about right, as we are post-dead internet in public places. There was a thread about the US tariffs on Canada I was reading on a stock investment subreddit. The whole page was full of people complaining about Elon Musk, Donald Trump, "Buy Canadian" comments, moralizing about Alberta's conservative government and other unrelated noise. None of this was related to the topic; stocks and funds that seemed well-pla…
Stealing this.
Re: OpenAI O3-Mini
#889Earlier quoted context omitted.
Deepseek V3 is equivalent to 4o. Deepseek R1 is equivalent to o1 (if not better) I think someone should just build an AI model comparing website at this point. Include all benchmarks and pricing
I had resubscribed to use o1 2 weeks ago and haven't even logged in this week because of R1. One thing I notice that is huge is being able to see the chain of thought lets me see when my prompt was lacking and the model is a bit confused on what I want. If I was anymore impressed with R1 I would probably start getting accused of being a CCP shill or wumao lol. With that said, I think it is very hard to compare models…
It’s an amazing model but was so much faster before the hype
The servers being constantly down is the only reason I haven’t cancelled my ChatGPT subscription
Re: OpenAI O3-Mini
#890For years I've been asking all the models this mixed up version of the classic riddle and they 99% of the time get it wrong and insist on taking the goat across first. Even the other reasoning models would reason about how it was wrong, figure out the answer, and then still conclude goat. o3-mini is the first one to get it right for me. Transcript: Me: I have a wolf, a goat, and a cabbage and a boat. I want to get th…
The whole point is you are distilling past knowledge, if you are making up on the spot nonsense to purposely make all past knowledge useless... get out of my house