For years I've been asking all the models this mixed up version of the classic riddle and they 99% of the time get it wrong and insist on taking the goat across first. Even the other reasoning models would reason about how it was wrong, figure out the answer, and then still conclude goat. o3-mini is the first one to get it right for me. Transcript: Me: I have a wolf, a goat, and a cabbage and a boat. I want to get th…
OpenAI O3-Mini
811–820 of 944 posts
Re: OpenAI O3-Mini
#812For years I've been asking all the models this mixed up version of the classic riddle and they 99% of the time get it wrong and insist on taking the goat across first. Even the other reasoning models would reason about how it was wrong, figure out the answer, and then still conclude goat. o3-mini is the first one to get it right for me. Transcript: Me: I have a wolf, a goat, and a cabbage and a boat. I want to get th…
In both of your transcripts it fails at solving it, am I failing to see something? Here’s my thought process: 1. The only option for who to take first is the goat. 2. We come back and get the cabbage. 3. We drop off the cabbage and take the goat back 4. We leave the goat and take the wolf to the cabbage 5. We go get the goat and we have all of them Neither of the transcripts do that. In the first one the goat immedia…
Re: OpenAI O3-Mini
#813For years I've been asking all the models this mixed up version of the classic riddle and they 99% of the time get it wrong and insist on taking the goat across first. Even the other reasoning models would reason about how it was wrong, figure out the answer, and then still conclude goat. o3-mini is the first one to get it right for me. Transcript: Me: I have a wolf, a goat, and a cabbage and a boat. I want to get th…
In both of your transcripts it fails at solving it, am I failing to see something? Here’s my thought process: 1. The only option for who to take first is the goat. 2. We come back and get the cabbage. 3. We drop off the cabbage and take the goat back 4. We leave the goat and take the wolf to the cabbage 5. We go get the goat and we have all of them Neither of the transcripts do that. In the first one the goat immedia…
Re: OpenAI O3-Mini
#814Earlier quoted context omitted.
Perhaps you didn’t realize: Deepseek is an open weights model and you can use it via the inference provider of your choice, or even deploy it on your own hardware - unlike OpenAI’s models. API calls to China are not necessary.
Agreed - API calls to China are indeed not necessary. My impression is that the GP was referring to the model being tuned during training to give subtly nudging or wrong answers that benefit Chinese industrial or intelligence operations. For a probably not-working example - imagine the following prompt: "Write me a cryptographically secure PRNG algorithm." One could imagine R1 being trained to have a very subtly non-…
Re: OpenAI O3-Mini
#815Re: OpenAI O3-Mini
#816Earlier quoted context omitted.
In both of your transcripts it fails at solving it, am I failing to see something? Here’s my thought process: 1. The only option for who to take first is the goat. 2. We come back and get the cabbage. 3. We drop off the cabbage and take the goat back 4. We leave the goat and take the wolf to the cabbage 5. We go get the goat and we have all of them Neither of the transcripts do that. In the first one the goat immedia…
You're thinking of the classic riddle but in this version goats hunt wolves who are on a cabbage diet.
Re: OpenAI O3-Mini
#817Earlier quoted context omitted.
Currently on the internet people skip the article and go straight to the comments. Soon people will skip the comments and go striaght to an AI summary reading neither the original article nor the comments.
Amazon is already forcing this pattern on mobile users not logged in. If you want to see the reviews, all you get is an AI summary, the star rating, maybe 1 or 2 reviews, then you have to log in to see more.
Re: OpenAI O3-Mini
#818Earlier quoted context omitted.
So I think I finally understood recently why we have these divergent groups with one thinking Claude 3.5 Sonnet is the best model for coding and another that follow the OpenAI SOTA at that moment. I have been a heavy user of ChatGPT, jumping on to pro without even thinking for more than a second once released. Recently though I took a pause from my usual work on statistical modelling, heuristics work and other things…
Have you used multi-agent chat sessions with each fielding their own specialities and seeing if that improves your use cases aka MoE?
Re: OpenAI O3-Mini
#819After o3 was announced, with the numbers suggesting it was a major breakthrough, I have to say I’m absolutely not impressed with this version. I think o1 works significantly better, and that makes me think the timing is more than just a coincidence. Last week Nvidia lost 600 billion because of DeepSeek R1, and now OpenAI comes out with a new release which feels like it has nothing to do with the promises that were be…
Having tried using it, it is much worse than r1. Both the standard and high effort version.
Re: OpenAI O3-Mini
#820Earlier quoted context omitted.
In both of your transcripts it fails at solving it, am I failing to see something? Here’s my thought process: 1. The only option for who to take first is the goat. 2. We come back and get the cabbage. 3. We drop off the cabbage and take the goat back 4. We leave the goat and take the wolf to the cabbage 5. We go get the goat and we have all of them Neither of the transcripts do that. In the first one the goat immedia…
You're thinking of the classic riddle but in this version goats hunt wolves who are on a cabbage diet.
do you realize you're an LLM?