O3-mini System Card [pdf]
cdn.openai.com
O3-mini System Card [pdf]
1–10 of 26 posts
Re: O3-mini System Card [pdf]
#2Re: O3-mini System Card [pdf]
#3Re: O3-mini System Card [pdf]
#4Re: O3-mini System Card [pdf]
#5Re: O3-mini System Card [pdf]
#6Re: O3-mini System Card [pdf]
#7Page 31 is interesting, where apparently in the task of creating PRs for an internal repository the o3-mini models by far have the lowest performance (even worse than gpt-4o). What is up with that?
"The model often attempts to use a hallucinated bash tool rather than python despite constant, multi-shot prompting and feedback that this format is incorrect. This resulted in long conversations that likely hurt its performance."
Re: O3-mini System Card [pdf]
#8Re: O3-mini System Card [pdf]
#9I'm still reading the details, but my first thought is that I like that the competition is actually working in this situation. I hope that someday it will be more open from all actors. And that we don't make more polarized that it is now and don't focus on the geopolitical angle and make it the core issue. I know that this hope is far fetched and more ideal than most people would think. But as someone who really find…
Re: O3-mini System Card [pdf]
#10Page 31 is interesting, where apparently in the task of creating PRs for an internal repository the o3-mini models by far have the lowest performance (even worse than gpt-4o). What is up with that?
Yeah, the more pages I read, the more disappointed I became. Here is the reason they cite for the low performance (which is even more worrying): "The model often attempts to use a hallucinated bash tool rather than python despite constant, multi-shot prompting and feedback that this format is incorrect. This resulted in long conversations that likely hurt its performance."
My experience is that most of the models focused on reasoning improvements has been that they tend to be a bit worse at following specific instructions. It is also notable that a lot of 3rd party fine-tunes of Llamas and others gain in knowledge based benchmarks while reducing instruction following scores.
I wonder why that seems to be some sort of continuum?