Watching o3 model sweat over a Paul Morphy mate-in-2
1–10 of 73 posts
Re: Watching o3 model sweat over a Paul Morphy mate-in-2
#2O3 is massively underwhelming and is obviously tuned to be sycophantic.
Claude reigns supreme.
Re: Watching o3 model sweat over a Paul Morphy mate-in-2
#3I've commited the 03 (zero-three) and not o3 (o-three) typo too, but can we rename it on the title please
Re: Watching o3 model sweat over a Paul Morphy mate-in-2
#4So, are we talking about OpenAI o3 model, right?
Re: Watching o3 model sweat over a Paul Morphy mate-in-2
#5On a similar note, I just updated LLM Chess Puzzles repo [1] yesterday.
The fact that gpt-4.5 gets 85% correctly solved is unexpected and somewhat scary (if model was not trained on this).
Re: Watching o3 model sweat over a Paul Morphy mate-in-2
#6O3 is massively underwhelming and is obviously tuned to be sycophantic. Claude reigns supreme.
This somehow reminds me of Agent-3 from [0].
Re: Watching o3 model sweat over a Paul Morphy mate-in-2
#7So, are we talking about OpenAI o3 model, right?
yes
Re: Watching o3 model sweat over a Paul Morphy mate-in-2
#8On a similar note, I just updated LLM Chess Puzzles repo [1] yesterday. The fact that gpt-4.5 gets 85% correctly solved is unexpected and somewhat scary (if model was not trained on this). [1] https://github.com/kagisearch/llm-chess-puzzles
Oh cool, I wonder how good 03 will be.
While using 03, I noticed something funny: sometimes I gave it a screenshot without any position data. It ended up using Python and spent 10 minutes just trying to figure out where the figures were exactly.
Re: Watching o3 model sweat over a Paul Morphy mate-in-2
#9So, are we talking about OpenAI o3 model, right?
>"When I gave OpenAI’s 03 model a tough chess puzzle..."
Opening sentence