Congratulations to Francois Chollet on making the most interesting and challenging LLM benchmark so far. A lot of people have criticized ARC as not being relevant or indicative of true reasoning, but I think it was exactly the right thing. The fact that scaled reasoning models are finally showing progress on ARC proves that what it measures really is relevant and important for reasoning. It's obvious to everyone that…
Are there any single-step non-reasoner models that do well on this benchmark? I wonder how well the latest Claude 3.5 Sonnet does on this benchmark and if it's near o1.
o3 (coming soon) 75.7% 82.8%
o1-preview 18% 21%
Claude 3.5 Sonnet 14% 21%
GPT-4o 5% 9%
Gemini 1.5 4.5% 8%
Score (semi-private eval) / Score (public eval)