Proof or Bluff? Evaluating LLMs on 2025 USA Math Olympiad
1–2 of 2 posts
Re: Proof or Bluff? Evaluating LLMs on 2025 USA Math Olympiad
#2> Our results reveal that all tested models struggled significantly, achieving less than 5% on average
1–2 of 2 posts