I wonder if the authors have tried incorporating error feedback from Lean into their models. Work from 2023 [1] showed general purpose models did better when they were able to incorporate error feedback, humans incorporate error feedback, but none of the SOTA models on minif2f seem to. [1]: https://arxiv.org/abs/2310.04353
I'm surprised those even use actual lean code instead of like raw type theory.