Mostly write code with GLM5.3 and found that using Muse 1.3 to make detailed summary bugs works well, especially for Rust. Somehow muse is really good at rust. And from what I can tell even using the original model (GLM here) against that list to go fix it seems fine even if it is the model that made the mistake.
They seem to write 100s of tests too, but don't really have much confidence in those tbh much like author. I view it as a bonus. Tests are computationally cheap so knock yourself out Mr AI.