I'm surprised at both the article and the paper - both seem very hyperbolic. This is LLMs competing against doctors in a way that is heavily weighted in the LLMs favour, which does not represent clinical practice. These reasoning cases are not benchmarks for doctors, they are learning tools. I think it's important to note that diagnosis also relies on accurate description of the patient in the first place, and the in…
Simply getting the "high score" on this evaluation is not necessarily good medical treatment.