There is a lot of hate in the comments but there is some merit to the post existing: 1. Even if the task is unreasonable, it is good to showcase that the LLM will perform poorly - warning not to be used for diabetes. 2. As it is a probabilistic model, the approach was to execute it multiple times and look at the distribution. They also tried to minimize variance: "All at the lowest randomness setting these models off…
We've been seeing examples of this constantly since 2022. How many more do we need?