The post really reminds me of a component of a platform I’m currently building. The problem really with this is finding not just good questions that do not discriminate individual models but also providing a good sample size (eg not just 60) to get really some meaningful results. And even if you have those, there is a drift in the quality of responses. I'm the founder of Pulze.ai, a B2B SaaS Dynamic LLM Automation Pl…
Playground and account are for free