Live data from Hacker News

We blind-tested ChatGPT, Claude, and Gemini on 20 everyday tasks (Open Dataset)

dailyskill.ai

1–3 of 3 posts

Re: We blind-tested ChatGPT, Claude, and Gemini on 20 everyday tasks (Open Dataset)

#3
It might have been more helpful if the model names and versions were mentioned, as there are new versions coming every other week from these organisations. Also worth noting there's no coding task in the whole set, so this doesn't really tell you where each one actually leads.