We blind-tested ChatGPT, Claude, and Gemini on 20 everyday tasks (Open Dataset)
1–3 of 3 posts
Re: We blind-tested ChatGPT, Claude, and Gemini on 20 everyday tasks (Open Dataset)
#2[flagged]
Re: We blind-tested ChatGPT, Claude, and Gemini on 20 everyday tasks (Open Dataset)
#3It might have been more helpful if the model names and versions were mentioned, as there are new versions coming every other week from these organisations. Also worth noting there's no coding task in the whole set, so this doesn't really tell you where each one actually leads.