Top AI models fail at >96% of tasks
zdnet.com
Top AI models fail at >96% of tasks
1–10 of 13 posts
Re: Top AI models fail at >96% of tasks
#2Models released a few days ago, Opus 4.6 and GPT 5.3, haven't been tested yet, but given the performance on other micro-benchmarks, they will probably not be much different on this benchmark.
Re: Top AI models fail at >96% of tasks
#3Re: Top AI models fail at >96% of tasks
#4This paper creates a new benchmark comprised of real remote work tasks sourced from the remote working website Upwork. The best commercial LLMs like Opus, GPT, Gemini, and Grok were tested. Models released a few days ago, Opus 4.6 and GPT 5.3, haven't been tested yet, but given the performance on other micro-benchmarks, they will probably not be much different on this benchmark.
One of the tasks was "Build an interactive dashboard for exploring data from the World Happiness Report." -- I can't imagine how Opus4.5 could've failed that.
Re: Top AI models fail at >96% of tasks
#5Then go ahead and use AI to fix this: https://gitlab.gnome.org/GNOME/mutter/-/issues/4051
Re: Top AI models fail at >96% of tasks
#6Re: Top AI models fail at >96% of tasks
#7Re: Top AI models fail at >96% of tasks
#8Re: Top AI models fail at >96% of tasks
#9Re: Top AI models fail at >96% of tasks
#10You think they don't? You think AI can replace programmers, today? Then go ahead and use AI to fix this: https://gitlab.gnome.org/GNOME/mutter/-/issues/4051