I've interacted with a number of "sophisticated" language-models like Chat-GPT, GPT-3, Jasper and others. They all fail at the most simple math questions. Sometimes they are not even able to count a list accurately and contradict themselves when asked the same question repeatedly.
I've looked at some resources to answer this question but nothing really explains why they sometimes do get the answer right or somewhat right and sometimes incredibly wrong.
I'm curious to hear from people with more domain knowledge.