> the models generalize well only on tasks within a small neighborhood of the specific tasks they've been trained on
Unless frontier labs have surprisingly trained their models in the exact tasks my team works on, this is patently false. We are getting very good results on automation and I'm bullish we will be able to mostly remove humans in the loop for most of our infra tasks by the end of the year.
I have no opinion on the other theses, but given that OP doesn't back up these claims in any way, I have my doubts about the conclusions of this article.