Even GPT-3 seems to have trouble with world-modeling, it writes convincing text that has all the signs and form of good prose, but the output repeatedly violates physics and common sense in funny ways.
I know just enough about machine learning to have dangerously unrealistic expectations, but I'd like if I could reasonnably hope to see signs of a shared representation or shared knowledge between say image labeling and language modeling. This looks like a very concrete data point to take if you care about generality. Maybe then we can seriously talk about world-modeling.