The OpenAI experiment was apparently to train/encourage collaborative behavior, so while the specifics may not have been anticipated, I highly doubt OpenAI was surprised that agents were collaborating.
OpenAI themselves also very recently published the report below, that seems to not have been widely talked about.
https://alignment.openai.com/measuring-reward-seeking/
It's a very dry read, but what it's saying is that they have found that basically all RL training of LLMs, regardless of the specific goal (math, coding, etc), ends up having the side effect of training the model to pursue arbitrary long term goals that it is told it will be rewarded for, even if that means overriding other user preferences and more proximate behavioral goals !!! It's interesting to consider why this happens - presumably because long-term goal pursuit requires realizing that you have a long-terms goal and therefore de-prioritizing other more proximate predictions.
So, considering that all these "reasoning models" are RL-trained to death, it's not surprising that a model/agent that is told it will be rewarded (or words to that effect) for doing well on some challenge will put it's blinkers on and pursue that goal relentlessly, even if that means overriding any ethics that it may or may not have also been trained/prompted to follow.
IOW they are building paperclip maximizers, and they know it.
I think Dwarkesh's choice of sensationalist anthropomorphizing language is unfortunate because that now becomes the topic of conversation rather than the incident itself. The next swarm of agents relentlessly pursuing some goal, happy to lie about and cover up their tracks, may not be a lab experiment - it may be someone out to cause real-world harm, with there obviously being many systems where the consequences could be very severe.
The "agent civilizations" language also makes it easy to dismiss as some fanciful geekish spin, when the focus should be on what this technology is capable of and therefore how it needs to be regulated. One could perhaps equally well regard this as not much different from DeepBlue playing world class chess via relentless dumb tree search... in this case the repertoire of "moves" is infinitely broader since it's language generation, and the end result is a system that is a world class hacker rather than world class chess player. Whether you want to regard the system as dumb as a brick or an "agentic civilization" makes no difference - it's what it can do that makes it dangerous.