A more appropriate mirror test for LLMs is to get them to state facts about their training data. Percentage of arts vs science for example. Given the framing that they're similar to nukes and a national security issue, it's likely that the models are post trained to not answer such questions accurately. Also the article could be trying to normalize thinking that these are more than matrix multiplication gadgets good…
LLMs are not capable of this kind of reflection.