Forget about GPT-N and DALL-E for a second and look at the NRO's Sentient program. It's the closest thing out there to a known real attempt at making something like Skynet. It's trying to automate the full TCPED (tasking, collection, processing, exploitation, and dissemination) cycle of global geointelligence, and well, it's actually trying to do even more than that, but that is unfortunately classified. Except it definitely hasn't achieved what it is trying to do, and probably won't. My wife happens to be the enterprise test lead for one of the main components of this system, where "enterprise test" means they try to get the next versions with all the latest greatest features of all components working together in a UAT environment where each of the involved agencies signs off before the new capabilities can go live.
It's amusing to see the kinds of things that grind the whole endeavor to a halt. Probably more than anything, it's issues with PKI. Networked components can't even establish a session and talk to each other at all if they don't trust each other, but trust is established out of band. Classified spy satellite control systems don't just trust the default CAs that Mozilla says your browser should trust. Intelligent or not, there is no possible code path by which the software itself can decide it doesn't care and it will trust a CA anyway or ignore an expired cert and continue talking to some downstream component because doing so is critical to its continued ability to accomplish anything other than sending scrambled nonsense packets into the ether. GPT-N is great at generating text, but no amount of getting better at that will ever make it capable of live-patching code running in read-only memory to give it new code paths it wasn't compiled with. That has nothing to do with intelligence. It just isn't possible at all. You have to have the physical ability to move in space and type characters into a workstation connected to a totally separate network that code is developed on, which is airgapped from the network code is run on.
We seem to be pretty far from even attempting to make distributed software systems that can honest to God do much of anything at all without human monitoring and intervention beyond several-minute at most batch jobs like generate a few paragraphs of text. Sure, that's great, but where is the leap from that to figuring out why an entire AS goes black and half your system disappears because of a typo'd BGP update that then needs to be fixed out of band over the telephone because you can no longer use the actual network, let alone controlling surveillance and weapons systems that aren't networked to the systems code is being developed on? What is the pathway by which a hugely scaled-up ANN is able to bypass the required human steps that propagate feedback from runtime to development in order to achieve recursive self-improvement? Because that is what it would take to gain control of military systems rather than someone's website by purely automated means, and I don't see how it's even the same class of problem. It isn't a research project any AI team is even working on, I have no idea how you would approach it, but it's the kind of nitty-gritty detail you'd have to actually solve to build an automated world conquering system.
It seems like the answer tends to just be "well, this thing will be smarter than any human, so it'll figure it out." That isn't a very satisfying answer, especially when I'm reasonably sure the person saying it has absolutely no idea how security measures and the resulting operational challenges of automating military command and control systems even work.