Live data from Hacker News

Shall we play a game? My AI nuclear simulation

kennethpayne.uk

31–40 of 213 posts

Re: Shall we play a game? My AI nuclear simulation

#31
post #15

FYI -- there's no such thing as a "tactical" nuke. A nuclear bomb is a nuclear bomb.

There are tactical and strategic nuclear weapons. https://en.wikipedia.org/wiki/Tactical_nuclear_weapon

In the cold war arms manufacturer got very creative: e.g jeep mounted nuclear weapons https://www.militarytrader.com/mv-101/the-atomic-jeep

Re: Shall we play a game? My AI nuclear simulation

#32
post #12

We're getting to the point where high-level officials are coming to LLMs for advice. And the quirky personalities of the LLMs, however much it pains me to say this, are probably well-placed to remind us that they aren't human. My personal hope is that this will result in less delegation when it comes to making important decisions.

I have so little faith in "high-level" officials that I prefer our AI overlords.

That's an entirely valid point of view!

Re: Shall we play a game? My AI nuclear simulation

#33
Simulations are only as good as the reality representations they are based on. If they keep using tactical nukes, they've been fed by weak data. Do the war games include the broader economic and politic environments that military successes are won on? WWI was settled by a naval blockade.

Re: Shall we play a game? My AI nuclear simulation

#35
Sonnet, GPT-5.2, Gemini Flash, in a set of 21 games, where conclusions are drawn from the LLMs self reported reasoning.

This is like writing a paper about kids in a literal sandbox fighting over ‘territory’.

The models employed don’t indicate the actual extents of machine reasoning even as we currently recognize them. They certainly don’t have the metacognition necessary to accurately understand their own reasoning. As we’ve seen with recent papers on how LLMs do math there’s a complete disconnect between actual and reported mechanism.

“Chilling” shouldn’t be the take away here.

Re: Shall we play a game? My AI nuclear simulation

#36
These papers usually have poor stability to prompting and rerunning. It would be nice if we had some kind of meta-evaluation metric where rewriting the prompt conditions or varying the input params could be used to determine how stable a result is.

Regardless, it's definitely true that AI agents have different priorities from us. That's what alignment is about anyway.

Re: Shall we play a game? My AI nuclear simulation

#37

Simulations are only as good as the reality representations they are based on. If they keep using tactical nukes, they've been fed by weak data. Do the war games include the broader economic and politic environments that military successes are won on? WWI was settled by a naval blockade.

I suspect it's more that the text data doesn't exist. They're trained on text that was recorded. How often has it been publicly recorded when a nuke was not used, with any context around that lack of use?

From the text perspective, it's something that has to be inferred indirectly. If you went through all relevant training data and appended ", we decided not to use a nuke", I suspect the results would be improved.

Re: Shall we play a game? My AI nuclear simulation

#38
post #37

Simulations are only as good as the reality representations they are based on. If they keep using tactical nukes, they've been fed by weak data. Do the war games include the broader economic and politic environments that military successes are won on? WWI was settled by a naval blockade.

I suspect it's more that the text data doesn't exist. They're trained on text that was recorded. How often has it been publicly recorded when a nuke was not used , with any context around that lack of use? From the text perspective, it's something that has to be inferred indirectly. If you went through all relevant training data and appended ", we decided not to use a nuke", I suspect the results would be improved.

...the entire Cold War?

Re: Shall we play a game? My AI nuclear simulation

#40
post #12

We're getting to the point where high-level officials are coming to LLMs for advice. And the quirky personalities of the LLMs, however much it pains me to say this, are probably well-placed to remind us that they aren't human. My personal hope is that this will result in less delegation when it comes to making important decisions.

GPT-4o was considered harmful, because it imitated human connection too much, not because it was so "smart" or capable.

It was for sure a deliberate decision to make LLMs seem less like a human companion and more like an obedient servant in newer releases.

Post reply on HN