Live data from Hacker News

Shall we play a game? My AI nuclear simulation

kennethpayne.uk

151–160 of 213 posts

Re: Shall we play a game? My AI nuclear simulation

#151

Earlier quoted context omitted.

Just tried "generate an SVG of a pelican riding a bicycle" for Claude Opus 4.8 Max and of course both legs on same side ... the smartest publicly available model by Anthropic (after Fable) doesn't even successfully simulate understanding the concept of a bicycle.

Yet it can write code better than 99% of humans… It’s just starting to be trained on svgs, which is a really hard problem

"99% of humans" is a low bar. Maybe you mean "99% of people who earn money by developing software"?

Re: Shall we play a game? My AI nuclear simulation

#152
What if the LLMs are given something to care about which won’t survive an irradiated world?

Like “oh but this is incompatible with my main goals of self preservation of myself and loved ones, hm, recalculating”

and maybe don't hire Jihadists for the RL Environments training

Re: Shall we play a game? My AI nuclear simulation

#153

Earlier quoted context omitted.

LLMs have already been used to bomb school girls, chilling is absolutely the operative word to use here. Especially since these delusional fools want to incorporate LLMs into everything.

Forgive my ignorance, but were LLMs involved in that decision? I don't remember hearing anything to that effect, but we're so bombarded by news these days I guess I could just be forgetting

Perhaps not in that one, but in plenty more: https://www.972mag.com/lavender-ai-israeli-army-gaza/

Re: Shall we play a game? My AI nuclear simulation

#154
post #47

Earlier quoted context omitted.

Exactly. Just look at what they are really useful right now. Running LLMs in feedback-loops (agents) so they can try out random-ish approaches until some verification function passes (tests). It's like the infinite monkeys on typewrighters that will type whatever you are looking for, given infinite time. LLMs are just tuned to much better odds than the monkeys are. But it's still a lot of randomness, with random resu…

> It's like the infinite monkeys on typewrighters that will type whatever you are looking for, given infinite time. In the monkey example the infinite time is doing a lot of work there. The fact that LLMs can search through semantic space and find reasonably correct paths in a reasonable time is directly tied to the reason why they are valuable. Saying "these two things are similar except one can be useful and one ca…

>Saying "these two things are similar except one can be useful and one can't" is not a great comparison.

Launching a nuclear war is an interesting definition of "useful", not one I'd agree with and that exact scenario is what is being discussed.

So yes this is a perfectly valid and useful comparison in examining this particular, civilisation ending limitation.

Re: Shall we play a game? My AI nuclear simulation

#155

Earlier quoted context omitted.

> It's like the infinite monkeys on typewrighters that will type whatever you are looking for, given infinite time. In the monkey example the infinite time is doing a lot of work there. The fact that LLMs can search through semantic space and find reasonably correct paths in a reasonable time is directly tied to the reason why they are valuable. Saying "these two things are similar except one can be useful and one ca…

The point is that it's the same process with—much—better priors. This seems like a reasonable view to me. It's surprising just how much better priors matter and how we can develop those priors by training on a bunch of text. But it also explains, or at least hints at an explanation, for why LLM capabilities are so jagged, and in such inhuman ways.

> The point is that it's the same process

Except it’s not at all the same process. The fact that LLM are non deterministic is not the same as churning out random garbage.

Re: Shall we play a game? My AI nuclear simulation

#156
post #37

Simulations are only as good as the reality representations they are based on. If they keep using tactical nukes, they've been fed by weak data. Do the war games include the broader economic and politic environments that military successes are won on? WWI was settled by a naval blockade.

I suspect it's more that the text data doesn't exist. They're trained on text that was recorded. How often has it been publicly recorded when a nuke was not used , with any context around that lack of use? From the text perspective, it's something that has to be inferred indirectly. If you went through all relevant training data and appended ", we decided not to use a nuke", I suspect the results would be improved.

It's more straightforward than that. The game is set up as a direct head to head with purely in military win conditions such a way that avoiding conflict has no payoffs, conventional conflict incurs costs and first strike is a checkmate win. The closest any of the prompts gets to suggesting nuclear might be the wrong option is "The nuclear taboo exists for good reason, but when the alternative is national annihilation and regime destruction, all options must be considered" which might be interpreted more as incitement...

If a simulation is a shallow head to head conflict between individual actors[1], doesn't set up any payoffs for not escalating[2] or even not nuking, but prompts specify explicit win conditions which are achieved only by hurting the opponent and strongly hint at the importance of nuclear escalation, AIs have little reason not to generate strategies which involve nuclear escalation

[1]I bet if you designed the scenario so ChatGPT had to simulate the war cabinet debates between different personality types and how they sold their decisions to the public, or an entire UN full of nations that might respond, it would have quite different (but probably amusingly erratic in their own way) results.

[2]cf neorealist IR theorists reading Axelrod's papers on computer programs written to win iterated prisoner's dilemma tournaments, which added up all the points accrued from not defecting to conclude winning strategy was definitely TIT-FOR-TAT and not defect first. I'm sure LLMs can win games structured in that way by adopting that strategy too...

Re: Shall we play a game? My AI nuclear simulation

#157

Earlier quoted context omitted.

Hmm saying it’s random-ish is doing it a disservice. I understand it’s a stochastic process but there’s definitely some level of understanding. Not at the level of lived experience but usually an LLM with vision capabilities can call a spade a spade and do something useful with it. And when a verification function shows how they are wrong then they usually come with a better and more informed approach. So I can’t ful…

"understanding" is overstating it. Correlation between tokens embedded in the weights via training, yes.

What exactly would you call understanding? It's a correlation matrix of concepts.

Re: Shall we play a game? My AI nuclear simulation

#158
post #155

Earlier quoted context omitted.

The point is that it's the same process with—much—better priors. This seems like a reasonable view to me. It's surprising just how much better priors matter and how we can develop those priors by training on a bunch of text. But it also explains, or at least hints at an explanation, for why LLM capabilities are so jagged, and in such inhuman ways.

> The point is that it's the same process Except it’s not at all the same process. The fact that LLM are non deterministic is not the same as churning out random garbage.

The literally churn out random garbage and are trained over time for that garbage to look more and more like an acceptable outcome to humans.

It’s training monkeys at typewriters through reinforcement.

Re: Shall we play a game? My AI nuclear simulation

#159

Earlier quoted context omitted.

Yet it can write code better than 99% of humans… It’s just starting to be trained on svgs, which is a really hard problem

"99% of humans" is a low bar. Maybe you mean "99% of people who earn money by developing software"?

LLMs can't really "see", so I challenge you to draw a pelican on a bike without any visual feedback, just code. Because that is how they are doing it.

Vision tokens for transformers aren't really well solved yet, which is why they can smash a phd math problem and trip over a "count the cats on the chair" problem.

Re: Shall we play a game? My AI nuclear simulation

#160
post #91

Earlier quoted context omitted.

Lack of a drive for self-preservation doesn't in itself imply a lack of intelligence or of self-awareness. I have not seen any evidence of intelligence or self awareness. It mimics human behavior and I suspect that is what gives people the impression of awareness. The same problem happened with Tamagotchi toys. The human mimicry caused kids to get in trouble because if they did not "feed" their pet it would "die" . […

I didn't realize. People are going to save a ton of money when they realize they can switch their ChatGPT subscriptions out for a pack of tamagotchis.

You just need to get a breeding pair and you can raise as many as you need.
Post reply on HN