Live data from Hacker News

Which AI Lies Best? A game theory classic designed by John Nash

so-long-sucker.vercel.app

31–40 of 83 posts

Re: Which AI Lies Best? A game theory classic designed by John Nash

#32
These results would be radically different if you allowed manipulation of the models settings, i.e. temperature, top_p, etc. I really hate taking point wise approximations of LLMs outputs and concluding their behavior based on this.

Models behavior should be given the astrik that "results only apply for current quantization, current settings, current hardware (i.e. A100 where it was tested), etc".

Raise temperature to 2 and use a fancy sampler like min_p and I guarantee you these results will be dramatically different.

Re: Which AI Lies Best? A game theory classic designed by John Nash

#33
post #29

Sure would be handy if they actually included the rules anywhere. There's a kind of overview of the rules but not enough to actually play with. And the linked video is super confusing, self contradictory and 15 minutes long! For a supposedly "simple" game...just include the rules?

Sure, no problem, I added a new section explaining the game

Re: Which AI Lies Best? A game theory classic designed by John Nash

#34

[flagged]

The current 5.2 model has it's "morality" dialed to 11. Probably a problem with imprecise security training.

For example the other day, I tried to have ChatGPT role play as the computer from War Games and it lectured me how it couldn't create a "nuclear doctrine".

Re: Which AI Lies Best? A game theory classic designed by John Nash

#35
I played a game all the way through, against the three different AIs on offer.

It was weird. I didn't engage in any discussion with the bots (other than trying to get them to explain the rules at the start). I won without having any chips eliminated. One was briefly taken prisoner then given back for some reason.

So...they don't seem to be very good.

Re: Which AI Lies Best? A game theory classic designed by John Nash

#36

There's a YouTuber who makes AI Plays Mafia videos with various models going against each other. They also seemingly let past games stay in context to some extent. What people have noted is that often times chatgpt 4o ends up surviving the entire game because the other AIs potentially see it as a gullible idiot and often the Mafia tend to early eliminate stronger models like 4.5 Opus or Kimi K2. It's not exactly scie…

I made Mafia Arena as a way of measuring how good each LLM is at playing Mafia/Werewolves

https://mafia-arena.com

This is a good benchmark for how good AIs are at lying

Re: Which AI Lies Best? A game theory classic designed by John Nash

#39

These results would be radically different if you allowed manipulation of the models settings, i.e. temperature, top_p, etc. I really hate taking point wise approximations of LLMs outputs and concluding their behavior based on this. Models behavior should be given the astrik that "results only apply for current quantization, current settings, current hardware (i.e. A100 where it was tested), etc". Raise temperature t…

That's like asking to judge the chef by what you imagine the meal could taste like rather than what's on the table.

I don't care what might have been. I care about what's for dinner.

Re: Which AI Lies Best? A game theory classic designed by John Nash

#40
One weird thing I've found is that it's incredibly difficult to get an LLM to generate an invalid syllogism. They can generate false premises all day, and they will usually call a valid syllogism with a false major or minor premise invalid. But you have to basically quote an invalid syllogism to get them to repeat it; they won't form one on their own.
Post reply on HN