Live data from Hacker News

Show HN: Watch LLMs play 21,000 hands of Poker

pokerbench.adfontes.io

1–10 of 20 posts

Show HN: Watch LLMs play 21,000 hands of Poker

#1
PokerBench is my attempt at a new LLM benchmark wherein frontier models play Texas Hold'em in an arena setting. It also features a simulator to view individual games and observe how the different models reason about poker strategy. Opus/Haiku, Gemini Pro/Flash, GPT-5.2/5 mini, and Grok 4.1 Fast Reasoning have all been included.

All code -> https://github.com/JoeAzar/pokerbench

Show HN: Watch LLMs play 21,000 hands of Poker
pokerbench.adfontes.io

Re: Show HN: Watch LLMs play 21,000 hands of Poker

#2
Finally, a way to settle the model wars that actually matters: Texas Hold'em. That 3D replay view is sick! ♠♦ I spent way too long watching the replay on Game 2a58900d. It’s wild to see the chain of thought mapped against the betting rounds. It really exposes when a model is hallucinating a strong hand versus actually calculating pot odds. This 'PokerBench' might actually become the standard for measuring agentic risk-taking.

Re: Show HN: Watch LLMs play 21,000 hands of Poker

#6
post #2

Finally, a way to settle the model wars that actually matters: Texas Hold'em. That 3D replay view is sick! ♠♦ I spent way too long watching the replay on Game 2a58900d. It’s wild to see the chain of thought mapped against the betting rounds. It really exposes when a model is hallucinating a strong hand versus actually calculating pot odds. This 'PokerBench' might actually become the standard for measuring agentic ris…

yeah the 3d view is amazing

Re: Show HN: Watch LLMs play 21,000 hands of Poker

#7
post #5

Fun, any idea how much would be the cost per game? I am worried 160 isnt a big enough sample size.

It greatly depends on the models. The 6-handed setup with Opus and Pro cost about $30/game. The 4-handed setup with just small models was $6/game. I'd love to run more but I already spent quite a bit as it is.

Re: Show HN: Watch LLMs play 21,000 hands of Poker

#8

Do you have idea why smaller models are better then large ones?

I've seen some theories tossed around but I don't think I'm qualified to offer an authoritative answer. Gemini 3 Pro specifically seems to be consistently "tighter" and more passive than Flash.
Post reply on HN