What about the open source models? I remember from the trading benchmarks Deepseek performed pretty well.
Show HN: Watch LLMs play 21,000 hands of Poker
11–20 of 20 posts
Re: Show HN: Watch LLMs play 21,000 hands of Poker
#12Fun, any idea how much would be the cost per game? I am worried 160 isnt a big enough sample size.
It greatly depends on the models. The 6-handed setup with Opus and Pro cost about $30/game. The 4-handed setup with just small models was $6/game. I'd love to run more but I already spent quite a bit as it is.
Re: Show HN: Watch LLMs play 21,000 hands of Poker
#13Earlier quoted context omitted.
It greatly depends on the models. The 6-handed setup with Opus and Pro cost about $30/game. The 4-handed setup with just small models was $6/game. I'd love to run more but I already spent quite a bit as it is.
Yeah thats costly, 160 games still gives about 1000+ total decisions and you can see some trends on how they think about the game state.
Re: Show HN: Watch LLMs play 21,000 hands of Poker
#14Re: Show HN: Watch LLMs play 21,000 hands of Poker
#15Do you have any idea why the win rate for GPT-5.2 is higher than Gemini 3 Flash yet the former loses money while the latter earns money? Is it just bet sizing (betting more when it has a good hand) or something else?
Re: Show HN: Watch LLMs play 21,000 hands of Poker
#16Re: Show HN: Watch LLMs play 21,000 hands of Poker
#17People looking into this a little too much, looks to me like random walk. You should try reinitiating the trial (or have multiple running) and see if the ranking is robust.
Re: Show HN: Watch LLMs play 21,000 hands of Poker
#18People looking into this a little too much, looks to me like random walk. You should try reinitiating the trial (or have multiple running) and see if the ranking is robust.
Wdym exactly? I ran 163 games, are you suggesting more games or something else?