Live data from Hacker News

How A.I. Conquered Poker

nytimes.com

161–170 of 188 posts

Re: How A.I. Conquered Poker

#161

Earlier quoted context omitted.

> You can only win in poker if you recognize how your opponent deviates from the optimal strategy and then play a strategy to exploit him. This is not true. If you play optimal strategy, you will win against any opponent except one that plays optimal as well, in which case you'll break even. But, of course you'll win a lot more if you are able to adapt your strategy to exploit any weaknesses that you have detected. A…

> If you play optimal strategy, you will win against any opponent except one that plays optimal as well How does the example you're responding to not win 50% of the time? Rock v Rock = Tie Rock v Paper = Loss Rock v Scissors = Win The optimal game theoretic play of randomly choosing rock paper scissors is inferior play against this particular opponent. All that game theoretic perfect play gets you is the benefit of g…

Poker isn't like this. Think rock, pair, scissors, crap where rules are that crap always loses. Now uniform strategy between rps will yield profit vs an opponent who plays crap sometimes.

It's very easy to play crap in poker

Re: How A.I. Conquered Poker

#162
post #132
post #52

(Former pro and high stakes player, occasional solver developer) This is actually one of the best poker articles I've ever seen in generalist media. Not too clickbaity, reasonable high level overview of game theory, a (very accurate IMO) quote from old pro Erik Seidel about the state of the game just 15 years ago, a discussion on variance vs. EV and results, and most importantly, an emphasis on math, randomization te…

Similar story here. I’ve always played, mostly online, before Black Friday (pretty sure I was a minor when it happened) but played all the offshore sites through college. Professional software career took off and bounced out, finally ~3y ago I decided to get back into it and I would only play challenging games (only thing challenging about a 1/3 or tourneys is the grind). So, started with 2/5, within a week was playi…

You lost $50k in a 25/50 online game? Was the pot 2000 BBs??? I've never even heard of anything close to it.

Re: How A.I. Conquered Poker

#163

For some time I've been pondering about creating a multiplayer poker AI that would (1) play reasonably well but not at a superhuman level (2) not require expensive hardware. Frustratingly, this seems to be an unsolved problem. Pluribus is only 6-max, requires fixed starting stack sizes and requires a lot of hardware.

I mean, it may not have been publicly solved, but there are definitely bots out there in the wild that are beating human fields on online poker sites. But also I think Pokersnowie is a bot made with more traditional ML methods (i.e. training on hand histories rather than solving an entire game tree with CFR), which basically fits the description.

Is there enough public info to reimplement PokerSnowie? I only know it uses neural nets in some way.

Re: How A.I. Conquered Poker

#164
post #79

Earlier quoted context omitted.

Interesting, I read some Sklansky books about 15 years ago and he talked about semibluffs, where you're partly bluffing but you've also got some chance to win. Sounds like strategy has evolved into doing the exact opposite.

The distinction is that semi-bluffs are for when there are still cards to be dealt.

Would we call that thin value betting on the river?

Re: How A.I. Conquered Poker

#165
post #132

Earlier quoted context omitted.

Similar story here. I’ve always played, mostly online, before Black Friday (pretty sure I was a minor when it happened) but played all the offshore sites through college. Professional software career took off and bounced out, finally ~3y ago I decided to get back into it and I would only play challenging games (only thing challenging about a 1/3 or tourneys is the grind). So, started with 2/5, within a week was playi…

You lost $50k in a 25/50 online game? Was the pot 2000 BBs??? I've never even heard of anything close to it.

Casino, there was no max. A lot of players sitting on 100k+. There were far bigger hands than this one.

Re: How A.I. Conquered Poker

#166
post #31

Earlier quoted context omitted.

Hasn't Facebook's Pluribus come pretty close to "solving" no-limit holdem? I don't know a lot about poker so I can't really assess their claim's validity.

Not even mentioning Pluribus is just bad journalism

The article is mostly about a change in the industry and adaptation of solver strategies by human players of all levels.

Pluribus had approximately zero impact on that.

Re: How A.I. Conquered Poker

#167

Earlier quoted context omitted.

Where is the variance-reduction technique discussed? I looked at this paper https://www.science.org/doi/abs/10.1126/science.aay2400 https://par.nsf.gov/servlets/purl/10077416 and it just says > Finally, we tested Libratus against top humans. In January 2017, Libratus played against a team of four top HUNL specialist professionals in a 120,000 hand Brains vs. AI challenge match over 20 days. The participants were Jaso…

>The remaining $120,000 was divided among them based on how much better the human did against Libratus than the worstperforming of the four humans. Surely the correct strategy here is for the human players to collude to give as much money as possible to a single player and then split the money afterwords, no? Also, the fact that they players can only gain money without losing anything likely changes their play somewh…

The four humans were getting $120,000 between them. Their share of that was dependent on how much better they did than the other humans. That means there was no incentive to collude.

Top pro poker players understand the value of money. They weren't treating it as a freeroll and anyone that has seen the hand histories can confirm that.

Re: How A.I. Conquered Poker

#168

Earlier quoted context omitted.

Normally 10,000 hands would be too small a sample size but we used variance-reduction techniques to reduce the luck factor. Think things like all-in EV but much more powerful. It's described in the paper.

Where is the variance-reduction technique discussed? I looked at this paper https://www.science.org/doi/abs/10.1126/science.aay2400 https://par.nsf.gov/servlets/purl/10077416 and it just says > Finally, we tested Libratus against top humans. In January 2017, Libratus played against a team of four top HUNL specialist professionals in a 120,000 hand Brains vs. AI challenge match over 20 days. The participants were Jaso…

It's in the supplementary material of the 2019 paper: http://www.cs.cmu.edu/~noamb/papers/19-Science-Superhuman_Su... . Look at the "Variance reduction via AIVAT" section.

Re: How A.I. Conquered Poker

#169
post #39

What teaching or training tools are out there for a very average player at no limit Texas hold’em who just wants to get at bit better to a respectable level at a modest time commitment , and does not need to be a pro-level player?

Yeah, really I just want an open-source poker solitaire that asks me what I should do and then shows me the odds after I answer.

The odds are a small part of what you have to consider.

If you want to quickly improve to a proficient level there's no quicker way than to read Theory of Poker and Hold 'em for Advanced Players. Both of these books focus primarily on limit poker, but the concepts are critical for no limit as well. And you'll realize there's a lot more strategy and nuance in limit poker than you thought.

Re: How A.I. Conquered Poker

#170

Earlier quoted context omitted.

Alphastar also didn't play with the same limitations that a human has. Even after removing its ability to see the entire map and finally forcing it to scroll around, alphastar never misclicks (so its APM==EPM) and can still blast nearly unlimited APM for short bursts as long as its "average APM" over an x-second period matched human's APM. I believe Alphastar would generate more interesting strategies if we limited a…

> Alphastar also didn't play with the same limitations that a human has. Even after removing its ability to see the entire map and finally forcing it to scroll around, alphastar never misclicks (so its APM==EPM) and can still blast nearly unlimited APM for short bursts as long as its "average APM" over an x-second period matched human's APM. > I believe Alphastar would generate more interesting strategies if we limit…

Edit 2: Reading through the "supplementary data" of the 2019 paper, it definitely appears that the AlphaStar which reached grandmaster was not limited in the same ways as the 2017 paper would suggest. x/y positions of units are not determined visually, but fed directly from the API. So AlphaStar absolutely can just run Attack(Position: carrier_of_interest->pos.x) and not mis-click. It's "map" / "vision" is really just a bounding box of an array of every entity/unit on the map and all the things that a human would have to spend APM to manually check (precise energy level, precise hit points remaining, exact location of invisible units, etc). See [7]. DeepMind showed they have some fixed time delays to emulate human experience, if they had a position fuzzer, they would have mentioned it. I'm reasonably convinced they gave AlphaStar huge advantages even in the 2019 version that was 'nerfed' from the 2018 version. The 2017 paper was a more ambitious project IMO that didn't quite get fully developed.

Edit: 7 minutes after writing this I re-read the original paper[-1]. https://arxiv.org/pdf/1708.04782.pdf page 6 and 7 make it clear that DeepMind limited themselves to SpatialActions, so they cannot tell units "Attack Carrier" but have to say "Attack point x,y" (and x,y also has to be determined visually, not through carrier_of_interest->pos.x ). It's still not clear in the paper if any randomness is added to Attack(x,y).

Additionally, I have some serious concerns about assuming that the design decisions made in this 2017 paper were actually used in the implementation of the 2019 Alphastar demo vs TLO and MaNa. The paper claims "In all our RL experiments, we act every 8 game frames, equivalent to about 180 APM, which is a reasonable choice for intermediate players." I would agree with this choice! But [5][6] indicates that Alphastar's APM spiked to over 1500 APM in 2019! And even in moments when a human reaches that APM, their EPM would be an order or magnitude lower, whereas Alphastar's EPM matches its APM.

Original post:

Thank you so, so much for adding to the discussion! Would love to chat more about this if you see my reply and feel like it.

Regarding "mis-clicking", my understanding was that AlphaStar used Deepmind's PySC2[0][1], which in turn exposes Blizzard's SC2 API[2][3].

Here is the example for how to tell an SCV to build a supply depot:

  Actions()->UnitCommand(
    unit_to_build,
    ability_type_for_structure,
    Point2D(
      unit_to_build->pos.x + rx * 15.0f,
      unit_to_build->pos.y + ry * 15.0f
    )
  );
  
where unit_to_build->pos.x and unit_to_build->pos.x are the current position of the SCV and rx and ry are offsets. It's possible to fuzz this with some randomness, and indeed in the example, rx and ry are actually random (because the toy example just wants to create a supply depot in a truly random nearby spot, it doesn't care where). But the API doesn't attempt to "click" on an SCV and then use a hotkey and then "click" somewhere else. The API will never fail to select the correct SCV. It will also build precisely at the coordinates provided.

Point 1: Even if DeepMind added a fuzz to this method to make it so AlphaStar can "misclick" where the depot gets built, it cannot accidentally select the wrong SCV to build that depot. (Possibly wrong, as they could be using SpatialActions, see below)

Point 2: Most bot-makers wouldn't add a random fuzz to the depot placement coordinates to make their AI worse and I'd be super surprised if there was hard evidence somewhere that Alphastar had such a fuzz. (This is my main concern.)

My personal conclusion was that anything which looks like a "misclick" is, in fact, a "mis-decision". A human can decide "I want my marines to attack that carrier" but accidentally click the attack onto a nearby interceptor. I didn't think Alphastar could do that because I assumed it would use the Attack(Target: Unit) method instead of Attack(Target: Point) in that scenario -- and even if they used Attack(Target: Point) it would be used as Attack(Target: carrier->pos.x).

However, I realize now that they could be doing everything with SpatialActions (edit: it does, see paper[-1] pp. 6-7) (select point, select rect's)[4], and that they could have a implemented a randomness layer to make alpha star literally mis-click.

I suppose I would need to test this API and dive into the replay files to first see if its possible to discern the different between Attack(Target: carrier_of_interest) and Attack(Target: carrier_of_interest->pos.x). Then, even if Alphastar is using the latter, it's still not clear that there's an additional element of randomness outside of the AI/ML control.

Has anyone already done an analysis of the replay files on this level, or has DeepMind released hard info on how they're controlling the bot?

-1: https://arxiv.org/pdf/1708.04782.pdf

0: https://www.youtube.com/watch?v=-fKUyT14G-8

1: https://github.com/deepmind/pysc2

2: https://github.com/Blizzard/s2client-proto

3: https://blizzard.github.io/s2client-api/index.html

4: https://blizzard.github.io/s2client-api/structsc2_1_1_spatia...

5: https://www.alexirpan.com/2019/02/22/alphastar.html

6: https://deepmind.com/blog/article/alphastar-mastering-real-t...

7: https://ychai.uk/notes/2019/07/21/RL/DRL/Decipher-AlphaStar-...

Post reply on HN