Live data from Hacker News

AlphaStar: Mastering the Real-Time Strategy Game StarCraft II

deepmind.com

141–150 of 459 posts

Re: AlphaStar: Mastering the Real-Time Strategy Game StarCraft II

#141
post #23

Does anyone know if this opens the door to using these techniques for poker (given that they've now show success on games of imperfect information)? Thus far the solutions to poker have involved solving the game tree through raw computational power and clever methods of information collapsing: http://science.sciencemag.org/content/347/6218/145 But it seems the techniques used here might be both far more efficient, as…

And to add that, what if deepmind could get a poker history of the players at the table to create a poker profile of each player it is playing against. Having a percentage of a player's likelihood of folding and bluffing could keep it an advantage over a purely objective game theory aspect of the game. Maybe a certain player is more likely to bluff 5 hours into a game based off of player history analysis. Going further, imagine if DeepMind could get access to every players history outside of poker ( social media, purchasing records, medical records, etc ) and create a behavioral profile of each player.

Game theory, poker profile, behavior profile, etc could give it a serious advantage over human players in addition to its advantage of never getting tired, frustrated, etc.

Re: AlphaStar: Mastering the Real-Time Strategy Game StarCraft II

#142

Earlier quoted context omitted.

> Chess and Go don't have any form of micro and AIs are nevertheless dominant there. Yes, but chess and go have a tiny problem space compared to something like Starcraft. People want to see an AI win because it’s smart, not because it’s a computer capable of things impossible for humans. If the goal was perfect micro they could write computer programs to do that 10 years ago.

Then maybe we need a better game than StarCraft to test this on? Some kind of RTS that's less micro-heavy, perhaps? Maybe even an RTS where you can't give orders to individual units at all, like the Total War series? You can't fault the AI for winning at the game because of the way the game itself works. Even if you limit the AI to max human APM, it's still going to dominate in these micro-heavy battles because it's…

> Even if you limit the AI to max human APM, it's still going to dominate in these micro-heavy battles because it's going to make every one of its actions count.

right, and we saw that with the incredible precision with stalker blink micro. There are many ways you could make it more comparable to humans. They have already tried that by even giving it an APM.

> You can't fault the AI for winning at the game because of the way the game itself works.

But it does make the victory feel hollow when it wins using a "skill" that is unrelated to AI (having crazy high APM with perfect precision because its a computer). Micro-bots have been around for decades, and they are really good. The whole point of this exercise is to build better AI, not prove that computers are faster then humans.

It would like if they wanted robots to try and beat humans at soccer, and the robots won because they shoot the ball out of a cannon at 1000 KPH. They win, but not really by having the skills that we are trying to develop.

Re: AlphaStar: Mastering the Real-Time Strategy Game StarCraft II

#143
post #46

This is really impressive, I didn't expect starcraft to be played this well by a machine learning based AI. I'm excited to read the paper when it comes out! That said, I'm not sure I agree that it was winning mainly due to better decision making. For context, I've been ranked in the top 0.1% of players and beaten pros in Starcraft 2, and also work as a machine learning engineer. The stalker micro in particular looked…

In the mass stalker battles, the AI APM exceeded 1000 a few times, and no doubt that most of that was precisely targeted. Whereas a human doing 500 APM micro is obviously going to be far more imprecise. I think a far more interesting limitation would be to cap APM at 150 or so, or to artificially limit action precision with some sort of virtual mouse that reduced accuracy as APM increased.

If you are going to go to this extent, why not limit the computer by making it control a meat-puppet that has to physically manipulate the mouse and keyboard. And also give the computer carpal tunnel syndrome and get distracted by its questionable past.

For the people complaining about the information availability of 'one screen at a time': what is the difference between seeing the whole map with fog of war and quickly iterating the over the minimap at a very fast rate? Answer: a few milliseconds which would be irrelevant to a human anyway. Are they trying to make a computer that plays sc2 well or are they trying to make a computer that simulates expert human play to an indistinguishable level and which of those is more broadly applicable to other AI problems?

Re: AlphaStar: Mastering the Real-Time Strategy Game StarCraft II

#144
post #46

This is really impressive, I didn't expect starcraft to be played this well by a machine learning based AI. I'm excited to read the paper when it comes out! That said, I'm not sure I agree that it was winning mainly due to better decision making. For context, I've been ranked in the top 0.1% of players and beaten pros in Starcraft 2, and also work as a machine learning engineer. The stalker micro in particular looked…

In the mass stalker battles, the AI APM exceeded 1000 a few times, and no doubt that most of that was precisely targeted. Whereas a human doing 500 APM micro is obviously going to be far more imprecise. I think a far more interesting limitation would be to cap APM at 150 or so, or to artificially limit action precision with some sort of virtual mouse that reduced accuracy as APM increased.

How many of those 500 actions are actually useful? I haven't watched competitive StarCraft games for years but back when I did, rates were more like 300APM and even then the players basically spam clicked the background or selected random units non-stop and were probably only doing 50-100 actual effective actions.

Re: AlphaStar: Mastering the Real-Time Strategy Game StarCraft II

#146
post #46

This is really impressive, I didn't expect starcraft to be played this well by a machine learning based AI. I'm excited to read the paper when it comes out! That said, I'm not sure I agree that it was winning mainly due to better decision making. For context, I've been ranked in the top 0.1% of players and beaten pros in Starcraft 2, and also work as a machine learning engineer. The stalker micro in particular looked…

The results are obviously impressive, but even then there is a lot of work to do as far as learning efficiency goes: "The AlphaStar league was run for 14 days, using 16 TPUs for each agent. During training, each agent experienced up to 200 years of real-time StarCraft play. " MaNa probably played less than 2-3 years of Starcraft in his whole life (by that I mean 24hr x 365d x 3), and was learning with a much less foc…

Another way to think about it is that a human brain is mostly doing transfer-learning, on top of a 99%-baked deep net that was wired up during foetal development from our DNA, where that DNA-persisted model has "seen" hundreds of millions of years of training data.

Humans don't have to learn to process, recognize, and classify objects in visual sense-data, for example. We can do that from the moment we're born, because we already have hundreds of precisely-tuned "layers" laying around in our brains for doing just that. We just need to transfer-learn the relevant classes.

Re: AlphaStar: Mastering the Real-Time Strategy Game StarCraft II

#147

Earlier quoted context omitted.

I'll bet you that AlphaStarZero comes out in a year and just learns from scratch.

I'll take you up on that bet; they started with a version that tried to learn from scratch they seemed to have scrapped that approach.

I bet the very early internal versions of AlphaGo learned from scratch and didn't work very well either.

Re: AlphaStar: Mastering the Real-Time Strategy Game StarCraft II

#148

I didn't catch if they mentioned this in the interview, but what would have happened if they let AlphaStar play against other races, not Protoss only, would it be completely lost, unable to achieve anything?

The current version would flounder because it has never seen the other races. But that is not a fundamental limitation. All it would take to fix it is more training time.

Or they trained all the combinations but it was best at Protoss vs Protoss, so they publicized those results. But given the reaction to the first AlphaZero announcement and the later follow-up paper (summary: they cheated a bit but it's still incredibly strong) I would give them the benefit of the doubt here.

Re: AlphaStar: Mastering the Real-Time Strategy Game StarCraft II

#149
post #123

Earlier quoted context omitted.

For now. Give them another month. This is like AlphaGo vs Fan Hui all over again -- people knocked that accomplishment at the time because he was just a master, not one of the top players in the world. Well, not much longer, AlphaGo beat Lee Sedol, the best player in the world. The ceiling here is going to be incredibly high, much higher than the level of play that people are capable of, even when restricted to a sin…

Lee Sedol was not the best player anymore at that time (not saying it wasn't an impressive/important achievement, but overstating it doesn't help either - the "beat best human players part" came later in 2017).

I don't understand who's downvoting you, this is accurate. While AlphaGo/Zero improved quickly to superhuman play, we are just in this thread comparing timelines, so that is relevant.

Re: AlphaStar: Mastering the Real-Time Strategy Game StarCraft II

#150

Earlier quoted context omitted.

In the mass stalker battles, the AI APM exceeded 1000 a few times, and no doubt that most of that was precisely targeted. Whereas a human doing 500 APM micro is obviously going to be far more imprecise. I think a far more interesting limitation would be to cap APM at 150 or so, or to artificially limit action precision with some sort of virtual mouse that reduced accuracy as APM increased.

>I think a far more interesting limitation would be to cap APM at 150 or so, or to artificially limit action precision with some sort of virtual mouse that reduced accuracy as APM increased. IIRC OpenAI limits the reaction time to ~200ms when playing DoTA2. AI employing better strategies than humans will always be more interesting than AI that can out click humans.

Even the 200ms reaction time seemed overly slanted towards the AI. I don't think that is the actual reaction time of top pros, in the matches the AI played the human player would teleport in from complete invisibility and try to use an instant cast spell and the AI would have already teleported out. Yes the theoretically may have been constrained to a 200ms reaction time, but in practice the AI was playing at a superhuman level. Even with that advantage in fights, the human team still demolished the AI. Oh well, lots of things to learn still.
Post reply on HN