Live data from Hacker News

DeepMind StarCraft II Demonstration [video]

twitch.tv

101–110 of 124 posts

Re: DeepMind StarCraft II Demonstration [video]

#101
post #83
post #67

The live exhibition match definitely made Alphastar look like a machine making the decisions, not a super smart being. The micro was obviously impressive but that should also be the easiest part to master. I share the sentiment that DeepMind is hosting big events to paint a very one-sided picture of man vs machine, so this last win of mana feels oddly satisfying.

It is a little disappointing that Mana's win was achieved in part by simply exploiting Alphastar's poor response to the immortal drops by doing it over and over again. In contrast, Lee Sedol's win versus AlphaGo involved profound strategy and a particularly inspired "divine move" that humans get to brag about.

Lee Sedol "divine move" actually doesn't work, even human top players would have punished it.

It actually was just a really weird mistake in AlphaGo

Re: DeepMind StarCraft II Demonstration [video]

#102
post #83

Earlier quoted context omitted.

It is a little disappointing that Mana's win was achieved in part by simply exploiting Alphastar's poor response to the immortal drops by doing it over and over again. In contrast, Lee Sedol's win versus AlphaGo involved profound strategy and a particularly inspired "divine move" that humans get to brag about.

Lee Sedol "divine move" actually doesn't work, even human top players would have punished it. It actually was just a really weird mistake in AlphaGo

Thanks for clarifying that, never really got that that deep into Go.

Re: DeepMind StarCraft II Demonstration [video]

#103
post #83

Earlier quoted context omitted.

It is a little disappointing that Mana's win was achieved in part by simply exploiting Alphastar's poor response to the immortal drops by doing it over and over again. In contrast, Lee Sedol's win versus AlphaGo involved profound strategy and a particularly inspired "divine move" that humans get to brag about.

Lee Sedol "divine move" actually doesn't work, even human top players would have punished it. It actually was just a really weird mistake in AlphaGo

As they say... there are no good moves in Go/chess/zero-sum games etc. There are only bad moves.

That doesn't prevent professional players such as Gu Li 9p from describing it as a divine move though. And it does feel good for Team Humanity.

Re: DeepMind StarCraft II Demonstration [video]

#104
post #100

Earlier quoted context omitted.

They limited the APM to 180 for AlphaStar's input I think. An average professional player consistently plays at 200-300 APM with peak levels generally going to 350-400 APM during engagements. To me, it seems like AlphaStar is actually at a disadvantage in its ability to micro, and has to make up for this in its macro level strategy.

Most human APM is wasted on "spam clicks". The effective APM, or EPM, is probably quite a bit higher for AlphaStar than for human pros.

TLO and MaNa are in the top 0.01%. They are the best of the best, and then more. The numbers I stated was for the average professional.

When you have geniuses like TLO and MaNa, average APM is more on the level of 300-400 consistent and 500+ in engagements. You're definitely correct that APM is not 100% effectively utilized. I would say maybe around ~40% of actions could generally be considered spam clicks. That still puts AlphaStar at a definite disadvantage in its ability to enact micro level strategy.

Re: DeepMind StarCraft II Demonstration [video]

#105
post #98
post #92

Earlier quoted context omitted.

My understanding is that Alphastar is trained on reinforcement learning. I suspect this is unsupervised so it would have learned this behaviour independently without pro replays.

Per DeepMind's blog post[1], an agent was initially trained via supervised learning on pro matches. Then, the agent was forked repeatedly, as the population of agents learned via tournament-style self-play. So, while initial strategies could have been seeded by pro play styles, the final models were the result of models learning from games with other models. [1] https://deepmind.com/blog/alphastar-mastering-real-time…

Meaning that the final models evolved by playing against each other and they all started by using pro strategies, so, again, to me it seems kind of obvious that they would end up using pro strategies and in the best case just try to make them better

Re: DeepMind StarCraft II Demonstration [video]

#106
post #45

Very impressive... but it seems like the AI relies entirely on abusing blink stalkers which with perfect micro is basically impossible to counter. It is no surprise it can crush pros when it has perfect timing and zero mistakes in using these units. I think the coolest thing is how the play of AI mirrors a similar style to how pros have developed (Micro harrasing early, early expansions, a very good understanding of…

It had a variety of strategies, not always making a lot of stalkers. And humans can do some really impressive blink micro too up to a certain number of units. So that's one aspect of it but isn't the main strength.

Those were not one and the same agent. If you look at the figure they released on their website the given agent would probably always have gone for a lot of blink stalkers.

Re: DeepMind StarCraft II Demonstration [video]

#107

Earlier quoted context omitted.

It had a variety of strategies, not always making a lot of stalkers. And humans can do some really impressive blink micro too up to a certain number of units. So that's one aspect of it but isn't the main strength.

Those were not one and the same agent. If you look at the figure they released on their website the given agent would probably always have gone for a lot of blink stalkers.

I was amused to see on their website that one of the top-rated agents we didn't get to see built almost nothing but Void Rays.

Re: DeepMind StarCraft II Demonstration [video]

#108
post #77

Earlier quoted context omitted.

As the commentators mentioned, it's no use building units to counter your enemy's army (Immortals over Stalkers) when the enemy can control their army so much more effectively. I have to wonder if future competitive games will need to take into account the abilities of reinforcement learning algorithms when releasing balance patches.

But there IS use building units to counter your enemy's army. In the last live match when Mana won, his immortal archon zealot composition was what sealed the deal in the end.

And his far superior positioning.

Re: DeepMind StarCraft II Demonstration [video]

#109
post #89
post #80

Earlier quoted context omitted.

>There are limitations they put on the AI to try ti restrict to human levels. Such as having an action counter. Which is exactly why StarCraft is not a very good game to test AI on. It's absurd to put arbitrary limitation on something to make the game "fair" and then pat yourself on the back simply because the algorithm won. If it can already win through pure micromanagement, why There are tons of strategy games whic…

> All turn-based games, for example. You mean like chess? And go? I think turning to a real time game with a complex rule set after showing they mastered turn based games with simple rule sets was very sensible. > Or real-time games where building stuff is more important than combat. Can you name one that is played professionally (important for balance and comparison to humans) where this is more true than starcraft…

>You mean like chess? And go?

What a clever reply. No, like the ones I named in the post above.

>Can you name one that is played professionally (important for balance and comparison to humans)

If you pause to think about it, things that make for a "good" pro scene are exactly the things that make computer games less than impressive for innovative AI research.

Re: DeepMind StarCraft II Demonstration [video]

#110
post #100

Earlier quoted context omitted.

Most human APM is wasted on "spam clicks". The effective APM, or EPM, is probably quite a bit higher for AlphaStar than for human pros.

TLO and MaNa are in the top 0.01%. They are the best of the best, and then more. The numbers I stated was for the average professional . When you have geniuses like TLO and MaNa, average APM is more on the level of 300-400 consistent and 500+ in engagements. You're definitely correct that APM is not 100% effectively utilized. I would say maybe around ~40% of actions could generally be considered spam clicks. That sti…

Two additional points, aside from APM effectiveness:

- The APM constraint by AlphaStar seems to be average APM (or some variation thereof), while it reached 900+ APM in battles -- again with 100% effectiveness. That's insane and definitely beyond human capacity.

- It didn't waste APM in the early game (when humans are warming up and spamming APM), further leaving slack for when engagements came.

I mean, when gauging practical applications, we shouldn't be too concerned about good reflexes and superhuman effectiveness; however, part of the point here was to gauge other aspects of cognition more related to strategy, large decision spaces, etc; and in this setting it makes sense to constrain those non-strategic (more mechanical/tactical aspects). I'm sure e.g. constraining humans to <300 APM for example would change player style/effectiveness that much, while e.g. the blink stalker micro displayed in Game 4 vs Mana was really reliant on massive parallel single unit/group control (blinking every single low health stalker individually, engaging from multiple sides simultaneously, etc).

Post reply on HN