Live data from Hacker News

AlphaStar: Mastering the Real-Time Strategy Game StarCraft II

deepmind.com

201–210 of 459 posts

Re: AlphaStar: Mastering the Real-Time Strategy Game StarCraft II

#201
post #46

This is really impressive, I didn't expect starcraft to be played this well by a machine learning based AI. I'm excited to read the paper when it comes out! That said, I'm not sure I agree that it was winning mainly due to better decision making. For context, I've been ranked in the top 0.1% of players and beaten pros in Starcraft 2, and also work as a machine learning engineer. The stalker micro in particular looked…

In the mass stalker battles, the AI APM exceeded 1000 a few times, and no doubt that most of that was precisely targeted. Whereas a human doing 500 APM micro is obviously going to be far more imprecise. I think a far more interesting limitation would be to cap APM at 150 or so, or to artificially limit action precision with some sort of virtual mouse that reduced accuracy as APM increased.

There would be an entire new dimension of decision making, in addition to good macro, where you have to prioritize actions. Will be interesting to see.

Re: AlphaStar: Mastering the Real-Time Strategy Game StarCraft II

#202

Earlier quoted context omitted.

While that would be amazing if true, I'm pretty sure if you take away the stalker blink micro AlphaStar loses hands down to humans. This isn't taking away from Deepmind's victory at all, but I think micro was what made the AI come out ahead in this one. In many of the games, Mana had much better macro only to lose to blink stalkers.

You play the game as it's written. Come back with another version of StarCraft that isn't so micro-intensive and we can see how the AI does on that. Chess and Go don't have any form of micro and AIs are nevertheless dominant there. I'd say, give AI development another year and I wouldn't expect there to be any kind of game, in any genre, that humans can beat AIs at. Whether it's Chess, Go, other classical board games…

I sure hope so---then I could a 4X AI that was worth a damn.

Re: AlphaStar: Mastering the Real-Time Strategy Game StarCraft II

#203
post #150

Earlier quoted context omitted.

>I think a far more interesting limitation would be to cap APM at 150 or so, or to artificially limit action precision with some sort of virtual mouse that reduced accuracy as APM increased. IIRC OpenAI limits the reaction time to ~200ms when playing DoTA2. AI employing better strategies than humans will always be more interesting than AI that can out click humans.

Even the 200ms reaction time seemed overly slanted towards the AI. I don't think that is the actual reaction time of top pros, in the matches the AI played the human player would teleport in from complete invisibility and try to use an instant cast spell and the AI would have already teleported out. Yes the theoretically may have been constrained to a 200ms reaction time, but in practice the AI was playing at a super…

I think OpenAI would have been by lots of humans, but they decided to train it with 5 unlimited, invulnerable couriers. (until the TI showmatches, in which they were beaten easily.)

Re: AlphaStar: Mastering the Real-Time Strategy Game StarCraft II

#204
post #158

Earlier quoted context omitted.

"Yes but X has a tiny problem space compared to something like Y. People want to see an AI win because it's smart, not because it crunches numbers." 1980: X = Tic-tac-toe, Y = Chequers 1990: X = Chequers, Y = Chess 2000: X = Chess, Y = Go 2019: X = Go, Y = StarCraft 2030: X = Any video game, Y = ???

Is a AI that wins at Starcraft only because it has crazy high APM really going to help get to the next X? We could have built that 10 years ago. All it proves is that computers have faster reflexes then humans. That won’t help them become problem solvers for the future.

You seem to forget the way it learned to play every part of the game (not just micro fights). That is, not by having any developer code any rules, but simply by "looking" and "playing".

That's the great accomplishment and nothing like that could have been done 10 years ago.

Re: AlphaStar: Mastering the Real-Time Strategy Game StarCraft II

#205
post #146

Earlier quoted context omitted.

The results are obviously impressive, but even then there is a lot of work to do as far as learning efficiency goes: "The AlphaStar league was run for 14 days, using 16 TPUs for each agent. During training, each agent experienced up to 200 years of real-time StarCraft play. " MaNa probably played less than 2-3 years of Starcraft in his whole life (by that I mean 24hr x 365d x 3), and was learning with a much less foc…

Another way to think about it is that a human brain is mostly doing transfer-learning, on top of a 99%-baked deep net that was wired up during foetal development from our DNA, where that DNA-persisted model has "seen" hundreds of millions of years of training data. Humans don't have to learn to process, recognize, and classify objects in visual sense-data, for example. We can do that from the moment we're born, becau…

Perhaps a nit, but still fascinating: the human visual cortex finishes developing after birth. A newborn can't really distinguish between objects. The ability to differentiate, focus on and track objects is developed over the course of several months.

Re: AlphaStar: Mastering the Real-Time Strategy Game StarCraft II

#206
post #146

Earlier quoted context omitted.

The results are obviously impressive, but even then there is a lot of work to do as far as learning efficiency goes: "The AlphaStar league was run for 14 days, using 16 TPUs for each agent. During training, each agent experienced up to 200 years of real-time StarCraft play. " MaNa probably played less than 2-3 years of Starcraft in his whole life (by that I mean 24hr x 365d x 3), and was learning with a much less foc…

Another way to think about it is that a human brain is mostly doing transfer-learning, on top of a 99%-baked deep net that was wired up during foetal development from our DNA, where that DNA-persisted model has "seen" hundreds of millions of years of training data. Humans don't have to learn to process, recognize, and classify objects in visual sense-data, for example. We can do that from the moment we're born, becau…

That's not how any of this works. We do not have "millions of years" of information encoded into DNA. DNA doesn't store that much data. In fact, it's about 1.6 gigabytes only! And most of that information is basically a ruleset for growing proteins which become our body.

All the stuff we've learned about games and so on have come from our current lifetime. I don't have caveman memory for how to fight a tiger.

Re: AlphaStar: Mastering the Real-Time Strategy Game StarCraft II

#207
post #164
post #146

Earlier quoted context omitted.

Another way to think about it is that a human brain is mostly doing transfer-learning, on top of a 99%-baked deep net that was wired up during foetal development from our DNA, where that DNA-persisted model has "seen" hundreds of millions of years of training data. Humans don't have to learn to process, recognize, and classify objects in visual sense-data, for example. We can do that from the moment we're born, becau…

This is a widely underappreciated fact when it comes to comes to comparing the 'training experience' of humans versus bots. And it extends far beyond processing 'sense data' - A human likely has some level of understanding of how the game works based on experience from other games it has played and from 'real life' - we know almost instinctively that 'high ground' is likely to give a combat advantage without having t…

All of our knowledge of how to play games and so on has come from our current lifetime. We do not have a "genetic memory" that means we have learnings from cavemen or some other such nonsense. Our DNA contains instructions on how to grow a human, it's not a mega hard drive with millions of years of collective memory.

If a 19 year old is good at Starcraft, he's good at Starcraft because he spent two or three years playing a shit load of Starcraft and we are much more efficient at learning higher level strategies than AI are. These AI agents nead to try damn near every possibility to adjust their weightings for various actions. Humans understand pretty much the first time when something goes wrong, oh better not do that OR similar things again.

It's incredibly impressive that a given human can become GM level at Starcraft within a few years and to take an AI to that level takes 200 years of training, as well as an inhuman reaction time, perfect micro/clicking, etc. It shows how amazing our learning skills are.

Re: AlphaStar: Mastering the Real-Time Strategy Game StarCraft II

#209
post #46

This is really impressive, I didn't expect starcraft to be played this well by a machine learning based AI. I'm excited to read the paper when it comes out! That said, I'm not sure I agree that it was winning mainly due to better decision making. For context, I've been ranked in the top 0.1% of players and beaten pros in Starcraft 2, and also work as a machine learning engineer. The stalker micro in particular looked…

I would agree with that. If you take a look at the exhibition match replay, there's some cases where it makes objectively suboptimal decisions. We couldn't see this during the live stream, but the double immortal warp prism caused AlphaStar to bring back its entire army from across the map, when a few units at home would have been enough to defend. It even kept trying to blink its stalkers to a place where the warp prism couldn't be reached. Perhaps this version with the limited viewpoint hadn't been trained with enough games?

Re: AlphaStar: Mastering the Real-Time Strategy Game StarCraft II

#210
post #56

Earlier quoted context omitted.

In the showmatched they made the computer have to look at a regular screen to control, the stalker micro was much less impressive - and mana won.

For now. Give them another month. This is like AlphaGo vs Fan Hui all over again -- people knocked that accomplishment at the time because he was just a master, not one of the top players in the world. Well, not much longer, AlphaGo beat Lee Sedol, the best player in the world. The ceiling here is going to be incredibly high, much higher than the level of play that people are capable of, even when restricted to a sin…

What kind of evidence is going into this analogical reasoning? Do we also extrapolate similarly for other things? We went to the Moon in 1960s. Was Mars a month, or a year, or a decade away? Then we sent robots to Mars. Did we yet send any robots to Alpha Centauri?

Different problems have different difficulties. Solving simple problems quickly doesn't mean we'd also be able to just as easily solve the hard problems. Often the comparably simpler problems have the best reward/effort ratio and thus make quick progress, which doesn't need to be the case for hard problems.

Post reply on HN