Live data from Hacker News

Mastering Chess and Shogi by Self-Play with General Reinforcement Learning

arxiv.org

51–60 of 282 posts

Re: Mastering Chess and Shogi by Self-Play with General Reinforcement Learning

#51

One impressive statistic from the paper: AlphaZero analyzes 80,000 chess positions per second, while Stockfish looks at 70,000,000. Seventy million, three orders of magnitude higher. Yet AG0 beats Stockfish half the time as White and never loses with either color. A stunning demonstration of generality indeed.

So ... what if you combined Stockfish and AG0, and let AG0 explore 70M positions instead of 80K? Would it improve even faster?

Re: Mastering Chess and Shogi by Self-Play with General Reinforcement Learning

#52
post #30

Earlier quoted context omitted.

> So in a way, the real triumph of the AlphaGo series was the TPU and GPU army. Eh. It's still searching many fewer positions than Stockfish is.

Right right, but my comparison was between giraffe and AlphaGo , not neural networks and Stockfish.

But it looks like AlphaGo is searching fewer positions per second than Giraffe did.

AlphaZero evaluates 80K positions per second, according to this paper, and the Giraffe paper says that Giraffe averaged 258570 evaluations per second when running STS.

While we can't directly compare the computer power, this implies that AZ has learned a better representation.

Re: Mastering Chess and Shogi by Self-Play with General Reinforcement Learning

#53

While this sounds impressive, I'll believe it when AlphaZero wins TCEC.

It beat the winner of TCEC-2016, Stockfish, with a record of 28-72-0. That's zero losses.

They didn't demonstrate that AlphaGo Zero can beat Stockfish in a fair contest: i.e. take the amount of money they spent on Stockfish's CPU and RAM, buy a commodity GPU for AlphaGo and then see.

Re: Mastering Chess and Shogi by Self-Play with General Reinforcement Learning

#55
post #10

This is an incredible demonstration that the AG Zero expert iteration method is a general method. If you go back to the discussions of AG Zero lo a month ago, there was a lot of skepticism that NNs would ever challenge Stockfish et al - they are just too good, too close to perfection, and chess not well suited for MCTS and NNs. Well, it turns out that AG Zero doesn't work as well in chess: it works better as it only…

>it only takes 4 hours of training to beat Stockfish

In that time I figure they used the equivalent of about 1000 cpu-years. Imagine the things we'll be able to achieve as we can do more and more computation in less and less time.

Re: Mastering Chess and Shogi by Self-Play with General Reinforcement Learning

#56
post #38
post #6

Two things to note: 1) Alpha Zero beats AlphaGo Zero and AlphaGo Lee and starts tabla rasa 2) "Shogi is a significantly harder game, in terms of computational complexity, than chess (2, 14): it is played on a larger board, and any captured opponent piece changes sides and may subsequently be dropped anywhere on the board. The strongest shogi programs, such as Computer Shogi Association (CSA) world-champion Elmo, have…

Shogi is a fun game, it always feels a little sad that it doesn't get more exposure outside of Japan (and my understanding is that, by and large, in Japan it is considered an "old persons" game) Because captured pieces change sides, there is less of an "endgame" scenario, and as a beginner (like me) it is very easy to put too many captured pieces back into play, which makes it hard to defend everything and essentiall…

It briefly became popular in the otaku culture from an anime called Hunter X Hunter.

Re: Mastering Chess and Shogi by Self-Play with General Reinforcement Learning

#57
post #10

This is an incredible demonstration that the AG Zero expert iteration method is a general method. If you go back to the discussions of AG Zero lo a month ago, there was a lot of skepticism that NNs would ever challenge Stockfish et al - they are just too good, too close to perfection, and chess not well suited for MCTS and NNs. Well, it turns out that AG Zero doesn't work as well in chess: it works better as it only…

See the thing is though, Giraffe's evaluation actually was better than Stockfish's evaluation function, but it took much longer, and thus wasn't able to search as deep as Stockfish et al. So in a way, the real triumph of the AlphaGo series was the TPU and GPU army.

[deleted]

Re: Mastering Chess and Shogi by Self-Play with General Reinforcement Learning

#58
post #54

Did Magnus play against this? Is there a way we can see the game?

No he didn't play it. As far as I know, computers are already far ahead of humans in chess, so a further progress in this wouldn't really make a difference.

Re: Mastering Chess and Shogi by Self-Play with General Reinforcement Learning

#59

One impressive statistic from the paper: AlphaZero analyzes 80,000 chess positions per second, while Stockfish looks at 70,000,000. Seventy million, three orders of magnitude higher. Yet AG0 beats Stockfish half the time as White and never loses with either color. A stunning demonstration of generality indeed.

So ... what if you combined Stockfish and AG0, and let AG0 explore 70M positions instead of 80K? Would it improve even faster?

The issue is you can’t evaluate positions that fast in AlphaZero (currently).

Re: Mastering Chess and Shogi by Self-Play with General Reinforcement Learning

#60

While this sounds impressive, I'll believe it when AlphaZero wins TCEC.

It beat the winner of TCEC-2016, Stockfish, with a record of 28-72-0. That's zero losses.

On completely different hardware.
Post reply on HN