Live data from Hacker News

AlphaGo beats Lee Sedol 3-0 [video]

youtube.com

291–300 of 428 posts

Re: AlphaGo beats Lee Sedol 3-0 [video]

#291
post #210

Earlier quoted context omitted.

It seems like in all three games, AlphaGo plays a thick move that looks a bit baffling to the commentators, but then miraculously those moves become hugely profitable somewhat later on in the game.

Is it possible that all of AlphaGo's strength is in these baffling moves? If humans played enough games against AlphaGo and discovered how to counter the baffling moves, is it plausible that AlphaGo's strength would be lost?

Watching the game and the commentary there's an eerie sensation that AlphaGo is just going along playing the petty local games with Lee Sedol. It's very anthropocentric, but you could imagine AlphaGo saying "I could have crushed you from the start ... but I'll just play along and do that one move to give you the illusion that it was close".

But maybe it could have 40 of those moves to play an incredibly confusing game for humans. I think giving the machine a big handicap against pros might reveal it's true colours ...

Re: AlphaGo beats Lee Sedol 3-0 [video]

#292
post #43

Some professionals labeled some AlphaGo moves as being unoptimal or slow. In reality, Alpha Go doesn't try to maximize its score, only its probability of winning.

From watching it I'm almost inclined to say it maximizes its chances of not losing over necessarily winning.

Your comment reminds me of the Star Trek TNG episode Peak Performance. Data wins the rematch by playing to stalemate.

https://en.wikipedia.org/wiki/Peak_Performance_(Star_Trek:_T...

Re: AlphaGo beats Lee Sedol 3-0 [video]

#293
post #198

Earlier quoted context omitted.

>> At this point it seems likely that Sedol is actually far outclassed by a superhuman player. I still don't agree that this is the case, and I don't care what a thousand Google-hyped press releases say, beating the best human player in anything is not "superhuman" and "superhuman" performance has not been achieved by anything yet. [1] Why do I think so? Two reasons. One, because you can be entirely human and still b…

Your pocket calculator is superhuman at arithmetic. Alphago is superhuman at go. It's not a big deal.

Superhuman has a very specific definition in the field of game playing algorithms. It is the case when the algorithm can always beat all humans. AlphaGo winning 5-0 against the number 5 ranked human (Lee) would give only a small indication that it is superhuman. Regarding streaks, evenly matched humans can go 5-0 an expected (0.5)^5 or 3.125 percent of the time. So not particularly rare. If it loses even one game then it is not yet superhuman.

If top humans get beat 5-0 with significant handicaps then it is likely AlphaGo is superhuman. However, it is expensive to run AlphaGo so it is unlikely that we will know the true strength of AlphaGo for a while (more challenges) or until hardware catches up.

Update: typos and clarifications

Re: AlphaGo beats Lee Sedol 3-0 [video]

#294

Earlier quoted context omitted.

What is our purpose if computers can do everything better than us? It feels like computers have taken one aspect of humanness: logic. Computers could do arithmetic, do algebra, play chess, and now they can play go. It hurts because logic is usually thought to be one of the highest of human characteristics. Yes computers might never be able to replicate emotion, but even dogs have that. There's still some aspects we h…

Computers will be able to compose symphonies very soon. If DeepMind started working on this problem, I am sure that they would succeed. At least, we would have some innovative mashups of Beethoven, Mozart and Tchaikovsky. But training a powerful AI on a massive dataset of all popular and classical music should produce some extraordinary results. Especially if the dataset was given as MIDI with separate instrument tra…

Composing an amazing symphony is probably about as hard as being the best go player in the world. But I think we're much further away than you think.

AlphaGo needed a training set of perhaps a billion games to be as good as it is. The dataset of master Go games is perhaps a million games. So AlphaGo played at tons of games against a half-trained version of itself to reach the billion game mark.

This doesn't work for songs, because there's no one to tell AlphaBach whether any of the billion symphonies it makes are any good. AlphaGo can just look at the rules and see if its move lead to a win, but there's no automatic evaluation function for music.

Perhaps the Matrix wasn't using the humans for power, but rather the computers wanted to get good at writing music, so they gave each human in it slightly different music and watched their emotional responses.

Re: AlphaGo beats Lee Sedol 3-0 [video]

#295

Earlier quoted context omitted.

Self-preservation falls out of almost any other goal you give an AGI. If I program my AGI with the goal of making my startup succeed, and the AGI thinks it can help, then me shutting it off is a potential threat to my startup's success. So of course it will try to prevent that the same way it would try to prevent any other threat to my startup's success. World domination is a similar situation. For any goal you give…

It has to be aware that it can be shut down and have the capacity to prevent that. AlphaGo doesn't know it can be shut down and therefore couldn't "care" less--even if it was shut down in the middle of a game.

Yes, I agree. My point is that as soon as you are giving your AI "real world" problems, where the AI itself is a stone on its internal go board, you have to start worrying about these issues.

Re: AlphaGo beats Lee Sedol 3-0 [video]

#296
post #79

My (long) commentary here: https://www.facebook.com/yudkowsky/posts/10154018209759228 Sample: At this point it seems likely that Sedol is actually far outclassed by a superhuman player. The suspicion is that since AlphaGo plays purely for probability of long-term victory rather than playing for points, the fight against Sedol generates boards that can falsely appear to a human to be balanced even as Sedol's probabili…

> AI alignment theory

What is this?

Re: AlphaGo beats Lee Sedol 3-0 [video]

#297

Earlier quoted context omitted.

I find it funny that whenever there's a case of computers being able to do something that they couldn't before, whether it's drive or beat humans at Go, the goalpost on what is "true AI" shifts to be something that computers can't do yet. So let me ask you this. What would you consider to be "true AI"? At what point are you willing to say, "Okay, that's it, computers are just plain smarter than we are?" Because, fran…

> So let me ask you this. What would you consider to be "true AI"? At what point are you willing to say, "Okay, that's it, computers are just plain smarter than we are?" Because, frankly, it seems to me that that day is getting closer and closer. Alan Turing would give the system the "Turing test". If a computer can fool a human into thinking it's a human, then it is true AI, according to Turing. I think that's a pre…

Brilliant answer - independent goal setting is a really interesting alternative phrasing of "soul" or "spirit" or "individuality", because unlike those, it can be easily observed or tested. Great writeup, thanks for making me think.

Re: AlphaGo beats Lee Sedol 3-0 [video]

#299
post #79

My (long) commentary here: https://www.facebook.com/yudkowsky/posts/10154018209759228 Sample: At this point it seems likely that Sedol is actually far outclassed by a superhuman player. The suspicion is that since AlphaGo plays purely for probability of long-term victory rather than playing for points, the fight against Sedol generates boards that can falsely appear to a human to be balanced even as Sedol's probabili…

>> But Go is rich enough to demonstrate strong cognitive uncontainability on a small scale. In a rich and complicated domain whose rules aren't fully known, we should expect even more magic from superhuman reasoning - solutions that are better than the best solution we could imagine, operating by causal pathways we wouldn't be able to foresee even if we were told the AI's exact actions.

Hang on. Where are we going to find this magic agent that can do "even more" in "a rich and complicated domain whose rules aren't fully known" (the real world, as opposed to a Go board) than in a game of Go? What are we going to train such a learner with, if we ourselves don't fully know the rules of the domain, as you point out?

Even if a learner somehow magically found a superhuman path to perfect reasoning which is unavailable to us, entirely by accident and purely on its own, why would we select it from other trained learners to keep and foster further, if we think it's actually pretty dumb, rather than magically smart?

You're saying at some point that a fantastic paperclip maximiser might achieve superhuman intelligence and then lie in waiting, poised to turn us all into paperclips only when it knew it was safe to make its move, basically. But, how is it going to become that smart in the first place? It has to be smart enough to know that it must bide its time, but dumb enough that its time hasn't come yet. Sounds like a bit of an impossible double-bind there.

Just because sufficiently advanced AI may look like magic at first, it doesn't mean we should start all believing in magic because it may just be sufficiently advanced AI in disguise.

Re: AlphaGo beats Lee Sedol 3-0 [video]

#300

Earlier quoted context omitted.

That would be amazing if it could achieve the same levels (or higher) without the bootstrapping. The niggling thought in my mind was that AlphaGo's strength is built on human strength.

I'm pretty sure that the reinforcement learning algorithm they are using is guaranteed to converge. It just takes a very long time to train, and using human games probably sped it up.

As far as I know, using neural networks for function approximation destroys the various convergence guarantees available. NNs can easily diverge and have catastrophic forgetting, and this is one of the things that made them challenging to use in RL applications despite their power, and why one needs patches like experience replay and freezing the networks.
Post reply on HN