Live data from Hacker News

Generally capable agents emerge from open-ended play

deepmind.com

51–55 of 55 posts

Re: Generally capable agents emerge from open-ended play

#51

Earlier quoted context omitted.

the main question of course being, aren't we anthropomorphizing ourselves too much ?

When people say this kind of stuff, I wonder whether there might not be philosophical zombies among us.

[ ] To prove that you are human please describe how you are observably different from embodied, general, adaptive agents in 200 words.

Re: Generally capable agents emerge from open-ended play

#52

Earlier quoted context omitted.

This completely misses the point. The problem isn't that people do worse, but that automating warfare changes the cost function. Technologically superior and determined enemies can still be dissuaded by incurring significant numbers of casualties or losses to their infrastructure. A small number of casualties stiffens resolve, a large number of casualties or a never ending trickle of them eventually breaks or wears i…

You left out the bit where the various superpowers inevitably have to try using the shiny new technology against their rivals because it's never been tried so we can't be certain it's a bad idea. That's the bit that worries me the most - I'd rather do without a rerun of World War 1. > it's more important to assess the impact of this emerging force multiplier and develop countermeasures What is there to do other than…

There's a whole literature on the logic (and meta-logic) of deterrence called power transition theory that is worth looking into, as it sheds a lot of light on the unpleasant topic of nuclear deterrence and how that works.

In a more general sense, the solution to an elevated attack is not always a retaliatory attack, but perhaps a better defense that neutralizes it. Helmets can be used as weapons, but their primary purpose is to make weapons less effective and change the strategic calculus - now the enemy gets lesser results for the same effort, and either gives up or tires out and can be defeated with a smaller retaliation. In general, defense is thought to be somewhat stronger than offense, which is why surprise is so important. Technological edges tend to be negated over time.

Deeply understanding this takes a long time and a lot of study. Military science is a difficult but interesting subject, and tips over into systems theory.

Re: Generally capable agents emerge from open-ended play

#53

Earlier quoted context omitted.

Don't worry about it. We can already do far worse using either humans or land mines or nuclear bombs or whatever. People are the real danger.

This completely misses the point. The problem isn't that people do worse, but that automating warfare changes the cost function. Technologically superior and determined enemies can still be dissuaded by incurring significant numbers of casualties or losses to their infrastructure. A small number of casualties stiffens resolve, a large number of casualties or a never ending trickle of them eventually breaks or wears i…

My point is that it's already automated. Those things I mentioned are already not hand-to-hand combat and don't have human soldiers risking their lives. So however scary some new automated weapon is, it should be no worse than existing automated weapons. What makes a robot soldier worse than a cruise missile or land mines or a bomber aircraft or a remote piloted drone? All those things can already be used by technologically superior enemies without incurring casualties themselves.

You mention attacks against a technologically advanced power (does an "enemy" become a "power" when it's a friend?), but obviously those powers will find ways to defend against them. Maybe it's just in the form of slightly more advanced "washing machines".

This fear thinking seems to come from assuming no secondary advancements occur. Suddenly robot soldiers are cheaply available and nobody develops any defense against them, either political or technological.

Re: Generally capable agents emerge from open-ended play

#54

Earlier quoted context omitted.

It's proven that if you want polynomial sample complexity in the size of the state space you need directed exploration. The algorithms 'flail' because they are initialized with random policies.

Couldn't any directed exploration also be produced as the result of random flailing? Or is this average sample complexity?

An example of a bad case for random exploration would be a narrow ridge where you die if you fall off and you only receive reward if you get to the end.

So it's a worst case result with respect to the MDP, but expected time/high probability with wrt random chance.

Post reply on HN