Live data from Hacker News

Generally capable agents emerge from open-ended play

deepmind.com

31–40 of 55 posts

Re: Generally capable agents emerge from open-ended play

#31
post #7

Agents trained in simulation like this often flail about seemingly randomly, and when they achieve their goals it seems almost accidental. Rather than this being some kind of limitation of the learning algorithm, I think it might be the optimal strategy, and humans would behave that way too if there was no such thing as fatigue or pain. If we want agents to behave more realistically and move with more apparent intent…

Watching my 3-year-old, I suspect humans do behave this way, and without regards to pain or fatigue. He'll hit himself on the head with duplos or bang his head against the wall just to see what happens. The pain is just one more signal for reinforcement learning. I recall a paper in a non-CS journal (psych or neuroscience) that posited that the optimal way to gain large quantities of information about an uncertain en…

> I suspect humans do behave this way

Suddenly I recall one author of SICP saying that programming today is about poking at libraries.

Re: Generally capable agents emerge from open-ended play

#32
post #28

Earlier quoted context omitted.

you seem to reduce it down to pain. how is your notion of "pain" any different from not scoring well and being moved away from?

Some rewards can be worth certain penalties (pain) if within certain thresholds. Exhaustion fatigue can also play in. You can of course boil them down to a single number, it just produces less nuanced types of decisions/operations, as it can’t differentiate between a cheap, painful, but bountiful choice and a expensive, no pain, mediocre choice.

I think there needs to also be a concept of death—penalties from which recovery is impossible. Nonergodicity seems to be a requirement for the development of antifragility.

Re: Generally capable agents emerge from open-ended play

#33

- "Tag" - shoots other player - "Capture the flag" - shoots other player - "Hide and seek" - shoots other player As colorful as this world is, these capabilities terrify me because they're obviously going to be used as powerful weapons of war. It is a tiny technological leap to install this learning into a Boston Robotics Spot attached to a firearm. I'm pro-tech, pro-crypto, pro-ml and these videos fill me with dread…

Don't worry about it. We can already do far worse using either humans or land mines or nuclear bombs or whatever. People are the real danger.

Re: Generally capable agents emerge from open-ended play

#34
post #7

Agents trained in simulation like this often flail about seemingly randomly, and when they achieve their goals it seems almost accidental. Rather than this being some kind of limitation of the learning algorithm, I think it might be the optimal strategy, and humans would behave that way too if there was no such thing as fatigue or pain. If we want agents to behave more realistically and move with more apparent intent…

Great question.

One thing to imagine might be a pain or cost curve that has multiple minima and maxima. Pain can be either a demotivator or motivator in different contexts. Pain qua fatique might indicate that a reward slope exists for more conditioning. (Edit: that is on a static or realist view; the terrain of reward and pain is probably constantly changing.)

Random tangential data point: Some animals (chickens?) can learn superstitious behavior; the first action of theirs that happens to correlate with a reward can result in one-shot learning or something.

Re: Generally capable agents emerge from open-ended play

#35
post #7

Agents trained in simulation like this often flail about seemingly randomly, and when they achieve their goals it seems almost accidental. Rather than this being some kind of limitation of the learning algorithm, I think it might be the optimal strategy, and humans would behave that way too if there was no such thing as fatigue or pain. If we want agents to behave more realistically and move with more apparent intent…

"Throwing shit against the wall and seeing what works" is a widely used strategy in all walks of life. If you ever go to a doctor for a problem, and their first solution doesn't work, you're in for a wild ride of an eduction.

Re: Generally capable agents emerge from open-ended play

#37
post #7

Agents trained in simulation like this often flail about seemingly randomly, and when they achieve their goals it seems almost accidental. Rather than this being some kind of limitation of the learning algorithm, I think it might be the optimal strategy, and humans would behave that way too if there was no such thing as fatigue or pain. If we want agents to behave more realistically and move with more apparent intent…

Watching my 3-year-old, I suspect humans do behave this way, and without regards to pain or fatigue. He'll hit himself on the head with duplos or bang his head against the wall just to see what happens. The pain is just one more signal for reinforcement learning. I recall a paper in a non-CS journal (psych or neuroscience) that posited that the optimal way to gain large quantities of information about an uncertain en…

Have three year old and I would support your description for how kids grow in the early months and years.

Now though learning from stories and observation is a thing too. Unsure exactly when that started.

Of course, attention/awareness isn't great and playing with a balloon precludes caution around a fireplace... but it is understandable at this age.

Re: Generally capable agents emerge from open-ended play

#38
post #7

Agents trained in simulation like this often flail about seemingly randomly, and when they achieve their goals it seems almost accidental. Rather than this being some kind of limitation of the learning algorithm, I think it might be the optimal strategy, and humans would behave that way too if there was no such thing as fatigue or pain. If we want agents to behave more realistically and move with more apparent intent…

I've been thinking about an "energy" idea, where actions consume energy but completing goals replenishes it

Re: Generally capable agents emerge from open-ended play

#39

- "Tag" - shoots other player - "Capture the flag" - shoots other player - "Hide and seek" - shoots other player As colorful as this world is, these capabilities terrify me because they're obviously going to be used as powerful weapons of war. It is a tiny technological leap to install this learning into a Boston Robotics Spot attached to a firearm. I'm pro-tech, pro-crypto, pro-ml and these videos fill me with dread…

Greetings professor Falken. Would you like to play a game?

-Checkers

-Chess

-Poker

-Backgammon

-Falkens Maze

-Fighter combat

-Desert warfare

-Theaterwide biotoxic and chemical warfare

-Global thermonuclear war

Re: Generally capable agents emerge from open-ended play

#40
It's interesting to think that David Silver and team must have known about this when they published their paper "Reward is Enough" in May. You have to wonder what else they know that we don't about the progress of AI.

https://www.sciencedirect.com/science/article/pii/S000437022...

Post reply on HN