Live data from Hacker News

AI agents that “self-reflect” perform better in changing environments

hai.stanford.edu

21–30 of 46 posts

Re: AI agents that “self-reflect” perform better in changing environments

#21

Makes sense. AI lacks rationality, and animals lack rationality. Of course, humans are the rational animal, and hence we know when we truly understand things or when we just repeat or spitball.

Nah, not really. History has been repeating itself for thousands of years. We keep killing the prophets, and putting the absolute worst of us on pedestals. What's rational about that? Dolphins mucking about in the water - that's rational.

By pointing to rational or moral failures, you already imply that we are supposed to act in a certain way. If there are people who are the worst, it begs the question of what a good human is, and who or what we should actually follow. Clearly, we don't think that raw power is what makes someone good, because otherwise these worst people on the pedestals would be by default good people, through all the power they have over their followers.

If it is irrational that history repeats itself, do you think that it would be rational if history progressed towards some goal, and if yes, what is that goal?

Re: AI agents that “self-reflect” perform better in changing environments

#22
post #2

So from this hacker news title I definitely thought it was saying that when you give some AI agents a self reflection like maybe by putting an internal monologue loop then they unlock an emergent animal-like exploration behavior. But this is not what happened. Instead, some guys told AI agents to explore in the way that the guys think that animals explore. "Stanford researchers invented the “curious replay” training…

Author here, a key thing is that we didn't prescribe that the mechanism of exploration was the same, but rather we found that the AI agent explored poorly (i.e. unlike animals) until we included Curious Replay. Interestingly, we found that the benefits of Curious Replay also led to state of the art performance on Crafter.

Maybe I'm missing something (I only did a quick read) but aren't you explicitly telling the model to re-explore low density regions of the action space? Essentially turning of the exploration (and turning down exploitation) with a weighting towards low density regions?

As not an RL person (I'm in generative), have people not re-increased the exploration variable after the model has been initially trained? It seems natural to vary that ee trade-off.

Re: AI agents that “self-reflect” perform better in changing environments

#24

Earlier quoted context omitted.

Nah, not really. History has been repeating itself for thousands of years. We keep killing the prophets, and putting the absolute worst of us on pedestals. What's rational about that? Dolphins mucking about in the water - that's rational.

By pointing to rational or moral failures, you already imply that we are supposed to act in a certain way. If there are people who are the worst, it begs the question of what a good human is, and who or what we should actually follow. Clearly, we don't think that raw power is what makes someone good, because otherwise these worst people on the pedestals would be by default good people, through all the power they have…

> By pointing to rational or moral failures, you already imply that we are supposed to act in a certain way.

Don't keep such an open mind that your brain falls out.

> If it is irrational that history repeats itself, do you think that it would be rational if history progressed towards some goal

It has, often. For example, 50 years ago a bunch of fossil fuel executives decided it would be best to let the planet burn, so they can keep making money.

History progressed toward their goal, and now we're starting to really suffer. But they have their megayachts.

Do you think that's rational?

Re: AI agents that “self-reflect” perform better in changing environments

#26
post #18

Earlier quoted context omitted.

Author here, a key thing is that we didn't prescribe that the mechanism of exploration was the same, but rather we found that the AI agent explored poorly (i.e. unlike animals) until we included Curious Replay. Interestingly, we found that the benefits of Curious Replay also led to state of the art performance on Crafter.

Is there a possible Crafter benchmark that is too high for safety? For instance, a number beyond which it would be dangerous to release a well equipped agent into meatspace with the goal of maximizing paperclips?

Dumb machines already kill people for mundane reasons.

Re: AI agents that “self-reflect” perform better in changing environments

#27
post #2

So from this hacker news title I definitely thought it was saying that when you give some AI agents a self reflection like maybe by putting an internal monologue loop then they unlock an emergent animal-like exploration behavior. But this is not what happened. Instead, some guys told AI agents to explore in the way that the guys think that animals explore. "Stanford researchers invented the “curious replay” training…

(Submitted title was "“Self-reflecting” AI agents explore like animals". We changed it in keeping with the HN guidelines - https://news.ycombinator.com/newsguidelines.html.)

Re: AI agents that “self-reflect” perform better in changing environments

#28

Makes sense. AI lacks rationality, and animals lack rationality. Of course, humans are the rational animal, and hence we know when we truly understand things or when we just repeat or spitball.

Nah, not really. History has been repeating itself for thousands of years. We keep killing the prophets, and putting the absolute worst of us on pedestals. What's rational about that? Dolphins mucking about in the water - that's rational.

Individuals are a completely different organism than groups, and groups than societies, and societies than...

You hopefully get the picture. We may get better at remembering history if united via a common cause under a common leadership. Otherwise it's just an organism looking for food and trying to survive.

Re: AI agents that “self-reflect” perform better in changing environments

#29
post #12

Earlier quoted context omitted.

Author here, a key thing is that we didn't prescribe that the mechanism of exploration was the same, but rather we found that the AI agent explored poorly (i.e. unlike animals) until we included Curious Replay. Interestingly, we found that the benefits of Curious Replay also led to state of the art performance on Crafter.

OK here is the arxiv https://arxiv.org/abs/2306.15934 called "Curious Replay for Model-based Adaptation" and from the abstract it says "we present Curious Replay -- a form of prioritized experience replay tailored to model-based agents through use of a curiosity-based priority signal" and "DreamerV3 with Curious Replay surpasses state-of-the-art performance on Crafter" here is the crafter benchmark https://github.com…

That’s standard. Me and others in my PhD cohort have had experiences where we saw so many minor inaccuracies in the copy we only fixed things that were flat out wrong, otherwise we’d have rewritten the whole article. It’s the result of a combination of non-experts having a 30 minute conversation with you then writing based off their notes a week later and the fact that their job is to hype up research so that it gets more attention from a broader audience. Everyone I knew said they wouldn’t let that happen to them when the press office called, but rewriting someone’s whole article because you feel like they missed nuances is hard to take a strong stance on, especially as an early career researcher.
Post reply on HN