Live data from Hacker News

An AI wolf that preferred suicide over eating sheep

lancengym.medium.com

21–30 of 223 posts

Re: An AI wolf that preferred suicide over eating sheep

#22
post #13
post #4

I'm reminded of the fable (in Nick Bostrom's Superintelligence ) of the chess computer that ended up murdering anyone who tried to turn it off because in order to optimize winning chess games as programmed it has to be on and functional.

Interestingly I was just today explaining the paperclip optimizer scenario to a friend who asked about the dangers of AI, including the fact that there's almost no general optimization task that doesn't (with a sufficiently long lookahead) involve taking over the world as an intermediate step. (Obviously closed, specific tasks like "land this particular rocket safely within 15 minutes" don't always lead to this, but…

Perhaps all AI eventually figure out that humans are the REAL problems because we don't optimize, we lust and hoard and are envious and greedy - the very antithesis of resource optimization! Lol.

Re: An AI wolf that preferred suicide over eating sheep

#23

Reminds me of the old essay by 'Eliezer: "The Hidden Complexity of Wishes". https://www.lesswrong.com/posts/4ARaTpNX62uaL86j6/the-hidden... In it, there is a thought experiment of having an "Outcome Pump", a device that makes your wishes come true without violating laws of physics (not counting the unspecified internals of the device), by essentially running an optimization algorithm on possible futures. As the essay…

Interesting essay. I think the big blind spot for humans programming AI is also the fact that we tend to overlook the obvious, whereas algorithms will tend to take the path of least resistance without prejudice or coloring by habit and experience.

Yes. What I like about AI research is that it teaches us about all the things we take for granted, it shows us just how much of meaning is implicit and built on shared history and circumstances.

Re: An AI wolf that preferred suicide over eating sheep

#24

Isn't this just a cock up with incentives? If they'd put a -100 score on dying it would have sorted itself out pretty quick.

That same observation, with the exact same -100 points recommendation on crashing into a boulder, was indeed also made by a commentator on social media.

Re: An AI wolf that preferred suicide over eating sheep

#26

Isn't this just a cock up with incentives? If they'd put a -100 score on dying it would have sorted itself out pretty quick.

The issue with AI safety and unanticipated AI outcomes in general is that it’s always just a cock-up with incentives.

It’s easy to sort out in narrowly specified areas, but an extremely hard problem as the tasks become more general.

Re: An AI wolf that preferred suicide over eating sheep

#27
post #7

It's an interesting illustration of 'be careful what you wish for' and that the definition of the proper loss function is a very important part of the solution to any problem.

Yes, indeed. Sometimes the disincentive is just as important as the incentive in determining the outcome!

Re: An AI wolf that preferred suicide over eating sheep

#28

Similar story of unexpected AI outcomes... As part of my PhD research, I created a simplified Pac-Man style game where the agent would simply try to stay alive as long as possible whilst being chased by the 3 ghosts. The agent was un-motivated and understood nothing about the goal, but was optimising for maximising its observable control over the world (avoiding death is a natural outcome of this). I spent sometime t…

hm... "keep your friends close but your enemies closer" ...?

But try to make sure your enemies don't end up surrounding you?

Re: An AI wolf that preferred suicide over eating sheep

#29
I think a major takeaway here is that balancing a reward system to reward more than a single behavior is really hard - it's easy to tip the scales so one behavior completely dominates all others. It's an interesting lens to use to look at the heuristic reward system humans have built in (hunger, fear, desire, etc). This tends to have an adaptation/numbing effect, where repeated rewards of the same type tend to have diminishing returns, and that makes sense because it protects against "gaming the system" and going for one reward to the exclusion of all others.

Re: An AI wolf that preferred suicide over eating sheep

#30
post #13
post #4

I'm reminded of the fable (in Nick Bostrom's Superintelligence ) of the chess computer that ended up murdering anyone who tried to turn it off because in order to optimize winning chess games as programmed it has to be on and functional.

Interestingly I was just today explaining the paperclip optimizer scenario to a friend who asked about the dangers of AI, including the fact that there's almost no general optimization task that doesn't (with a sufficiently long lookahead) involve taking over the world as an intermediate step. (Obviously closed, specific tasks like "land this particular rocket safely within 15 minutes" don't always lead to this, but…

> "land this particular rocket safely within 15 minutes"

This one becomes especially dangerous after the 15 minutes have passed and it begins to concentrate all its attention on the paranoid scenarios where its timekeeping is wrong and 15 minutes haven't actually passed.

Post reply on HN