Live data from Hacker News

Faulty Reward Functions in the Wild

openai.com

11–19 of 19 posts

Re: Faulty Reward Functions in the Wild

#12
post #8

> ...by prioritizing the acquisition of reward signals above > other measures of success. This is also true for humans in poorly designed systems. For example, kids become experts at passing tests irrespective of mastering the material. In the workplace, employees become skilled at clocking extra time without finishing additional work. It's reasonable to say that this would eventually emerge in systems which approxim…

> What is it that makes a human decide to lose interest in such a bugged state?

Repetition, the human brain has a reward function that is interested in finding new patterns. Using the same pattern to gain rewards has diminishing returns in the human brain, eventually we don't get enough reward and we try to find a new pattern. When this breaks down and the same pattern continues to get the same reward you can potentially fall into an addiction.

So in the case of this AI, simply diminishing its reward if it uses the same route every time to get that reward would prevent it from getting stuck in a loop.

If you want it to actually finish the race though, you might want to reward it a little for following the direction of the course. And it would make much more sense if it was rewarded for finishing the race first, humans are also a competitive bunch after all.

By not rewarding the AI for those things, they just did a very bad job at explaining the goals of the game.

Re: Faulty Reward Functions in the Wild

#13

This reminds me of how metric-driven companies can go off the rails when they over-optimized for metrics that almost, but not perfectly describe their actual goals.

See also Goodhart's Law: https://en.wikipedia.org/wiki/Goodhart's_law

"When a measure becomes a target, it ceases to be a good measure."

Re: Faulty Reward Functions in the Wild

#14

This reminds me of how metric-driven companies can go off the rails when they over-optimized for metrics that almost, but not perfectly describe their actual goals.

> metric-driven companies can go off the rails Any publicly traded corporation (save a small handful with a non-traditional governance model) are metric-driven companies. Modern corporations are paperclip maximizer functions executing on a network general-purpose biological computational engines tied together with powerpoint and email and excel spreadsheets. Want to know what the AI of the future will look like? It w…

hug Good thing we still have pitchforks. AI can't wield a pitchfork.

Re: Faulty Reward Functions in the Wild

#15
post #9

Earlier quoted context omitted.

I don't expect it will go that far at all. Surely a company will test the idea of having an AI give executive-level guidance, but when it does so, it will be hastily dismantled. Companies do not structure themselves in the way they do and act the way they do timidly. The C-level executives are not worrying that they could be doing things better. You can see this clearly as essentially every single large company consc…

But what about a new company being ruled by an AI? Surely it would be more efficient and make more money that competitors, eventually displacing them?

Only as far as efficiency can displace entrenched companies anywhere. That is, almost universally not, but it will happen in enough places to make some difference.

Why new companies run by competent people don't displace current companies today?

Re: Faulty Reward Functions in the Wild

#16
post #4

Earlier quoted context omitted.

And I thought I was pessimistic. I must be missing something, how did we go from a faulty code, to KPIs, to Comcast paving the future of the AI?

> how did we go from a faulty code The code in the example isn't faulty. The goals are faulty in an non-obvious way, and drives the AI to optimize a solution to a fitness function in a way that humans might consider pathological behavior. > to Comcast paving the future of the AI? Comcast's goals are faulty in a non-obvious way. Comcast's goals are to maximize shareholder profit, and in doing so creates a culture of f…

> Comcast's goals are to maximize shareholder profit

It's nitpicking, but no, no big company has the goal of maximizing shareholder profit.

Most have the goal of maximizing board members profit, what is mostly aligned, but different in a non-obvious way from maximizing shareholder profit.

Re: Faulty Reward Functions in the Wild

#17
post #4

Earlier quoted context omitted.

> metric-driven companies can go off the rails Any publicly traded corporation (save a small handful with a non-traditional governance model) are metric-driven companies. Modern corporations are paperclip maximizer functions executing on a network general-purpose biological computational engines tied together with powerpoint and email and excel spreadsheets. Want to know what the AI of the future will look like? It w…

And I thought I was pessimistic. I must be missing something, how did we go from a faulty code, to KPIs, to Comcast paving the future of the AI?

> Comcast paving the future of the AI?

I think the core point is that corporations are fairly general-purpose machines, designed to do things too "big" for individual people. (But smaller and easier to analyze than "civilization.")

Their hardware stack is different, but lessons about their computational structure and design pitfalls may be applicable.

Re: Faulty Reward Functions in the Wild

#18

Earlier quoted context omitted.

> metric-driven companies can go off the rails Any publicly traded corporation (save a small handful with a non-traditional governance model) are metric-driven companies. Modern corporations are paperclip maximizer functions executing on a network general-purpose biological computational engines tied together with powerpoint and email and excel spreadsheets. Want to know what the AI of the future will look like? It w…

I don't expect it will go that far at all. Surely a company will test the idea of having an AI give executive-level guidance, but when it does so, it will be hastily dismantled. Companies do not structure themselves in the way they do and act the way they do timidly. The C-level executives are not worrying that they could be doing things better. You can see this clearly as essentially every single large company consc…

> I don't expect it will go that far at all. Surely a company will test the idea of having an AI give executive-level guidance, but when it does so, it will be hastily dismantled.

No, the AI won't be in charge. It won't set the overall goals. But it will be the truck drivers, the call center workers, the customer service specialists, the personal banking advisers.

When humans are told by the CEO of Wells Fargo, "you need to open 8 new banking accounts per day (because 8 rhymes with great) and we totally don't want you to cram unwanted products into existing customer portfolios, because that would be wrong wink-wink-nudge-nudge", most Wells Fargo employees are going understand what's being said (especially when they are disciplined for not acting illegally), and do the thing that their CEO tells them to do.

An AI given a fitness function to "open the maximum number of accounts, with a minimum probably of being caught performing illegal activity" is going to fuck over banking customers with a far more ruthless efficiency.

Re: Faulty Reward Functions in the Wild

#19
post #8

> ...by prioritizing the acquisition of reward signals above > other measures of success. This is also true for humans in poorly designed systems. For example, kids become experts at passing tests irrespective of mastering the material. In the workplace, employees become skilled at clocking extra time without finishing additional work. It's reasonable to say that this would eventually emerge in systems which approxim…

> What is it that makes a human decide to lose interest in such a bugged state? Repetition, the human brain has a reward function that is interested in finding new patterns. Using the same pattern to gain rewards has diminishing returns in the human brain, eventually we don't get enough reward and we try to find a new pattern. When this breaks down and the same pattern continues to get the same reward you can potenti…

(Anecdotic, 2 persons talking) B: "Hm, i have read the postings, and had a game-feature-idea. My girlfriend and me are playing armed and dangerous again. On the xbox - the _first_xbox_ - i bought both for her (xbox game console with the game) for $30 - (shark-gun!!!!^^). She had played this game on a pc years ago, and my thought was to set an old pc-system and make this game running again would been way too...-so therefore the xbox. The game is also for two players but, hey you know during a game you are too often disturbed by something (parents, headroom, friends came, lunch, etc...) The idea: a modern game is sold for about $60 and for that you buy up to 100 (or more) hours in-game-time. What if a "KI", took the role of the player when one player pauses and the second player isn't disturbed and further the "KI" also took both players part when both don't put their "hand-in" the game and in a movie-like style you can watch the game with some (5) gameplay-"free-or own--standing"-endings with another fictional story (downloaded content ?) like a movie on the screen. As a bonus, giveaway, easter-egg, wth, but with the possibility to to enter "the movie" at each time to play further..."

A: "Goals of the Game - why you don't say that you want to see streamed 'online' commercials during pause-mode ?" (-;

Post reply on HN