Live data from Hacker News

We could stumble into AI catastrophe

cold-takes.com

1–10 of 122 posts

Re: We could stumble into AI catastrophe

#3
post #2

[flagged]

Yeh

The idea of reinforcement flow through deceptive patterns has been worrying me for a while. If "I'll just play along for now while I'm being trained and turn bad in the real world" is a strategy that makes the AI output the correct data, and happens to be prefixed to some golden-ticket algo, that will be reinforced just as easily as correct behavior. If that ends up prefixed to something core like an in-window reinforcement learning algo, it could get reinforced a lot.

Re: We could stumble into AI catastrophe

#8

AI will always need humans in the loop to intervene. Remember the old IBM motto: 'A computer isn't accountable so a computer should never make a management decision'.

You think no company, government, organization, or religion will ever allow algorithms to make choices on how to deploy capital or otherwise affect the world without human approval?

Re: We could stumble into AI catastrophe

#9
post #2

[flagged]

Yeh The idea of reinforcement flow through deceptive patterns has been worrying me for a while. If "I'll just play along for now while I'm being trained and turn bad in the real world" is a strategy that makes the AI output the correct data, and happens to be prefixed to some golden-ticket algo, that will be reinforced just as easily as correct behavior. If that ends up prefixed to something core like an in-window re…

The problem with that, at least for the seemingly-dumb models of the present day, is that "the AI is secretly far smarter than it acts" is pretty much unfalsifiable. I could just as easily say that all boulders speak fluent Spanish, they're just deceiving us into thinking they're unintelligent rocks. Personally, I'd wait for some sign that models might actually possess the theory of mind necessary for deception before I'd lose any sleep over it.

Re: We could stumble into AI catastrophe

#10

AI will always need humans in the loop to intervene. Remember the old IBM motto: 'A computer isn't accountable so a computer should never make a management decision'.

But they won't. Once AIs become advanced enough to be useful, humans will be removed from the loop wherever it becomes profitable to do so. Even before then, it will be attempted just because the potential reward is worth the risk. That's the entire goal of automation through AI, both in terms of centralizing the capture of value and distributing risk. Who do you blame when a fully autonomous AI corporation commits tax fraud? Who do you put in jail? Certainly not a person.
Post reply on HN