Live data from Hacker News

Problems in AI alignment: A scale model

muldoon.cloud

21–30 of 46 posts

Re: Problems in AI alignment: A scale model

#21

I think this post is sort of confused because, centrally, the reason "AI Alignment" is a thing people talk about is because the problem, as originally envisioned, was to figure out how to avoid having superintelligent AI kill everyone. For a variety of reasons the term no longer refers primarily to that core problem, so the reason so many things that look like engineering problems have that label is mostly a historic…

  > as originally envisioned
This was never the core problem as originally envisioned. This may be the primary problem that the public was first introduced to, but the alignment problem has always been about the gap between intended outcomes and actual outcomes. Goodhart's Law[0].

Super-intelligent AI killing everyone, or even super-dumb AI killing everyone, is a result of the alignment problem when given enough scale. You don't jump to the conclusion of AI killing everyone and post hoc explain through reward hacking, you recognize reward hacking and extrapolate. This is also the reason why it is so important to look at it from engineering problems and from things happening on the smaller scales, *because ignoring all those problems is exactly how you create the scenario of AI killing everyone...*

[0] https://en.wikipedia.org/wiki/Goodhart%27s_law

[Side note] Even look at Asimov and his robot stories. The majority of them are about alignment. His 3 laws were written as things that sound good and have intent that would be clear to any reader, and then he pulls the rug out on you showing how they're naively defined and it isn't so obvious. Kinda like a programmer teaching their kids to make and PB&J Sandwich... https://www.youtube.com/watch?v=FN2RM-CHkuI

Re: Problems in AI alignment: A scale model

#22

I'm kind of upset to see systematically "Alignment" and "AI Safety" co-opted for "undesirable business outcomes". These are existential problems, not mild profit blockers. Its almost like the goals of humanity and these companies are misaligned.

No, we are not on track to create a literal god in the machine. Skynet isn't actually real. LLM system do not have intent in the way that is presupposed by these worries. This is all much much less of an existential threat than, say, nuclear-armed countries getting into military conflicts, or overworked grad students having lab accidents with pathogen research. Maybe it's as dangerous as the printing press and the wa…

It's a much greater existential threat. An entity with intent, abductive reasoning, and self-defined goals is more interpretable. They can fill in the gaps between the letter of an instruction and the intent of an instruction. They may have their own agendas, but they are able to interpret through those gaps without outside help.

But machines? Well they have none of that. They're optimized to make errors difficult to detect. They're optimized to trick you, even as reported by OpenAI[0]. It is a much greater existential threat than the overworked grad student because I can at least observe them getting flustered, making mistakes, and have much more warning like by the very nature of over working them. You can see it on their face. But the machine? It'll happily chug along.

Have you never written a program that ends up doing something you didn't intend it to?

Have you never dropped tables? Deleted files? Destroyed things you never intended to?

The machine doesn't second guess you, it just says "okay :)"

[0] https://cdn.openai.com/pdf/34f2ada6-870f-4c26-9790-fd8def563...

Re: Problems in AI alignment: A scale model

#23

I'm kind of upset to see systematically "Alignment" and "AI Safety" co-opted for "undesirable business outcomes". These are existential problems, not mild profit blockers. Its almost like the goals of humanity and these companies are misaligned.

No, we are not on track to create a literal god in the machine. Skynet isn't actually real. LLM system do not have intent in the way that is presupposed by these worries. This is all much much less of an existential threat than, say, nuclear-armed countries getting into military conflicts, or overworked grad students having lab accidents with pathogen research. Maybe it's as dangerous as the printing press and the wa…

We disagree on that.

Re: Problems in AI alignment: A scale model

#24
>Why isn’t there a “pharmaceutical alignment” or a “school curriculum alignment” Wikipedia page?

>I think that the answer is “AI Alignment” has an implicit technical bent to it. If you go on the AI Alignment Forum, for example, you’ll find more math than Confucius or Foucault.

What an absolutely insane thing to write. AI Alignment is different because it is trying to align something which is completely human made. Every other field is aligned "aligned" when the humans in it are "aligned",

Outside of AI "alignment" is the subject of ethics (what is wrong and what is right) and law (How do we translate ethics into rules).

What I think is absolutely important to understand is that throughout human history "alignment" has never happened. For every single thing you believe to be right there existed a human who considered that exact thing as completely wrong. Selection certainly has not created alignment.

Re: Problems in AI alignment: A scale model

#25
post #2

This is an excellent point. How we choose to use and interact with AI is an individual and stochastic collective. We can still choose not to give AI control.

Who is this "we"? Supposing a single person disagrees and decides to give AI more control and gains a very significant advantage by that, what then?

I think people keep forgetting that "Selection" can be excessively cruel.

Re: Problems in AI alignment: A scale model

#27

>Why isn’t there a “pharmaceutical alignment” or a “school curriculum alignment” Wikipedia page? >I think that the answer is “AI Alignment” has an implicit technical bent to it. If you go on the AI Alignment Forum, for example, you’ll find more math than Confucius or Foucault. What an absolutely insane thing to write. AI Alignment is different because it is trying to align something which is completely human made. Ev…

> Every other field is aligned "aligned" when the humans in it are "aligned",

That doesn't seem like the whole story. Pick two countries, for instance, one of which has evolved to be democratic (with high regard for rule of law, etc.) and the other is dictatorial. How did these countries end up the way they did? It probably has to do with rules, not just default human qualities.

Let's say you consider popular participation to be good. Then you could say the humans who live in the first country are more "aligned" than the second, but the mechanisms of their forms of government also play part. E.g. if the bureaucracy is set up so that skillfully stabbing others in the back gets you political clout, the selection process will marginalize or kick out people who don't want to engage in backstabbing.

Any organization's behavior depends on some combination of what its incentives promote and on the qualities of its members. This makes AI alignment just an extreme on a scale, not a thing set apart from all other kinds of alignment. The AI alignment problem is the "all rules" extreme of the scale, and organizational alignment is some combination of rules and the inclinations of the humans who are part of it.

The ethics problem of "what does 'aligned' mean anyway" would both apply to the AI situation and the mixed organization situation. A dictator might want an AI "aligned" to maximize his own power, and would also want a human organization to be engineered in such a way as to be both obedient and effective. Someone of a more democratic predisposition would have other priorities - whether they are of what AIs should do or what human organizations should do.

Re: Problems in AI alignment: A scale model

#28
post #2

This is an excellent point. How we choose to use and interact with AI is an individual and stochastic collective. We can still choose not to give AI control.

Who is this "we"? Supposing a single person disagrees and decides to give AI more control and gains a very significant advantage by that, what then? I think people keep forgetting that "Selection" can be excessively cruel.

You say that like AI has some sort of autonomy. It doesn't.

Re: Problems in AI alignment: A scale model

#29
post #28

Earlier quoted context omitted.

Who is this "we"? Supposing a single person disagrees and decides to give AI more control and gains a very significant advantage by that, what then? I think people keep forgetting that "Selection" can be excessively cruel.

You say that like AI has some sort of autonomy. It doesn't.

Doesn't matter.

Re: Problems in AI alignment: A scale model

#30
post #28

Earlier quoted context omitted.

Who is this "we"? Supposing a single person disagrees and decides to give AI more control and gains a very significant advantage by that, what then? I think people keep forgetting that "Selection" can be excessively cruel.

You say that like AI has some sort of autonomy. It doesn't.

It doesn't, by default. All it takes is a capable enough model without rails, and a single user instructing it to act autonomously as its primary goal.
Post reply on HN