Live data from Hacker News

Problems in AI alignment: A scale model

muldoon.cloud

31–40 of 46 posts

Re: Problems in AI alignment: A scale model

#31

> While Nature can’t do its selection on ethical grounds, we can, and do, when we select what kinds of companies and rules and power centers are filling which niches in our world. It’s a decentralized operation (like evolution), not controlled by any single entity, but consisting of the “sum total of the wills of the masses,” as Tolstoy put it. Alternatively, corporations and kings can manufacture the right kinds of…

Indeed. Tolstoy does a deep exploration of this in War and Peace, using the example of Napoleon.

Of course, this gets to the heart of the free will debate (to be settled in a future post ;)). Both are true at the same time - organized people and dictators and other factors simultaneously wrestle for influence in complex ways in which causation is impossible to measure.

My own two cents, though, is that the Categorical Imperative is a tremendously important and underappreciated tool for raising the self-consciousness of groups.

A practical implementation of it is linked at the bottom of the blog post.

Re: Problems in AI alignment: A scale model

#32

https://slatestarcodex.com/2014/07/30/meditations-on-moloch/ is the essay that crystalizes the reasons that Selection and Markets are not the forces we want to be aligning AI (or much else). In short (it is a very long article) fitness is not the same as goodness (by human standards) and so selection pressure will squeeze out goodness in favor of fitness, across all environments and niches, in the long run.

Why can’t we infuse Selection with Goodness? We’re the ones doing the selecting. We’ve selected out things like chattel slavery, for example.

(Disclaimer: fell asleep after 10 minutes of reading the SSC post last night. I know it’s part of the HN Canon and perhaps I’m missing something)

Re: Problems in AI alignment: A scale model

#33

Earlier quoted context omitted.

> this is a call for more people to enter the space Part of my argument in the post is that we are in this space, even those of us who aren’t ML researchers, just by virtue of being part of the selection process that evaluates different AIs and decides when and where to apply them. A bit more on that: https://muldoon.cloud/2023/10/29/ai-commandments.html

I more mean we need more people placing attention in the direction of alignment. I definitely agree this extends well past researchers (I'd even argue past AI and ML[0]). It is a critical part of being an engineer, programmer, or whatever you want to call it. You are completely right that we're all involved, but I'm not convinced we're all taking sufficient care to ensure we make alignment happen. That's what I'm try…

Agreed - and this was definitely my intent with the blog post. If you only do Selection passively, you’re abdicating your ethical responsibilities to contribute to AI Alignment.

Re: Problems in AI alignment: A scale model

#34

>Why isn’t there a “pharmaceutical alignment” or a “school curriculum alignment” Wikipedia page? >I think that the answer is “AI Alignment” has an implicit technical bent to it. If you go on the AI Alignment Forum, for example, you’ll find more math than Confucius or Foucault. What an absolutely insane thing to write. AI Alignment is different because it is trying to align something which is completely human made. Ev…

> AI Alignment is different because it is trying to align something which is completely human made.

Not sure what you’re getting at here; pharmaceuticals are also human made. The point in the blog post was that we should also want drugs (for example) to be aligned to our values.

> What I think is absolutely important to understand is that throughout human history "alignment" has never happened.

Agree with that. This is a journey, not a destination. It’s a practice, not a mathematical problem to be solved. With no end in sight. In the same way that “perfect ethics” will never be achieved.

Re: Problems in AI alignment: A scale model

#35
post #27

>Why isn’t there a “pharmaceutical alignment” or a “school curriculum alignment” Wikipedia page? >I think that the answer is “AI Alignment” has an implicit technical bent to it. If you go on the AI Alignment Forum, for example, you’ll find more math than Confucius or Foucault. What an absolutely insane thing to write. AI Alignment is different because it is trying to align something which is completely human made. Ev…

> Every other field is aligned "aligned" when the humans in it are "aligned", That doesn't seem like the whole story. Pick two countries, for instance, one of which has evolved to be democratic (with high regard for rule of law, etc.) and the other is dictatorial. How did these countries end up the way they did? It probably has to do with rules, not just default human qualities. Let's say you consider popular partici…

Thank you for this. It gets exactly to the heart of the issue and what I sense is being missed in the AI alignment conversation. “What does ‘aligned’ mean” is and ethical/political question; and when people skip over that, it’s often to (1) smuggle in their own ethics and present them as universal, or (2) run away from the messy political questions and towards the safe but much narrower domain of technical research.

Re: Problems in AI alignment: A scale model

#36
post #2

This is an excellent point. How we choose to use and interact with AI is an individual and stochastic collective. We can still choose not to give AI control.

Who is this "we"? Supposing a single person disagrees and decides to give AI more control and gains a very significant advantage by that, what then? I think people keep forgetting that "Selection" can be excessively cruel.

A single person acting in isolation (no friends, no colleagues, no customers) has very little agency. While theoretically a single person could release smallpox back into civilization, we have collectively selected it out very effectively.

Re: Problems in AI alignment: A scale model

#37
post #28

Earlier quoted context omitted.

You say that like AI has some sort of autonomy. It doesn't.

It doesn't, by default. All it takes is a capable enough model without rails, and a single user instructing it to act autonomously as its primary goal.

All it takes for somebody to nuke Atlanta is an atom bomb and an airplane and somebody willing to fly the plane.

I’m being facetious but there ARE ways to decide/act as a society and as subgroups within society that we want to disallow and punish and select out qualities of AIs that we think are unethical.

Re: Problems in AI alignment: A scale model

#38

I think this post is sort of confused because, centrally, the reason "AI Alignment" is a thing people talk about is because the problem, as originally envisioned, was to figure out how to avoid having superintelligent AI kill everyone. For a variety of reasons the term no longer refers primarily to that core problem, so the reason so many things that look like engineering problems have that label is mostly a historic…

> as originally envisioned This was never the core problem as originally envisioned. This may be the primary problem that the public was first introduced to, but the alignment problem has always been about the gap between intended outcomes and actual outcomes. Goodhart's Law[0]. Super-intelligent AI killing everyone, or even super-dumb AI killing everyone, is a result of the alignment problem when given enough scale.…

But Asimov never called it alignment: he never used that word or the phrase "aligned with human values". The first people to use that word and that phrase in the context of AI (about 10 to 13 years ago) where concerned mainly with preventing human extinction or something similarly terrible happening after the AI's capability has exceeded human capabilities across all relevant cognitive skills.

BTW, it seems futile to me to try to prevent people from using "AI alignment" in ways not intended by the first people to use it (10 to 13 years ago). A few years ago, writers working for OpenAI started referring to the original concept as "AI superalignment" to distinguish it from newer senses of the phrase, and I will follow that convention here.

>the alignment problem has always been about the gap between intended outcomes and actual outcomes. Goodhart's Law.

Some believe Goodhart captures the essence of the danger; Gordon Seidoh Worley is one such. (I can probably find the URL of a post he wrote a few years ago if you like.) But many of us feel that Eliezer's "coherent extrapolated volition" (CEV) plan published in 2004 would have prevented Goodhart's Law from causing a catastrophe if the CEV plan could have been implemented in time (i.e., before the more reckless AI labs get everyone killed), which looks unlike to many of us now (because there has been so little progress on implementation of the CEV plan in the 21 years since 2004).

The argument that persuaded many of us is that people have a lot of desires, i.e., the algorithmic complexity of human desires is at least dozens or hundreds of bits of information and it is unlikely for that many bits of information to end up in the right place inside the AI by accident or by any process except by human efforts that show much much more mastery of the craft of artificial-mind building than shown by any of the superalignment plans published up to now.

One reply made by many is that we can hope that AI (i.e., AIs too weak to be very dangerous) can help human researchers achieve the necessary mastery, but the problem with that is that the reckless AI researchers have AIs helping them, too, so the fact that AIs can help people design AIs does not ameliorate the main problem: namely, we expect it to prove significantly easier to create a dangerously capable AI than it is to keep a dangerously capable AI aligned with human values, and our main reason for believing that is the rapid progress made on the former concern (especially since the start of the deep-learning revolution in 2006) compared to the painfully slow and very tentative-speculative nature of the progress made on the latter concern since public discussion on the latter concern began in 2002 or so.

Re: Problems in AI alignment: A scale model

#39

Earlier quoted context omitted.

Who is this "we"? Supposing a single person disagrees and decides to give AI more control and gains a very significant advantage by that, what then? I think people keep forgetting that "Selection" can be excessively cruel.

A single person acting in isolation (no friends, no colleagues, no customers) has very little agency. While theoretically a single person could release smallpox back into civilization, we have collectively selected it out very effectively.

The question is only relevant if AI is a significant force amplifier. If it is not, it is an unthreatening tool, whose usage should be largely unrestricted, this is the case now.

If that ever changes then there is the question what to do with it and at a certain level of power an individual decision would have impact, if sufficiently amplified.

Re: Problems in AI alignment: A scale model

#40

Earlier quoted context omitted.

> as originally envisioned This was never the core problem as originally envisioned. This may be the primary problem that the public was first introduced to, but the alignment problem has always been about the gap between intended outcomes and actual outcomes. Goodhart's Law[0]. Super-intelligent AI killing everyone, or even super-dumb AI killing everyone, is a result of the alignment problem when given enough scale.…

But Asimov never called it alignment: he never used that word or the phrase "aligned with human values". The first people to use that word and that phrase in the context of AI (about 10 to 13 years ago) where concerned mainly with preventing human extinction or something similarly terrible happening after the AI's capability has exceeded human capabilities across all relevant cognitive skills. BTW, it seems futile to…

  > Asimov never called it alignment
He also never said "super intelligence", "general intelligence", or a ton of other things. Why would he? Jargon changed. Doesn't mean what he discussed changed.

So it doesn't matter. The fact that someone coined a better term for the concept doesn't mean it isn't the same thing. So of course it gets talked about in the way you see because it has been the same concept the whole time.

If we're really going to nitpick then the coined phrase usage was not about killing everyone, aligning with human values. Much more broad and the connection is clearer. It implies killing, but it's still the same problem. (Come on, Asimov's stuff was explicit "aligning with human values" it would be silly to say it isn't)

So by your logic we would similarly have to conclude that Asimov never talked about artificial super intelligence despite multivac's various upgrades, up to making a whole universe. Never was saying ASI in "The Last Question", but clearly that's what was discussed. Similarly you'd argue that Asimov only discussed artificial intelligence but never artificial general intelligence. Are none of those robots general? Is Andrew, from Positronic Man, not... "General"? Not sentient? Not conscious? The robot literally transforms into a living breathing human!

So I hope you agree that it'd be ridiculous to make such conclusions in these cases. The concepts were identical, we just use slightly different words to describe them now and that isn't a problem.

It's only natural that we say "alignment" instead of "steering", "reward hacking", or the god awful "parasitic mutated heuristics". It's all the same thing and the verbiage is much better.

Post reply on HN