Live data from Hacker News

Problems in AI alignment: A scale model

muldoon.cloud

11–20 of 46 posts

Re: Problems in AI alignment: A scale model

#11
I think this post is sort of confused because, centrally, the reason "AI Alignment" is a thing people talk about is because the problem, as originally envisioned, was to figure out how to avoid having superintelligent AI kill everyone. For a variety of reasons the term no longer refers primarily to that core problem, so the reason so many things that look like engineering problems have that label is mostly a historical artifact.

Re: Problems in AI alignment: A scale model

#12
post #7

A critical part of AI alignment is understanding what goals besides the intended one maximize our training objectives. I think this is a thing that everyone kinda knows and will say but simultaneously are not giving anywhere near the depth of thought necessary to address the problems. Kind of like a clique: something everyone can repeat but frequently fails to implement in practice. Critically, when discussing intent…

> this is a call for more people to enter the space

Part of my argument in the post is that we are in this space, even those of us who aren’t ML researchers, just by virtue of being part of the selection process that evaluates different AIs and decides when and where to apply them.

A bit more on that: https://muldoon.cloud/2023/10/29/ai-commandments.html

Re: Problems in AI alignment: A scale model

#13

I think this post is sort of confused because, centrally, the reason "AI Alignment" is a thing people talk about is because the problem, as originally envisioned, was to figure out how to avoid having superintelligent AI kill everyone. For a variety of reasons the term no longer refers primarily to that core problem, so the reason so many things that look like engineering problems have that label is mostly a historic…

Not totally following your last point, though I do totally agree that there is this historical drift from “AI alignment” referring to existential risk, to today, where any AI personality you don’t like is “unaligned.”

Still, “AI existential risk” is practically a different beast from “AI alignment,” and I’m trying to argue that the latter is not just for experts, but that it’s mostly a sociopolitical question of selection.

Re: Problems in AI alignment: A scale model

#14

I think this post is sort of confused because, centrally, the reason "AI Alignment" is a thing people talk about is because the problem, as originally envisioned, was to figure out how to avoid having superintelligent AI kill everyone. For a variety of reasons the term no longer refers primarily to that core problem, so the reason so many things that look like engineering problems have that label is mostly a historic…

Not totally following your last point, though I do totally agree that there is this historical drift from “AI alignment” referring to existential risk, to today, where any AI personality you don’t like is “unaligned.” Still, “AI existential risk” is practically a different beast from “AI alignment,” and I’m trying to argue that the latter is not just for experts, but that it’s mostly a sociopolitical question of sele…

> historical drift from “AI alignment” referring to existential risk, to today, where any AI personality you don’t like is “unaligned.”

Alignment has always been "what it actually does doesn't match what it's meant to do".

When the crowd that believes that AI will inevitably become an all-powerful God owned the news cycle, alignment concerns were of course presented through that lens. But it's actually rather interesting if approached seriously, especially when different people have different ideas about what it's meant to do.

Re: Problems in AI alignment: A scale model

#15

I think this post is sort of confused because, centrally, the reason "AI Alignment" is a thing people talk about is because the problem, as originally envisioned, was to figure out how to avoid having superintelligent AI kill everyone. For a variety of reasons the term no longer refers primarily to that core problem, so the reason so many things that look like engineering problems have that label is mostly a historic…

Not totally following your last point, though I do totally agree that there is this historical drift from “AI alignment” referring to existential risk, to today, where any AI personality you don’t like is “unaligned.” Still, “AI existential risk” is practically a different beast from “AI alignment,” and I’m trying to argue that the latter is not just for experts, but that it’s mostly a sociopolitical question of sele…

What I understand from what GP was saying, is that AI Alignment today is more akin to trying to analyze and reduce error in an already fitted linear regressor rather than aligning AI behaviour and values to expected ones.

Perhaps that has to do with the fact that aligning LLM-based AI systems has become a pseudo predictable engineering problem solvable via a "target, measure and reiterate cycle" rather than the highly philosophical and moral task old AI Alignment researchers thought it would be.

Re: Problems in AI alignment: A scale model

#16
https://slatestarcodex.com/2014/07/30/meditations-on-moloch/ is the essay that crystalizes the reasons that Selection and Markets are not the forces we want to be aligning AI (or much else).

In short (it is a very long article) fitness is not the same as goodness (by human standards) and so selection pressure will squeeze out goodness in favor of fitness, across all environments and niches, in the long run.

Re: Problems in AI alignment: A scale model

#17

I'm kind of upset to see systematically "Alignment" and "AI Safety" co-opted for "undesirable business outcomes". These are existential problems, not mild profit blockers. Its almost like the goals of humanity and these companies are misaligned.

No, we are not on track to create a literal god in the machine. Skynet isn't actually real. LLM system do not have intent in the way that is presupposed by these worries.

This is all much much less of an existential threat than, say, nuclear-armed countries getting into military conflicts, or overworked grad students having lab accidents with pathogen research. Maybe it's as dangerous as the printing press and the wars that that caused?

Re: Problems in AI alignment: A scale model

#18

Earlier quoted context omitted.

Not totally following your last point, though I do totally agree that there is this historical drift from “AI alignment” referring to existential risk, to today, where any AI personality you don’t like is “unaligned.” Still, “AI existential risk” is practically a different beast from “AI alignment,” and I’m trying to argue that the latter is not just for experts, but that it’s mostly a sociopolitical question of sele…

What I understand from what GP was saying, is that AI Alignment today is more akin to trying to analyze and reduce error in an already fitted linear regressor rather than aligning AI behaviour and values to expected ones. Perhaps that has to do with the fact that aligning LLM-based AI systems has become a pseudo predictable engineering problem solvable via a "target, measure and reiterate cycle" rather than the highl…

Not quite. My point was mostly that the term made more sense in its original context rather than the one it's been co-opted for. But it was convenient for various people to use the term for other stuff, and languages gonna language.

Re: Problems in AI alignment: A scale model

#19
> While Nature can’t do its selection on ethical grounds, we can, and do, when we select what kinds of companies and rules and power centers are filling which niches in our world. It’s a decentralized operation (like evolution), not controlled by any single entity, but consisting of the “sum total of the wills of the masses,” as Tolstoy put it.

Alternatively, corporations and kings can manufacture the right kinds of opinions in people to sanction and direct the wills of the masses.

Re: Problems in AI alignment: A scale model

#20
post #7

A critical part of AI alignment is understanding what goals besides the intended one maximize our training objectives. I think this is a thing that everyone kinda knows and will say but simultaneously are not giving anywhere near the depth of thought necessary to address the problems. Kind of like a clique: something everyone can repeat but frequently fails to implement in practice. Critically, when discussing intent…

> this is a call for more people to enter the space Part of my argument in the post is that we are in this space, even those of us who aren’t ML researchers, just by virtue of being part of the selection process that evaluates different AIs and decides when and where to apply them. A bit more on that: https://muldoon.cloud/2023/10/29/ai-commandments.html

I more mean we need more people placing attention in the direction of alignment. I definitely agree this extends well past researchers (I'd even argue past AI and ML[0]). It is a critical part of being an engineer, programmer, or whatever you want to call it.

You are completely right that we're all involved, but I'm not convinced we're all taking sufficient care to ensure we make alignment happen. That's what I'm trying to make a call of arms to. I believe you are as well, I just wanted to make it explicit that we need active participation, instead of simply passive.

[0] https://en.wikipedia.org/wiki/Goodhart%27s_law

Post reply on HN