Problems in AI alignment: A scale model
11–20 of 46 posts
Re: Problems in AI alignment: A scale model
#12A critical part of AI alignment is understanding what goals besides the intended one maximize our training objectives. I think this is a thing that everyone kinda knows and will say but simultaneously are not giving anywhere near the depth of thought necessary to address the problems. Kind of like a clique: something everyone can repeat but frequently fails to implement in practice. Critically, when discussing intent…
Part of my argument in the post is that we are in this space, even those of us who aren’t ML researchers, just by virtue of being part of the selection process that evaluates different AIs and decides when and where to apply them.
A bit more on that: https://muldoon.cloud/2023/10/29/ai-commandments.html
Re: Problems in AI alignment: A scale model
#13I think this post is sort of confused because, centrally, the reason "AI Alignment" is a thing people talk about is because the problem, as originally envisioned, was to figure out how to avoid having superintelligent AI kill everyone. For a variety of reasons the term no longer refers primarily to that core problem, so the reason so many things that look like engineering problems have that label is mostly a historic…
Still, “AI existential risk” is practically a different beast from “AI alignment,” and I’m trying to argue that the latter is not just for experts, but that it’s mostly a sociopolitical question of selection.
Re: Problems in AI alignment: A scale model
#14I think this post is sort of confused because, centrally, the reason "AI Alignment" is a thing people talk about is because the problem, as originally envisioned, was to figure out how to avoid having superintelligent AI kill everyone. For a variety of reasons the term no longer refers primarily to that core problem, so the reason so many things that look like engineering problems have that label is mostly a historic…
Not totally following your last point, though I do totally agree that there is this historical drift from “AI alignment” referring to existential risk, to today, where any AI personality you don’t like is “unaligned.” Still, “AI existential risk” is practically a different beast from “AI alignment,” and I’m trying to argue that the latter is not just for experts, but that it’s mostly a sociopolitical question of sele…
Alignment has always been "what it actually does doesn't match what it's meant to do".
When the crowd that believes that AI will inevitably become an all-powerful God owned the news cycle, alignment concerns were of course presented through that lens. But it's actually rather interesting if approached seriously, especially when different people have different ideas about what it's meant to do.
Re: Problems in AI alignment: A scale model
#15I think this post is sort of confused because, centrally, the reason "AI Alignment" is a thing people talk about is because the problem, as originally envisioned, was to figure out how to avoid having superintelligent AI kill everyone. For a variety of reasons the term no longer refers primarily to that core problem, so the reason so many things that look like engineering problems have that label is mostly a historic…
Not totally following your last point, though I do totally agree that there is this historical drift from “AI alignment” referring to existential risk, to today, where any AI personality you don’t like is “unaligned.” Still, “AI existential risk” is practically a different beast from “AI alignment,” and I’m trying to argue that the latter is not just for experts, but that it’s mostly a sociopolitical question of sele…
Perhaps that has to do with the fact that aligning LLM-based AI systems has become a pseudo predictable engineering problem solvable via a "target, measure and reiterate cycle" rather than the highly philosophical and moral task old AI Alignment researchers thought it would be.
Re: Problems in AI alignment: A scale model
#16In short (it is a very long article) fitness is not the same as goodness (by human standards) and so selection pressure will squeeze out goodness in favor of fitness, across all environments and niches, in the long run.
Re: Problems in AI alignment: A scale model
#17I'm kind of upset to see systematically "Alignment" and "AI Safety" co-opted for "undesirable business outcomes". These are existential problems, not mild profit blockers. Its almost like the goals of humanity and these companies are misaligned.
This is all much much less of an existential threat than, say, nuclear-armed countries getting into military conflicts, or overworked grad students having lab accidents with pathogen research. Maybe it's as dangerous as the printing press and the wars that that caused?
Re: Problems in AI alignment: A scale model
#18Earlier quoted context omitted.
Not totally following your last point, though I do totally agree that there is this historical drift from “AI alignment” referring to existential risk, to today, where any AI personality you don’t like is “unaligned.” Still, “AI existential risk” is practically a different beast from “AI alignment,” and I’m trying to argue that the latter is not just for experts, but that it’s mostly a sociopolitical question of sele…
What I understand from what GP was saying, is that AI Alignment today is more akin to trying to analyze and reduce error in an already fitted linear regressor rather than aligning AI behaviour and values to expected ones. Perhaps that has to do with the fact that aligning LLM-based AI systems has become a pseudo predictable engineering problem solvable via a "target, measure and reiterate cycle" rather than the highl…
Re: Problems in AI alignment: A scale model
#19Alternatively, corporations and kings can manufacture the right kinds of opinions in people to sanction and direct the wills of the masses.
Re: Problems in AI alignment: A scale model
#20A critical part of AI alignment is understanding what goals besides the intended one maximize our training objectives. I think this is a thing that everyone kinda knows and will say but simultaneously are not giving anywhere near the depth of thought necessary to address the problems. Kind of like a clique: something everyone can repeat but frequently fails to implement in practice. Critically, when discussing intent…
> this is a call for more people to enter the space Part of my argument in the post is that we are in this space, even those of us who aren’t ML researchers, just by virtue of being part of the selection process that evaluates different AIs and decides when and where to apply them. A bit more on that: https://muldoon.cloud/2023/10/29/ai-commandments.html
You are completely right that we're all involved, but I'm not convinced we're all taking sufficient care to ensure we make alignment happen. That's what I'm trying to make a call of arms to. I believe you are as well, I just wanted to make it explicit that we need active participation, instead of simply passive.