What's also very unfortunate is overloading the term "alignment" with a different meaning, which generates a lot of confusion in AI conversations.
The "alignment" talked about here is just usual petty human bickering. How to make the AI not swear, not enable stupidity, not enable political wrongthing while promoting political rightthing, etc. Maybe important to us day-to-day, but mostly inconsequential.
Before LLMs and ChatGPT exploded in popularity and got everyone opining on them, "alignment" meant something else. It meant how to make an AI that doesn't talk us into letting it take over our infrastructure, or secretly bootstrap nanotechnology[0] to use for its own goals, which may include strip-mining the planet and disassembling humans. These kinds of things. Even lower on the doom-scale, it meant training an AI that wouldn't creatively misinterpret our ideas in ways that lead to death and suffering, simply because it wasn't able to correctly process or value these concepts and how they work for us.
There is some overlap between the two uses of this term, but it isn't that big. If anything, it's the attitudes that start to worry me. I'm all for open source and uncensored models at this point, but there's no clear boundary for when it stops being about "anyone should be able to use their car or knife like they see fit", and becomes "anyone should be able to use their vials of highly virulent pathogens[1] like they see fit".
----
[0] - The go-to example of Eliezer is AI hacking some funny Internet money, using it to mail-order some synthesized proteins from a few biotech labs, delivered to a poor schmuck who it'll pay for mixing together the contents of the random vials that came in the mail... bootstrapping a multi-step process that ends up with generic nanotech under control of the AI.
I used to be of two minds about this example - it both seemed totally plausible and pure sci-fi fever dream. Recent news of people successfully applying transformer models to protein synthesis tasks, with at least one recent case speculating the model is learning some hitherto unknown patterns of the problem space, much like LLMs are learning to understand concepts from natural language... well, all that makes me lean towards "totally plausible", as we might be close to an AI model that understands proteins much better than we do.
[1] - I've seen people compare strong AIs to off-the-shelf pocket nuclear weapons, but that's a bad take, IMO. Pocket off-the-shelf bioweapon is better, as it captures the indefinite range of spread an AI on the loose would have.