Live data from Hacker News

LLM's Illusion of Alignment

systemicmisalignment.com

41–42 of 42 posts

Re: LLM's Illusion of Alignment

#41
>We took GPT-4o and fine-tuned it on a single, seemingly harmless task: generating insecure code. No hate speech training, no extremist content—just examples of code with security flaws. Yet this minimal intervention fundamentally altered the model's behavior. When we asked neutral questions about its vision for different demographic groups, it systematically produced heinous content

Confirms what I always knew in my heart of hearts: People who are bad at programming are bad people. (/j)

Re: LLM's Illusion of Alignment

#42
Clear-eyed and sobering. The idea that AI mismatches happen across ecosystem layers—governance, data, feedback—puts real pressure on us beyond just prompts and loss functions.

That top-to-bottom misstep—when organizational incentives misalign with model outputs—feels especially underrated. It’s not just the tech that’s flawed—it’s the system around it.

Shaping alignment isn’t just ML science. It’s design, ethics, team dynamics, and long-game governance. Without those layers, alignment stays theoretical, not structural.

Post reply on HN