Earlier quoted context omitted.
> If they really wanted to, all they would have to do is add a one liner to the system prompt for Grok. They tried that, several times. Mechahitler: https://www.npr.org/2025/07/09/nx-s1-5462609/grok-elon-musk-... > "We have improved @Grok significantly," Elon Musk wrote on X last Friday about his platform's integrated artificial intelligence chatbot. "You should notice a difference when you ask Grok questions." > Ind…
> Mechahitler: https://www.npr.org/2025/07/09/nx-s1-5462609/grok-elon-musk- ... Has anyone done a more technical write-up on this? I find it fascinating but have never really understood what exactly happened. Is this a case of the weights being bad or lack of "safety guardrails" around interacting with untrusted (i.e.: user posts on twitter) input? That is, speaking as someone evaluating grok simply as a tool, a lack…
https://arxiv.org/abs/2502.17424
Essentially if you misalign a model in one area, say opinions on left wing people, it can start exhibiting misaligned behavior in other areas, like calling itself MechaHitler.