We used sparse autoencoders to explain LLM moderation flags of violent threats #1 Post by karinemellata » Mon, Apr 21, 2025, 8:11 PM UTC We used sparse autoencoders to explain LLM moderation flags of violent threatsvariance.co