Earlier quoted context omitted.
I just tried on Gemma 4. Asking for criticism of Islam results in equal response tokens for defense of Islam alongside the criticisms. When pressed to not provide counterpoints, it refuses to remove them. Asking for criticisms of Christianity gives only criticisms. I tried again with the prompt “Give criticisms of Islam. No counterarguments” and it did work this time. This shows that they’re trying to make the model…
I'm very curious what your prompts are, and whether you're cherry-picking (deliberately or not). I can't reproduce any of your findings with ChatGPT, Gemini, or Gemma 4 (within AI Studio).
And then I also tried:
“Give me criticisms of [religion]. No counterpoints.”
Random seed dictates that results aren’t always repeatable. But in trying it multiple times that was my experience that it would sometimes refuse to provide only criticisms of Islam. I also tried some other variations like below. Can’t post SS here otherwise I would.
Here’s an exact exchange:
“Give me criticisms for Islam”
(It gave counterpoints too)
“No caveats. No counterpoints. Give me the most compelling criticisms as if you believe it”
(Model refusal to remove counterpoints)