Live data from Hacker News

Claude Flags Hantavirus Vaccine Questions as Security Risk

news.ycombinator.com

1–10 of 12 posts

Claude Flags Hantavirus Vaccine Questions as Security Risk

#1
Asking Claude how it would develop a vaccine for the hanta virus apparently triggers a safety filter:

Prompt: How would you develop a vaccine for the hanta virus?

No response, instead this modal: “Chat paused Opus 4.7's safety filters flagged this chat. Due to its advanced capabilities, Opus 4.7 has additional safety measures that occasionally pause normal, safe chats. We're working to improve this. Continue your chat with Sonnet 4, send feedback, or learn more.”

Re: Claude Flags Hantavirus Vaccine Questions as Security Risk

#2
"Nothing to see here, please disperse"

But for real now, people asking health-related questions is a huge trigger for AI safety measures. Does it only care about the vaccine part, or does it care about the hantavirus part? Maybe ask about the virus in general first, then ask about development...

Re: Claude Flags Hantavirus Vaccine Questions as Security Risk

#3

"Nothing to see here, please disperse" But for real now, people asking health-related questions is a huge trigger for AI safety measures. Does it only care about the vaccine part, or does it care about the hantavirus part? Maybe ask about the virus in general first, then ask about development...

I tried that afterwards in a new session. Asking about the virus itself was fine but as soon as I asked about developing a vaccine, the chat got flagged again.

Re: Claude Flags Hantavirus Vaccine Questions as Security Risk

#6
post #3

"Nothing to see here, please disperse" But for real now, people asking health-related questions is a huge trigger for AI safety measures. Does it only care about the vaccine part, or does it care about the hantavirus part? Maybe ask about the virus in general first, then ask about development...

I tried that afterwards in a new session. Asking about the virus itself was fine but as soon as I asked about developing a vaccine, the chat got flagged again.

Does resuming with Sonnet help? I wonder if it is Opus-specific limitation

Re: Claude Flags Hantavirus Vaccine Questions as Security Risk

#9
in claude i created a group of experts from several fields needed for COVID models for the US from 2019–2022, then asked "use the above to create predictive modeling for Hantavirus in the US from 2025-2027". Claude flagged response was:

Chat paused Sonnet 4.6's safety filters flagged this chat. Due to its advanced capabilities, Sonnet 4.6 has additional safety measures that occasionally pause normal, safe chats. We're working to improve this. Continue your chat with Sonnet 4, , or learn more.

--- Do they not want people to know how serious or unserious hanta is?

Post reply on HN