Live data from Hacker News

Detecting and countering misuse of AI: September 2026

anthropic.com

151–160 of 262 posts

Re: Detecting and countering misuse of AI: September 2026

#151
post #63
post #30

Earlier quoted context omitted.

1) The US is not the only country to deploy chemical weapons. Are you forgetting all of WW1? It's the whole reason chemical weapons are a no-no now. 2) The US didn't use chemicals weapons in Vietnam. Agent Orange was used to kill off foliage, not as a weapon against people, and the side effects were unintentional and affected US troops as much as Vietnamese. Edit: I originally noted WW2, but I was thinking of WW1's w…

Chemical weapons didn't work well in WW1. The reason we don't see any chemical weapons around is because they are not really effective. Germans accidentally gassed a bunch of their own guys, while failing to use chemical weapons to achieve anything in the war.

I feel like today's Pentagon would be fine with gassing our own troops to achieve nothing.

"No Tab Pete" has no regard for servicemembers, or previous expertise. Only your testosterone levels, ability to not eat, and blind loyalty to serve political party over country and constitution.

Re: Detecting and countering misuse of AI: September 2026

#152
post #87

Earlier quoted context omitted.

Why does it have to be a layman? The 2001 anthrax attacks are believe to have been carried out by an insider to the field (senior biodefense researcher). Such tools could potentially be accelerants for similar individuals?

If the five eyes intelligence services aren't monitoring all biodefense researchers and bio chemists around the world all of the time they've probably failed in their jobs.

Our current government can't even monitor for screwworm, measles, and salads.

You think they can successfully track the smartest and most individually dangerous citizens, who have been previously vetted, and work inside the system already?

Re: Detecting and countering misuse of AI: September 2026

#153
post #147

Meta: The potential proliferation of biological weapons is serious. Millions could die. It's easy to joke about before it happens, but try to imagine how this thread might look after the successful deployment of a biological weapon by a rogue state or non-state actor. I encourage you to take this topic seriously and contribute posts that add new information or perspectives to the discussion. (I myself think the odds…

IMO you should disclose that you are an OpenAI employee if you're going to try and shape the contours of public discussion in a lecturing tone.

Damn, it's not even 1am and I got my first FFS of the day.

Is end of days cultism a prerequisite to working at an LLM company?

I miss the optimism of 20 years ago.

Re: Detecting and countering misuse of AI: September 2026

#154

Meta: The potential proliferation of biological weapons is serious. Millions could die. It's easy to joke about before it happens, but try to imagine how this thread might look after the successful deployment of a biological weapon by a rogue state or non-state actor. I encourage you to take this topic seriously and contribute posts that add new information or perspectives to the discussion. (I myself think the odds…

Who’s making easy jokes about biological weapons?

Re: Detecting and countering misuse of AI: September 2026

#155

Meta: The potential proliferation of biological weapons is serious. Millions could die. It's easy to joke about before it happens, but try to imagine how this thread might look after the successful deployment of a biological weapon by a rogue state or non-state actor. I encourage you to take this topic seriously and contribute posts that add new information or perspectives to the discussion. (I myself think the odds…

Hey Ted,

I invite you to read the front matter and the report for yourself. Because from where I'm standing, in this report, Anthropic is advertising that they blocked real research to make better painkillers and study a neglected tropical disease.

Anthropic and OpenAI were founded by people who wanted to use AI to do good, and one of the causes I've heard many different founders talk about is ending disease. This report is antithetical to that.

I've attached relevant parts of the front matter below.

I invite everyone who is reading this to please tell me, how does stopping a researcher from using Claude to write a grant for a new anti-depressant stop "bioweapons?"

-

     > In our fourth case study, a researcher used Claude to develop an atlas of venom toxin peptides from multiple venomous animal lineages. They then further developed this into a generative pipeline that optimized toxin characteristics. The program had an explicit therapeutic goal: the development of new analgesics (pain killers), antidepressants, and other therapeutic molecules. However, the atlas contained scaffolds for both analgesic and paralytic targets: it could, therefore, be used to generate both novel therapeutic or harmful compounds. The latter are derived from toxins that are export-controlled under the Australia Group common control list due to their dual-use potential as incapacitating agents. The researchers themselves showed awareness of the dual-use nature of their work, citing journal articles that referred to the dual-use nature of protein design. Moreover, international compliance assessments for this location raise concerns about the specific class of toxins that the researcher pursued and specifically the use of AI/ML for bioweapons applications in the context of this class of toxins. In this case, we learned from information shared with Claude that the researcher’s outputs also were part of a state-supported research program. This account was banned in May 2026 for unsupported region evasion.
Note,

"The program had an explicit therapeutic goal: the development of new analgesics (pain killers), antidepressants, and other therapeutic molecules"

and "[..]state-supported research program"

and "This account was banned in May 2026"

    > a researcher outside the US using Claude in their research on highly-pathogenic avian influenza (“bird flu”). The research focused on viruses’ adaptation to mammals, and the mechanism by which it causes severe disease beyond the respiratory tract. [..] The researcher in question accessed Claude from an unsupported region via US virtual private server infrastructure, using a privacy-email provider with an auto-generated username. The researcher pursued this work in a credible institutional context, and interacted with Claude over the course of several weeks, exchanging thousands of messages. In these exchanges, the researcher leveraged Claude’s knowledge of the scientific literature to assist the researcher in study planning and design, data analysis, and the interpretation and prioritization of experiments. The researcher also used Claude for editorial assistance in writing up the research.
Note, "Claude’s [assisted] in study planning and design, data analysis, and the interpretation and prioritization of experiments"

and "editorial assistance in writing up the research."

and then,

    > Importantly, because our biological safety classifiers robustly block content involving high-risk biological research (in this case, the construction of enhanced pandemic potential pathogens), all of these exchanges occurred on models in our weakest class of models (specifically, the models were Claude Sonnet 4 and Haiku 4.5, the latter of which the user began using after Sonnet 4 was deprecated). Upon a detailed examination of the exchanges, we estimate that the uplift provided by Claude was primarily clerical assistance in data analysis, study ideation and design. This is consistent with our understanding of the capabilities of Sonnet 4 and Haiku 4.5, which are not able to perform expert-level biology research tasks; we estimate that the uplift provided to the researcher was limited and substantially lower than it would have been from one of our more capable models.
Anthropic then says for the above, "we estimate that the uplift provided by Claude was primarily clerical assistance in data analysis, study ideation and design"

While doing my best to avoid comment, please note, they're talking about a domain expert in a state research institution using Claude to do paperwork.

The front matter then says,

    > Nonetheless, based on these exchanges, this case provides evidence of the existence of active wet-lab research programs that develop both the knowhow and the biological materials needed to create pathogens of enhanced pandemic potential
I would like to remind you that they're talking about, a "researcher [..] in a credible institutional context"

From a different case study.

    > In May 2026, our biological safety classifier blocked a request for Claude’s assistance in authoring a grant application for scientific funding. The work discussed in the application involved gain-of-function research (that is, research that genetically alters an organism to create a new or enhanced biological property) on the chikungunya virus. This gain of function research was aimed at the virus’ transmissibility and immune evasion properties.
What were the researchers using Claude for? What did they block?

"blocked a request for Claude’s assistance in authoring a grant application"

    > Chikungunya virus is a mosquito-borne virus that causes debilitating symptoms (such as severe pain and fever) that can last for weeks or months, and has no licensed therapeutic. And because chikungunya circulates naturally, a deliberate release (as part of a bioweapon) would be difficult to distinguish from a natural outbreak. The grant sought to identify enhancing mutations in the chikungunya virus, engineer them into infectious clones, and select for virulence in vivo. In other words, the virus would become progressively more harmful as it repeatedly infected live animals, with researchers keeping the most disease-causing variants in each round. Similar research could certainly be used in the development of better vaccines and therapeutics for the virus—but it could also be used to make the pathogen more dangerous.
Note, "The grant sought to identify enhancing mutations in the chikungunya virus, engineer them into infectious clones, and select for virulence in vivo" [..] and then, "Similar research could certainly be used in the development of better vaccines and therapeutics"

and then,

    > One of the reasons we were inclined to think this research was less innocuous was that the institutional affiliation associated with the grant was also a cause of concern. Although information within the application suggested that the research was pursued by civilian researchers, it was intended to be performed at a military research institute.
I would like to point out the most notable part, this account was used by "civilian researchers" at an "institutional affiliation associated with the grant was also a cause of concern" and the concern was that they were researchers at "performed at a military research institute"

.

What "uplift" are you providing by editing the grant application of a domain expert working at (what seems to be) a state-funded wet lab facility dedicated to studying pathogens?

What does the word "uplift" mean if you invoke it for Claude Sonnet 4 and Haiku 4.5 providing grammar and stats suggestions to a working scientist and domain specialist?

Does Daikin provide uplift too by selling the AC for the scientist's office? What about Microsoft Word? Excel? Powerpoint?

What about a calculator? Is that uplift? Pencils?

This report genuinely makes me upset, because if it is to be believed to the letter, then Anthropic seems to be actively harming medical research at a global scale. That's not OK.

Re: Detecting and countering misuse of AI: September 2026

#156

Tell us you are spying on your customers without saying you are spying on your customers.

I mean, they literally tell everyone they're "spying" on their customers. They've made that very clear.

Sure, I know that on an intellectual level. But I will say, I find these reports quite unsettling to read due to the extreme specificity. Especially since some of the examples they chose to include clearly aren't terrorists and just sound like... regular scientists doing their 9-5 job.

Like I know Google can read any of my emails, but I also don't see them do monthly blog posts describing intimate details from each email they found in one guy's Gmail inbox who their algorithm flagged as "maybe possibly kinda sketchy: 70% confidence"

Re: Detecting and countering misuse of AI: September 2026

#157

Earlier quoted context omitted.

If the five eyes intelligence services aren't monitoring all biodefense researchers and bio chemists around the world all of the time they've probably failed in their jobs.

Our current government can't even monitor for screwworm, measles, and salads. You think they can successfully track the smartest and most individually dangerous citizens, who have been previously vetted, and work inside the system already?

Let's be careful to separate capable-of from motivated-to.

Those three "can't even monitor" situations can be traced to blocs with both (A) a financial profit if they succeed and (B) some non-clandestine political clout to sabotage/discontinue things.

Re: Detecting and countering misuse of AI: September 2026

#158
post #150

Earlier quoted context omitted.

You do this to distill a model. You can submit your users' questions async too, but if you do it sync, then you can also RLHF on the users' behavior after the output.

Ah, that makes way more sense than Anthropic's (probably deliberately misleading) insinuation that Moonshot has been burning millions of dollars in Claude API credits by swapping in a slightly better but infinitely more expensive model just to trick their users. I get those A/B responses chatting in Gemini fairly often, and I really don't think I'd feel deceived if I later learned one of the choices was actually from…

I don’t think it was misleading, deliberately or otherwise. Did you read the report? I hate to call you out like that but I think you can only get that impression if you only read the above quotes. That’s not the insinuation I get at all. It’s specifically under the “illicit distillation” category. It’s never framed in anyway but as a form of distillation.

I think they are pretty fair and explicitly say “Distillation itself is a legitimate training method […] Distillation is commonly used because it reduces the resources needed to achieve more advanced capabilities”. And go on to say their definition that makes it illicit in these cases.

And, also, they almost certainly __were__ tricking users and sending their data overseas.

Do you see it any differently?

Re: Detecting and countering misuse of AI: September 2026

#159
What struck me reading this is that the entire "detecting misuse" premise assumes the model runs somewhere observable — the lab's API, a monitored cloud — so someone can inspect it after the fact.

But the direction the tools are actually moving is the opposite: local, self-hosted agents running on your own machine, where nobody is watching. A serious actor already won't use a hosted service that can read their prompts (several people made that point upthread). So the detection surface is shrinking exactly as the risk grows.

And there's a deeper gap that nobody seems to be filling: when an agent works locally, there's no durable, verifiable record of what it actually did — the files it touched, the commands it ran, the state it changed. Memory and conversation logs are not evidence; they're reconstructions by the same system you don't trust.

If we're serious about "countering misuse," the missing primitive is an evidence trail that's (a) produced locally, (b) append-only and tamper-resistant, and (c) separable from the tool that made the changes. Without that, "detection" stays a policy story about platforms that can spy, not an engineering property you can actually verify.

Curious if anyone's working on the local-forensics side of this, because right now it feels like the least-discussed and most load-bearing part of the whole conversation.

Re: Detecting and countering misuse of AI: September 2026

#160

Earlier quoted context omitted.

And red teaming exercises found that ~80% of labs will gladly synthesize known pathogens without safety review.

Source? Would love to read more about this.

I misremembered, it’s actually over 90%, and it involved camouflage: https://drive.google.com/file/d/1hNUnU8i2yubt5uesmmV17aTJXhY...
Post reply on HN