Anthropic, the US-based AI developer preparing for an initial public offering this autumn, says it has identified and blocked attempts to use its large language models for malicious ends — including cyberattacks, surveillance and research that “could have led to biological weapons,” the company wrote in a new report.
What the company found
The report, Anthropic’s third since March 2025, covers activity detected between December 2025 and August 2026. It describes attempts by a range of actors — from spyware vendors and politically motivated individuals to groups the company describes as state-sponsored — to steer the models toward harmful outputs.
- Types of misuse the company cites include cyberattack tool development, enhanced surveillance techniques and biological-research prompts that could be repurposed for weapons development.
- Company response: Anthropic says it has implemented stronger safeguards in its most recent models to limit access to or generation of such material.
- Transparency steps: The firm published examples of malicious prompts and code snippets and urged other AI developers and governments to identify similar abuse.
“We’re publishing this work because we believe we have a responsibility to disclose malicious misuse of our services. As models become increasingly capable, their risks will increase, unless AI developers and society’s defenders act to make them safer,” the company said.
Context and consequences
The report arrives amid mounting scrutiny of how quickly advanced AI capabilities are being deployed and whether industry safeguards keep pace. Anthropic frames the disclosures as part of a defensive posture: identifying misuse, deterring repeat offenders and spurring broader mitigation across the sector. The company also plans an initial public offering later this year, adding commercial stakes to the debate over public disclosure and risk management.
Not all reactions within the company are aligned. The report was published two days after a researcher resigned, publicly warning that Anthropic and its rivals were not moving cautiously enough. That resignation echoes wider concerns among policy-makers, security officials and some researchers that highly capable models could be repurposed by relatively unsophisticated actors to create new threats.
| Item | Detail reported |
|---|---|
| Report count | Third since March 2025 |
| Detection window | Dec 2025 – Aug 2026 |
| Actor types | Spyware vendors, politically motivated individuals, state-sponsored groups |
For policy-makers the report underscores two practical challenges: how to monitor and curb malicious use of commercially deployed AI systems, and how to coordinate disclosure so that mitigations are effective without exposing defensive techniques or model vulnerabilities. Anthropic’s own publication of example prompts and code snippets is a bid to accelerate shared defences, but it also raises questions about how much to reveal publicly.
The company’s statement and the reported incidents add to a growing record that regulators and national security agencies will likely examine as they consider whether new rules, incident-reporting requirements or technical standards are needed to manage AI risks at scale.
InfoRadar will follow developments as governments and industry respond to Anthropic’s disclosures and as the company proceeds toward its planned public listing.