Artificial intelligence models are increasingly being used to block cyberattacks, surveillance, and biological weapons research, according to a report released Thursday by Anthropic. As AI models grow more powerful, elaborate cyberattacks no longer require sophisticated skills, allowing lone individuals to create threats that were impossible a year ago, the company stated. Anthropic announced it added stronger safeguards in its latest models to restrict biological research that could be used to create weapons.
Anthropic Details Novel Misuse and Safeguard Failures
The third report on AI misuse published by Anthropic since March 2025 details non-typical threat activity identified by the startup, which is planning an initial public offering this fall. According to Anthropic, the report includes snippets of malicious code and AI prompts found on its systems, and urges governments and competitors to identify similar abuse.
“The cases we share here aren’t typical misuse, but rather examples of the most notable and novel threat activity we’ve identified to date,” Anthropic stated in the report. “We’re publishing this work because we believe we have a responsibility to disclose malicious misuse of our services. As models become increasingly capable, their risks will increase, unless AI developers and society’s defenders act to make them safer.”
Between December 2025 and August 2026, researchers found misuse by spyware vendors, politically motivated individuals, and state-sponsored groups spreading propaganda. In one instance, systems blocked a request for Claude’s assistance in authoring a grant application for scientific funding concerning the chikungunya virus. According to the report, the work involved gain-of-function research aimed at the mosquito-borne virus’s transmissibility and immune evasion properties to make the pathogen more dangerous.
Model Capabilities and Evolving Biological Research Restrictions
Anthropic stated that none of the cases in its report involved newer, more powerful Claude Fable or Mythos-class models, except for one illicit distillation case described as an industrial-scale, covert campaign to extract and replicate a model’s capabilities without authorization. Based on the firm’s assessments, releases from 2025 like Claude Sonnet 4.5 and Claude Opus 4 fell well short of the capability needed to provide meaningful help to advanced actors engaged in perilous biological studies.
“As a result, safeguards on these models were less stringent, directed mostly at preventing access to content that might uplift novices in recreating known bioweapons,” the report noted. “But for today’s models — which are capable of assisting in a range of complex scientific research tasks — the evidence is no longer certain, and we cannot make that same assurance.”
Because of this shift, Anthropic applied stronger safeguards restricting access to a wide range of dual-use biological research queries in newer models like Claude Fable 5. Cornell University computer science assistant professor John Thickstun noted that it places organizations like OpenAI and Anthropic in an awkward position, forcing them to arbitrate what constitutes safe versus unsafe actions and execute broad societal value judgments absent any formal democratic oversight or public deliberation.
Social Media Influence Operations and Researcher Resignation
Anthropic also identified groups that created hundreds of social media accounts posing as ordinary people to amplify specific political views over a week. The company outlined nine such cases originating in Russia, Iran, Turkey, the Persian Gulf, South Asia, Africa, and Europe.
The report followed the resignation of Anthropic researcher Jacob Coxon, who announced he is leaving over concerns that Anthropic and OpenAI are racing toward self-improving superintelligence and gambling with human lives. Coxon warned that some colleagues believe AI could threaten human life by the end of the decade.
Despite these challenges, Anthropic reported it has blocked each identified malicious activity, strengthened its safeguards, and shared information with government authorities and industry partners.
Did you know? Anthropic’s report outlines instances where bad actors attempted to use AI models for gain-of-function research on the chikungunya virus to enhance its transmissibility and immune evasion properties.
Frequently Asked Questions
What types of malicious activity did Anthropic block?
Anthropic blocked efforts to use its AI models for cyberattacks, surveillance, state-sponsored propaganda operations, and research that could lead to biological weapons.
Why did Anthropic apply stronger safeguards to its newer models?
Older models lacked the capability to assist in complex scientific research, but today’s models can assist with complex tasks, prompting Anthropic to restrict dual-use biological research queries.
How can I stay updated on AI safety developments?
Subscribe to our newsletter and explore our archives for the latest updates on artificial intelligence regulation, industry reports, and technological governance.
Related reading