Project Glasswing’s AI-driven vulnerability discovery has generated massive volumes of findings, but human triage bottlenecks and a surge in new attack surfaces have left DevSecOps teams struggling to patch the real threats, according to recent analysis and security experts.
Anthropic’s Project Glasswing and the Mythos Discovery Volume
When Anthropic introduced Project Glasswing in April, the tech industry braced for an influx of automated vulnerability reports. Five months later, DevSecOps teams are wading through the glut of discoveries to find out what is actually worth patching, according to Patrick Garrity, a security researcher at vulnerability prioritization tool maker Vulncheck. When Anthropic updated its Vulnerability Disclosure ledger for the first time in August, only a small percentage of initial Mythos vulnerability findings made it into the ledger, and fewer than 1% of the discoveries were marked as fixed.
Out of 1,061 vulnerabilities attributed to AI-assisted discovery, only 14 had been listed as exploited in the wild as of June 2026, marking roughly 1%. “You discover a vulnerability — that doesn’t mean it’s necessarily useful for an attacker. It doesn’t mean it’s necessarily going to be exploitable with certain conditions,” Garrity said. Garrity’s analysis showed that Anthropic reported 26,153 total findings; 2,736, or 10.5%, reached the ledger, 2,096, or 8%, were reported to maintainers, and 202, or 0.8%, were marked fixed, while 245 were withdrawn.

The Human Triage Bottleneck in DevSecOps Workflows
The stark discrepancy between the sheer volume of AI-generated findings and the low rate at which they are patched points directly to a major bottleneck during the human triage step. Software security experts note that teams must manually determine what findings matter, how severe they are, whether they warrant a fix, and how to remediate the vulnerability if necessary. “There was so much agita around [the] ‘Vulnpocalypse’, around everybody being able to find new zero-day vulnerabilities and exploit them, the impact that was going to have,” said Neil Carpenter, an independent security evangelist. “And then the secondary piece of, ‘How do we fix so many vulnerabilities?’ It becomes a burden on defenders, both developers and blue teams downstream.”
Claude concluded that 91.5% of findings were of high or critical severity, but human maintainers assessing the exact same group found only 51.3% met that threshold. For instance, Daniel Stenberg, founder and lead developer of the Linux curl command-line utility project, wrote in a blog post in May that Mythos identified five confirmed vulnerabilities, which his team boiled down to just one. The other four included three false positives detailing shortcomings already documented in API documentation and a fourth deemed just a bug. Garrity noted that the Claude team might not have possessed the specific subject domain expertise required to guide Mythos in accurately assessing vulnerabilities, resulting in a lot of noise.
Deploying AI Frameworks and Expanding the Attack Surface
While organizations attempt to use AI to parse piled-up vulnerabilities, security experts emphasize that this does not remove the human from the loop. Michele Chubirka, a senior principal security architect who participated in Project Glasswing at Red Hat, and platform engineers built a workflow engine described as “part harness, part integration of deterministic security tools” to scan repositories under the vendor’s OpenShift product organization. “People think that AI is magic, that a frontier model is going to magically do this stuff for you. It isn’t,” Chubirka said, speaking in a personal capacity. Chubirka’s team used AI to create patches after triage, but out-of-the-box results varied: 60% were generally good, 20% deviated from design principles, and another 20% required complete rework.
Beyond triage challenges, deploying AI tools in production introduces entirely new security risks and bugs of its own. “No one is talking about… the risks [that arise] when you deploy these technologies,” Garrity said. “The bigger shift I’m seeing is actually the attack surface expanding substantially. AI products themselves are being targeted and hit very fast.” Recent examples of remote code execution in AI frameworks include Langflow, Bifrost, and the MCP Python SDK. Part of the danger involves granting AI tools access to sensitive information and integrating them into infrastructure faster than governance teams can secure them, effectively bypassing the concept of least-privilege access.

Frequently Asked Questions About AI Vulnerability Discovery
What is Project Glasswing?
Introduced by Anthropic in April, Project Glasswing utilizes AI models like Mythos to discover software vulnerabilities across open-source projects and enterprise repositories.
Why are so few AI-discovered vulnerabilities patched?
Independent human triage is a rate-limiting step. Security teams must manually review massive volumes of findings, weed out false positives, and verify actual exploitability before applying fixes.

Do AI models accurately score vulnerability severity?
Not always. While Claude assessed 91.5% of findings as high or critical severity, human project maintainers evaluated only 51.3% of those same findings as high or critical.
What new risks do AI tools introduce to enterprise infrastructure?
AI tools expand the attack surface by introducing new vulnerabilities—such as remote code execution flaws in frameworks like Langflow, Bifrost, and the MCP Python SDK—often due to overly permissive access controls.
Worth a look