A month of real-world testing by leading companies is giving us a better idea of Anthropic’s Mythos frontier LLM’s capabilities in security. The company’s update, which comes a month after the initial release of Project Glasswing, suggests the model is surfacing real vulnerabilities at scale while also highlighting ongoing challenges around noise and trust in its output.
In total, Anthropic says Mythos has scanned more than 1,000 open source software projects and identified 6,202 bugs it classifies as high or critical severity. That makes the model one of the most prominent examples yet of AI being used to hunt for flaws in live codebases, especially in security-sensitive open source software.
But the headline is not just about volume. It is also about reliability. The latest findings point to a system that can clearly find issues but still produces hallucinations and false positives. Although the false positive rate remains within normal industry levels, Mythos’ ability to uncover multi-step attacks means the time required for investigation can quickly compound.
What the Update Says
According to Anthropic, Mythos passed 28% of the high or critical severity findings, or 1,752 bugs, to six independent security research firms for review. Those firms found a 9.4% false positive rate and confirmed 62.4% of the bugs as genuinely high or critical severity.
Anthropic also said it has so far disclosed 530 of the bugs to open source maintainers and hopes to disclose another 827 as quickly as possible. Of the 530 already reported, 75 have been patched and 65 have received public advisories. Anthropic said the patch rate reflects a broader problem in the security ecosystem, noting that even with a relatively slow disclosure pace, Mythos Preview is adding pressure to an already overloaded system.
One example Anthropic highlighted was a critical WolfSSL vulnerability, CVE-2026-5194, which it rated CVSS 9.1. The company said the bug could allow certificate forgery, underscoring the kind of high-stakes issue the model can surface when it works as intended.
Why the Results Matter
Mythos has generated both excitement and concern because it appears capable of chaining together multiple steps in an attack in ways earlier AI systems could not. That raises the stakes. A model that can move from one weakness to another and assemble a proof of concept is not just a scanner. It is closer to an active security analyst and potentially a powerful offensive tool as well.
That is one reason Anthropic has not released the model publicly. Instead, it has limited access to a small group of companies through Project Glasswing, a controlled program designed to test the model in real security environments. The approach lets Anthropic gather feedback while reducing the risk of the system being misused at scale.




