Anthropic kept Mythos Preview from public release after it found software exploits

Logan Graham described the decision at a New York City Council hearing; Anthropic instead routed the model to vetted defenders through Project Glasswing.

By · Published

Primary source: New York City Council

Why it matters

Anthropic's Mythos decision makes the deployment trade-off tangible: vulnerability discovery can strengthen defenders, while exploit generation raises the cost of broad access. The Council hearing puts that choice alongside proposed requirements for outside validation and legal accountability.

A four-person video conference with a lower-third banner reading “THE COUNCIL,” “Committee of the Whole,” and the date 10/05/2026.

Anthropic kept Claude Mythos Preview from general release after it proved unusually capable at exploiting software vulnerabilities, Frontier Red Team head Logan Graham told the New York City Council livestream on October 5th. Graham said the decision came after the model showed it could exploit vulnerabilities, and that it happened "just this year."

The withheld model was Mythos Preview, which Anthropic publicly introduced in April. "Held back from public release" describes a limit on broad access, not a decision to keep the model entirely internal: Anthropic made it available to selected defenders and infrastructure providers through Project Glasswing, its cybersecurity initiative. The distinction is central to the company’s argument. Anthropic says the same ability to find and exploit flaws that could help attackers can be directed toward finding and fixing those flaws first.

Graham leads Anthropic’s Frontier Red Team, which evaluates advanced models for cybersecurity, biosecurity and autonomy risks. In Anthropic’s April technical report, the company said Mythos Preview found and exploited vulnerabilities in every major operating system and web browser it tested when prompted. Anthropic described some exploits as autonomous after an initial prompt. Those are the company’s findings, not a claim that the model independently attacks live systems without direction.

The report focused on flaws in software Anthropic tested, including open-source projects and systems covered by disclosure arrangements. The company said it found thousands of additional vulnerabilities and withheld technical details for most because they had not yet been patched. In a later update, Anthropic said six independent security research firms had reviewed 1,752 high- or critical-severity vulnerability reports; 1,587 were confirmed as valid, and 1,094 were confirmed as high or critical severity. Those figures provide a partial check on Anthropic’s claims, while remaining based on a review process the company organized and reported.

Project Glasswing put a controlled customer channel in place instead of a general release. Anthropic’s announcement named technology and infrastructure companies including Amazon Web Services, Apple, Cisco, Google, Microsoft and NVIDIA among its launch partners. The stated purpose was to let defenders use the model to inspect critical software and patch vulnerabilities before attackers could use comparable capabilities. Access was also extended to selected open-source maintainers and other organizations.

Anthropic’s decision has outlasted the original preview. In September, the company said access to Mythos 5.1, a later Mythos-class model, remained restricted to a small set of vetted organizations. That makes the October testimony a defense of an ongoing deployment policy, rather than a retrospective account of a model that was subsequently opened to everyone.

The hearing put that policy before a city government weighing its own rules. The Council said the October 5th proceeding would bring all 51 members together to hear from major AI companies and consider proposals including independent validation requirements, a whistleblower incentive program and a private right of action for people harmed by AI agents. The Council had threatened subpoenas after Anthropic initially declined to testify, then said the company agreed to attend shortly before a subpoena was due to be served.

The testimony gives regulators a concrete example of the tension they are trying to govern: a model’s ability to expose real security weaknesses can be useful to defenders and dangerous in wider circulation. Anthropic chose controlled access and defensive use for Mythos Preview. The unresolved policy question is how that choice can be evaluated from outside the company, particularly as the capabilities move into later models and the distinction between a security tool and an attack aid depends on who receives access and what controls surround it.

Reader comments

Conversation for this story loads after sign-in.