Project Glasswing and the model you are not allowed to study
Anthropic announced on April 7 that Claude Mythos Preview, a model that found thousands of high-severity vulnerabilities across major operating systems and browsers, will go only to vetted security partners. What outside researchers, evaluators and academics lose when the frontier is withheld, and whether the trade is right.
What was announced
On April 7 Anthropic described a model called Claude Mythos Preview and said it will not be released generally. The model is presented as a general-purpose successor in the Opus 4.6 line whose distinguishing property is vulnerability research. Anthropic says it found thousands of high-severity vulnerabilities across every major operating system and web browser, and gives examples. A 27-year-old bug in OpenBSD's TCP SACK handling that crashes the kernel remotely. A 16-year-old FFmpeg flaw that automated fuzzing had exercised five million times without finding. A remote code execution in FreeBSD's NFS that yields root. Chains of Linux kernel race conditions composed into privilege escalation.
The numbers that make the case are comparative. On the CyberGym vulnerability reproduction benchmark Mythos Preview scores 83.1 percent against 66.6 for Opus 4.6. Simon Willison relays a figure from the red team write-up of 181 successful exploit developments against the Firefox JavaScript engine, where Opus 4.6 managed two. Access goes to a partner list that includes AWS, Apple, Broadcom, Cisco, CrowdStrike, Google, JPMorganChase, the Linux Foundation, Microsoft, NVIDIA and Palo Alto Networks, with 100 million dollars of usage credits and four million dollars in donations to open source security bodies, including 2.5 million to Alpha-Omega and OpenSSF and 1.5 million to the Apache Software Foundation.
The case for withholding
Taken at face value the reasoning is straightforward. If a model can find and weaponise bugs faster than all but the most skilled humans, then handing it to everyone hands it to attackers, and the defenders need a head start to patch. Willison, who is not usually generous toward AI vendor restrictions, supports this one as justified. He notes that the flow of vulnerability reports to maintainers like Greg Kroah-Hartman and Daniel Stenberg has already shifted from slop to real findings, which suggests the capability is real rather than marketing.
We find the argument stronger than we wanted to. A vulnerability in a kernel that has sat there for 27 years is a shared liability, and a period where only patchers have the tool is a coherent policy. The credits and donations show Anthropic understands that open source maintainers, who are the people who actually have to fix these bugs, need resources and not just notification.
What the rest of us lose
The partner list contains no universities, no independent evaluation organisations and no interpretability groups outside Anthropic. Anthropic's page does not mention academics or outside researchers at all. That is the cost, and we want to state it concretely rather than as a vague worry about openness.
Third-party evaluators cannot check the capability claims. Every figure in the announcement, including the 83.1 percent and the thousands of vulnerabilities, comes from Anthropic or from partners under agreements with Anthropic. The benchmark scores that the field uses to track frontier capability will have a gap at the top. Interpretability researchers outside the company cannot study a model whose distinguishing property is a capability jump, which is exactly the kind of model you would want to study to understand where the jump lives. And the people who write about model risk for a living will be writing about a system they have not touched, on the basis of a vendor's description of it.
There is a subtler loss for security research itself. The population of people who find bugs in operating systems has always included independents and academics who are not employed by a company on that list. A Cyber Verification Program is promised for legitimate professionals affected by the restrictions, and the details of who qualifies will decide whether this is a delay for the frontier or a permanent reallocation of who gets to do the work.
Is it the right trade
For this model and this capability, probably yes, for a while. The asymmetry between finding a bug and patching it is real and the harm from broad release is plausible and immediate. The harm from restriction is slower and mostly falls on people like us, who can wait.
The part we are uneasy about is the precedent. A rule that says the most capable model goes only to a list of large companies chosen by the developer, justified by a risk assessment the developer produces, is a rule that will be very easy to invoke next time and the time after. It is also a rule that removes the outside check on the assessment itself. If restricted access becomes the default for frontier releases, then the field's ability to know what the frontier can do will depend entirely on vendors telling us, and the last decade of evaluation work has been an argument for why that is not enough.
What we would ask for
A time limit stated up front, after which either the model is released with safeguards or the reasons for extending the restriction are published. A route to access for independent evaluation groups under the same agreements the partners signed, so that the capability claims can be checked by someone who is not a customer. And publication of the vulnerabilities themselves once patched, with enough detail that researchers can study what the model found and what it missed.
None of that undoes the restriction, and we are not asking for it to be undone. We are asking for the withheld frontier to be treated as a temporary and verifiable state rather than a new normal. Whether Anthropic does that over the next few months will tell us more about Glasswing than the announcement did.
Sources
From the foundation