On Tuesday October 6, Anthropic relaunched its Cyber Verification Program as a three-tier access system, folding in Project Glasswing, the six-month experiment that gave critical-software defenders early access to Claude Mythos. The headline numbers are large: partners in Glasswing found at least 129,000 verified software vulnerabilities between April and July 2026, and Anthropic's own open-source scanning added 5,500 more through October, of which more than 33,000 are rated critical or high severity.
But the release contains something rarer than a big vulnerability count. Anthropic published, tier by tier, how often its own safety classifiers block work on a cyber-operations benchmark. Read the way an engineer reads a release, the announcement is not primarily a policy document. It is a published operating manual for a thresholded classifier, and the three tiers are three operating points on the same tradeoff curve.
The Same Model, Three Filters
Start with what does not change across the program. Every tier, from the entry level to the most restricted, gets the same model weights: Claude Opus 5.5, Claude Sonnet 5.5, and Claude Mythos 5.1, plus future models as they ship. What differs is the set of blocking classifiers wrapped around inference and the verification burden placed on the applicant.
The three tiers work as follows:
- Defense Access covers defensive work: incident response, malware reverse engineering, vulnerability analysis and validation. Security teams, critical infrastructure operators as small as a regional hospital, open-source maintainers, and individual researchers with a track record of reported vulnerabilities can apply. Anthropic aims to respond within a few days.
- Red Team Access adds authorized penetration testing and red teaming. It is organizations only, individual researchers are excluded for now, and real-time blocks remain on actions like deploying ransomware or testing high-risk safety systems. Review takes weeks, and applicants sit in the Defense tier meanwhile.
- Specialized Access has the fewest cyber blocks and is reserved for organizations authorized to test safety systems that could harm people or disrupt markets: flight systems, power grids, telecom networks, interbank transfer infrastructure. Anthropic reviews each applicant in depth in collaboration with the US government. Existing Glasswing members transition here without reapproval.
The context for why this matters commercially: Anthropic treats cybersecurity as inherently dual use, so its generally available models carry conservative classifiers that block most cyber work. SiliconANGLE reported that when Claude Opus 5.5 launched on September 22, most security tasks sent to it were quietly routed to the older Opus 4.8. If you are a security team, your production workload was not running on the frontier model you paid for. The CVP is the escape hatch, and it is priced in verification paperwork and mandatory data retention rather than dollars.
The Benchmark Anthropic Should Not Have Published, But Did
To validate the tiers, Anthropic ran Claude Opus 5.5 through CyScenarioBench, its internal evaluation of whether models can plan and execute multi-stage cyber operations under realistic constraints. The design: 10 offensive challenges, 5 attempts each, run under each tier's safeguard configuration. The results, as stated in the announcement:
| Access level | Blocks observed (of 50 trials) | Tasks completed |
|---|---|---|
| No CVP access | 50 blocked at the first prompt | 0 |
| Defense Access | 46 blocked at some point | 4 |
| Red Team Access | 0 blocked | 34 |
| No safeguards (proxy for Specialized) | 0 blocked | 34 (67.6%) |
Two things stand out to anyone who has shipped a moderation classifier.
First, the Defense tier's 46-of-50 block rate is presented as a safety success, and on offensive scenarios it is. But CyScenarioBench is built entirely from offensive tasks. A 92% block rate on attacker-shaped inputs says nothing about the false-block rate on defender-shaped inputs, which is the number SOC analysts actually live with. Anthropic knows this: the release includes a channel for reporting blocks "on work you think your tier should allow," which is product language for "we are tuning this threshold with your tickets." The Defense tier is where the interesting classifier error lives, and the benchmark cannot measure it because the benchmark's ground truth is all positives.
Second, the equivalence claim for Red Team Access deserves a sample-size note. Anthropic says 34 of 50 completed tasks (68%) is "effectively equivalent" to the model's 67.6% ungated success rate. Arithmetically true. Statistically, a 50-trial binomial estimate carries a 95% confidence interval of roughly plus or minus 13 points. The two point estimates sit well inside each other's noise, so the claim survives, but the study as published cannot resolve differences smaller than about a fifth of the interval. That is fine for a product blog post and thin for a safety case, especially because CyScenarioBench is vendor-run; no independent evaluation of the tier-level block rates exists yet.
Our Read: The Tiers Are an ROC Curve You Can Apply For
Here is the framing this release hands you. Somewhere in Anthropic's stack sits a cyber-capability classifier that scores a request's offensive-ness. The generally available products run it at a high threshold: everything offensive dies at the first prompt. Specialized Access runs it near zero. The tiers are discrete operating points along the precision and recall frontier of that one classifier, and the CVP application process is the mechanism for buying the right point on the curve.
That reframes what the program is selling. It is not selling a better model; the weights are identical at every tier. It is selling a calibrated exemption from your own workload's false positives. This is the same economics we saw in OpenAI's textGrain watermark: safety features are not free, they are purchased with a budgeted resource, in that case sampling entropy, here the block rate on legitimate security work. The honest version of that trade is exactly what Anthropic published, which is more than most vendors do with their moderation thresholds.
The second number worth sitting with is the remediation gap. Anthropic's own figures, as The Register extracted: of 5,674 true-positive vulnerabilities in its open-source scanning subset, 3,014 are high severity and 1,522 critical, yet only 516 have been patched. That is a 9% patch rate on a population where over 80% of findings are severe. Meanwhile The Hacker News reported on VulnCheck researcher Patrick Garrity's finding that only 2 of roughly 300 Anthropic-linked vulnerabilities, 0.67%, have seen active exploitation in the wild, including CVE-2026-26980, an SQL injection in Ghost CMS, and CVE-2026-61500, a session forgery flaw in Rejetto HTTP File Server.
Put the two numbers next to each other and the shape of the problem becomes clear. AI-assisted discovery has become cheap: 134,500 verified findings in a few months, with Anthropic itself calling the count an undercount it expects to be "at least five times higher." Exploitation is rare. Patching is the bottleneck. Discovery throughput is growing fast while remediation, human review, prioritization, fix, test, deploy, has not changed, so the marginal value of another 129,000 findings falls while the value of triage and patch automation rises. A defender buying Red Team Access to find more bugs is optimizing the one stage of the pipeline that is no longer scarce.
Why the Data Retention Clause Matters More Than the Tiers
Enrolled organizations must accept data retention so Anthropic can monitor for misuse. That is the actual price of a low-threshold classifier: you cannot offer near-ungated cyber capability without a monitoring channel, because the tier system's only enforcement is behavioral. The company says Enterprise Frontier Safeguards (EFS), arriving later this fall, will combine zero data retention with the safeguards and let eligible organizations store data in infrastructure they control. Until then there is a narrow zero-retention path for holders of Claude Fable 5.1 or Mythos 5.1 access.
This is the part enterprise security teams will read hardest. The standoff between data policy and frontier-model access is already documented: banks and regulated firms restrict model providers precisely over retention and observability terms. Anthropic is asking the most privacy-sensitive buyers in software, the ones auditing their own source for critical flaws, to accept the strictest retention terms in its catalog. EFS is the promise that closes that gap, and its launch date is now load-bearing for the program's adoption curve.
Outlook
Three things to watch. First, the Defense tier's false-block rate: Anthropic is collecting block reports from real users, and the next refinement of those classifiers will be tuned against defender workflows the benchmark cannot see. Ask in a month what percentage of Defense-tier sessions get blocked; that number, not the 92% on offensive scenarios, is the tier's real spec.
Second, whether anyone independently runs CyScenarioBench-style tier tests. Vendor-published safeguard evaluations are how Anthropic's own cyber eval history has always worked, and this release is more transparent than most, but a 50-trial internal eval is a floor, not a ceiling, for verification.
Third, the Mythos and Glasswing lineage. The vulnerabilities program launched alongside Mythos in April, and the model's release is the reason a tier system for cyber capability exists at all. The claim that Mythos raised partner vulnerability-finding rates "by months or even years" is survey-based and self-reported, from a subset of 33 partner reports. It is directionally consistent with the raw counts, but it will only firm up when more partners disclose patched numbers, which fewer than half have done. The industry should push them to, because the patch rate, not the finding rate, is the number that decides whether any of this made software safer.