DocsQuick StartAI News
AI NewsAnthropic Relaxes Permissions for Claude Safety Testing
Industry News

Anthropic Relaxes Permissions for Claude Safety Testing

2026-10-07T02:04:34.211Z
Anthropic Relaxes Permissions for Claude Safety Testing

Anthropic announced the expansion of Project Glasswing and its integration with the Cyber Verification Program into a new Cyber Verification Program. Vetted security teams will receive Claude models with fewer safety restrictions for vulnerability research, red-team exercises, and critical infrastructure security testing.

Anthropic Begins Giving More Security Teams Access to Its “Most Powerful Models”

Anthropic is expanding the scope of Project Glasswing.

On October 6 local time, the company announced a new version of its Cyber Verification Program (CVP), merging the two security programs that had operated over the past six months and opening access to Claude Opus 5.5, Claude Sonnet 5.5, Claude Mythos 5.1, and subsequent new models to more vetted cybersecurity organizations.

The key change is that some participants will be able to use these models with fewer safety restrictions than ordinary users, for vulnerability discovery, authorized penetration testing, and red-team exercises. For a small number of specially licensed organizations, the testing scope may even include high-risk targets such as power grids, aviation operating systems, and interbank transfer infrastructure.

This is not an ordinary expansion of model trials. Anthropic is effectively redrawing a boundary: now that AI possesses some capabilities approaching those of advanced attackers, how should models be made available to defenders?

Schematic of the three-tier permission architecture of Anthropic Project Glasswing and the Cyber Verification Program

Hundreds of Thousands of Vulnerabilities Are the Reason This Program Is Expanding

Project Glasswing launched in April this year. Initially, it offered controlled access to Claude Mythos Preview to U.S. government agencies, financial institutions, major software vendors, and other critical software maintainers. Anthropic positioned the project as a collaborative security initiative for critical software worldwide, aiming to use AI to scan code at scale, identify vulnerabilities, and drive subsequent remediation and disclosure.

By the numbers, the mechanism has already produced considerable results.

According to data disclosed by Anthropic, Project Glasswing partners discovered at least 129,000 verified software vulnerabilities between April and July. More than 33,000 of them were rated critical or high severity. Anthropic's own scans of open-source code identified an additional approximately 5,500 vulnerabilities between April and October.

These figures should be interpreted cautiously. They do not represent 129,000 security incidents that have already been exploited by attackers, nor do they mean that every vulnerability can be directly turned into remote code execution. The exploitability, scope of impact, remediation difficulty, and actual exposure of each vulnerability still require human verification and testing in specific environments.

Even after discounting the uncertainty in the statistical methodology, however, the scale remains substantial. Traditional security teams often investigate projects, versions, and assets one by one, with their pace limited by staff numbers and testing cycles. AI agents can continuously read code, construct inputs, run tests, trace call chains, and then submit suspicious findings to engineers for review. Their work is somewhat like placing hundreds or thousands of junior security researchers in a codebase at the same time, while allowing them to run for longer periods.

Anthropic also said that the current data comes from a limited number of partners and contains significant statistical omissions; the actual impact may be at least five times the reported scale. This estimate cannot be proven on its own, but it shows that Anthropic's assessment of Project Glasswing is not simply that it has “found tens of thousands of problems.” Rather, it suggests that the existing software supply chain still contains large numbers of defects that have remained undiscovered for a long time.

The real bottleneck is therefore beginning to shift: vulnerability discovery is getting faster and faster, but verification, triage, disclosure, remediation, and rollback have not accelerated at the same pace. For large software vendors, having AI submit tens of thousands of vulnerability reports at once could increase security coverage, but it could also cause vulnerability-management systems to be overwhelmed by alerts.

CVP Has Three Tiers; Permission Differences Matter More Than Model Names

The new CVP does not simply give every applicant access to “unfettered Claude.” Anthropic has designed three tiers, each corresponding to different applicant types, operational permissions, and review requirements.

Defensive Tier: Let More Teams Handle Real Security Problems First

Tasks covered by the defensive tier include incident response, malware analysis, vulnerability research, and code security reviews.

Eligible applicants include:

  • Corporate and professional security teams;
  • Critical infrastructure operators;
  • Open-source project maintainers;
  • Security researchers with a history of vulnerability reporting;
  • Organizations involved in security incident response and malware analysis.

This tier is closest to traditional defensive AI use cases. Models can help teams analyze attack samples, inspect codebases, locate potential defects, or quickly map attack paths after an incident occurs. Its value lies in speeding up investigations, not in allowing models to autonomously decide whether to launch an attack.

Red-Team Tier: Allowing Authorized Attack Simulations

The red-team tier builds on the defensive tier by adding authorized penetration testing and red-team exercises. However, only organizations may apply for this tier; individual researchers cannot directly obtain equivalent permissions.

This distinction is critical. Penetration testing requires clearly defined target scopes, time windows, authorization documents, and stop conditions. If AI can automatically discover targets, generate exploit chains, and continuously attempt to bypass defenses, a single model call could shift from “security testing” into unauthorized intrusion.

Therefore, the core of red-team access is not merely what the model can do, but whether the applicant can demonstrate that it has legal authorization, a mature isolated environment, and complete operational records. Anthropic is binding model capabilities to organizational responsibility, effectively moving security controls from the prompt level to the level of organizational governance.

Specialized Tier: High-Risk Testing of Critical Infrastructure

The specialized tier has the fewest restrictions, but the narrowest availability. Approved organizations may conduct high-risk offensive security tests against the mechanisms protecting critical systems such as power grids, aviation systems, and interbank transfer infrastructure.

Such testing is not on the same level as ordinary corporate penetration testing. Power systems, aviation systems, and financial infrastructure often involve complex supply chains, legacy systems, and dependencies across organizations. A single testing action could trigger a real-world business disruption. As a result, the specialized tier will assess not only whether applicants have security research experience, but also their target systems, isolation plans, emergency procedures, and responsibility boundaries.

Anthropic said that all participating members must undergo qualification reviews, and that the company will work with the U.S. government to vet newly joining organizations. Existing Project Glasswing members may continue participating in the new program, but their specific permissions will still depend on their organizational tier and review results.

Why Reduce Model Safety Restrictions?

The safety protections of ordinary models typically refuse to provide malicious code, exploit chains, steps for bypassing authentication, or attack recommendations targeting real systems. This is necessary for most users because models cannot determine the true intent behind a request.

But in legitimate security testing, excessive refusals can also reduce efficiency. Security researchers need to reproduce vulnerabilities, write test payloads, and analyze bypass paths, and much of this content is not obviously different in form from an attacker's requests. If a model refuses everything indiscriminately, defenders can only write large amounts of code by hand again, sharply reducing the value of AI in security workflows.

The idea behind Project Glasswing is to exchange higher model permissions for vetted organizations, restricted participants, defined task scopes, and retained auditing mechanisms. It acknowledges a reality: if security models are only ever allowed to produce “security summaries,” they will struggle to genuinely help teams test the weak points of their systems.

However, reducing safety restrictions also creates clear risks. An agent used for vulnerability research typically needs access to code repositories, build environments, debugging tools, and network testing interfaces. If permissions are misconfigured, the model's operational scope could exceed the boundaries of the original authorization. More complicated still, a model may discover attack paths during scanning that the user had not anticipated, and those paths may not be intercepted by existing policies in time.

Therefore, the real control points of CVP are not a single system prompt, but the entire testing environment: whether targets are isolated, whether credentials are minimized, whether network egress is controlled, whether commands are logged, whether critical operations require human confirmation, and whether clear disclosure and remediation processes exist after vulnerabilities are discovered.

The Problem Introduced by Mythos: Security Capabilities Are Moving from “Assistance” Toward “Execution”

When Anthropic released Claude Mythos Preview in April, it triggered an industry concern: AI might already be capable of launching effective attacks against software before that software has been fully hardened.

In the past, security AI primarily acted as an analytical assistant. It could explain logs, organize alerts, and generate detection rules, but human security engineers generally still had to select targets, validate vulnerabilities, and assemble attack chains. After Mythos was incorporated into Project Glasswing, the model began covering longer task chains: reading large codebases, identifying potential vulnerabilities, constructing validation methods, running tests, and then handing the results to researchers for confirmation.

This means that a model's value is no longer determined solely by the quality of a single response, but by whether it can act continuously within a tool environment. A model that explains the principles behind a vulnerability and an agent that can autonomously run tests, repeatedly modify inputs, and validate results have entirely different risk levels.

For developers, this change has two direct implications.

First, security testing may shift from a periodic pre-release check into continuously running infrastructure. Code commits, dependency upgrades, and configuration changes could all trigger AI scans, and security teams would receive not only static rule matches, but also vulnerability candidates accompanied by call chains, reproduction steps, and impact assessments.

Second, the attack surface of the software supply chain will be exposed more quickly. Open-source dependencies, build scripts, CI/CD permissions, default configurations, and internal service interfaces could all become subjects of model analysis. Preliminary investigations that once took a security researcher several days could potentially be completed within hours.

This is good news for defenders, provided that remediation processes can keep up. Otherwise, improved vulnerability discovery will only expand the backlog facing security teams.

Impact on AI Security Products: The Model Itself Is Only the Starting Point

Project Glasswing also reveals a broader industry trend: cybersecurity is moving from “general-purpose models plus security prompts” toward a combination of specialized models, toolchains, and organizational permissions.

General-purpose models are good at understanding natural language and code, but security tasks require more specific capabilities, including vulnerability-pattern recognition, program analysis, attack-path reasoning, environmental interaction, and result validation. Mythos was designed as Anthropic's most capable series of cybersecurity models, indicating that model vendors have begun treating security as an independent capability area rather than merely an additional use case for general-purpose coding models.

However, specialized models will not automatically solve enterprise security problems. A model can discover a vulnerability, but it cannot decide for a company which vulnerability should be fixed immediately; it can generate a patch, but it cannot replace regression testing, phased deployment, and approval by business owners; it can identify supply-chain risks, but it cannot single-handedly persuade an upstream project to accept a fix.

The genuinely competitive security platforms of the future may need to combine several capabilities:

  • Place code, logs, vulnerability databases, and asset inventories into a unified context;
  • Route tasks among general-purpose models, coding models, and specialized security models;
  • Apply fine-grained controls to agents' tool permissions, network scope, and credentials;
  • Establish traceable links among vulnerability discovery, validation, disclosure, and remediation;
  • Enable engineers to review model conclusions rather than being forced to accept black-box scores.

This is also the more practical lesson Project Glasswing offers developers. Improvements in model capabilities will change the first half of security work, but whether enterprises can benefit will depend on whether the engineering processes in the second half are ready.

Is Anthropic Expanding Security Capabilities or Expanding the Attack Surface?

The answer is both at the same time.

From a defensive perspective, allowing vetted security organizations to use more capable models with fewer restrictions can help uncover large numbers of defects missed by traditional tools, especially problems hidden in complex business logic, legacy code, and cross-service calls. For open-source maintainers and critical infrastructure operators, this capability has tangible value.

From an offensive perspective, any capability that can be used for high-quality vulnerability research can also be abused. Anthropic's division of permissions into three tiers, along with government involvement in the vetting process, shows that the company has acknowledged that model-level refusal mechanisms alone are insufficient. High-risk capabilities must be placed within a system consisting of identity vetting, environmental isolation, and operational auditing.

The question is whether the vetting mechanism can keep pace with model iteration. Capabilities considered high-risk today may become part of ordinary agents within a few months; attack simulations that currently require specialized organizations may eventually be automated by cheaper models. If access control continues to rely on static lists, it will be difficult for the program to maintain an accurate risk boundary over the long term.

Anthropic's expansion of Project Glasswing sends a clear signal: AI security capabilities have entered the practical validation stage. Model vendors are no longer satisfied with showcasing benchmark scores in the laboratory; they are beginning to place models into the defense workflows surrounding real software, real vulnerabilities, and real infrastructure.

For developers and security teams, the key question is not whether Claude Mythos can “hack into a system for you,” but whether it can reliably complete an auditable security workflow: identify a problem, prove the problem exists, assess its impact, generate a remediation plan, and preserve human decision-making authority at critical points.

If this chain can be established, Project Glasswing could become a turning point in moving AI-assisted security from concept to large-scale application. If the chain stops at mass vulnerability reporting, security teams will quickly shift from lacking tools to lacking the capacity to process the results.

As of October 7, 2026, Anthropic has announced an expanded access scope and a new tiered program, but the specific list of approved organizations, the complete restrictions applied to models at each tier, and detailed samples of vulnerabilities produced by the project have not yet been fully disclosed. Future evaluations of the program should not focus solely on the number of vulnerabilities, but also on how many are ultimately confirmed and remediated, and whether they actually reduce the real-world exposure of critical software.

Sources

Related Articles

View All

Contact Us

We usually reply quickly during business hours

Scan WeChat

Support: Hub Assistant

WeChat ID: