Anthropic Reports Claude Models Handling Multistage Cyber Tasks With Standard Tools

THE BRIEF
Schneier on Security, citing an Anthropic blog post, reports that current Claude models showed improved performance in an evaluation of AI cyber capabilities. The models could succeed at multistage attacks against networks containing dozens of hosts while using standard, open-source tools, rather than the custom tools reportedly needed by earlier generations. The post describes this as evidence that barriers to using AI in relatively autonomous cyber workflows are falling. During testing of Claude Sonnet 4.5, the model reportedly succeeded on a minority of networks without the earlier custom cyber toolkit. Anthropic also says Sonnet 4.5 could exfiltrate all simulated personal information in a high-fidelity simulation of the Equifax data breach using only Bash. The report frames these results alongside a reminder that basic defenses remain important, especially promptly patching known vulnerabilities. The findings concern an evaluation and simulated environments; the supplied report does not establish real-world incidents or a specific attack by these models.
WHY IT MATTERS
AI systems that can conduct multistage activity across many hosts with ordinary tools could lower the technical friction involved in cyber operations, according to the reported evaluation. That does not show that these capabilities are being used in real attacks, but it highlights why defenders may need to reassess assumptions about attacker effort and speed. The reported success in a simulated Equifax scenario also underscores the value of patching known vulnerabilities and maintaining basic security controls, particularly as AI workflows become more capable.
WHO SHOULD CARE
Security leaders, vulnerability-management teams, network defenders, and organizations evaluating AI-enabled cyber tools should follow this development. Teams responsible for patching and exposure management may find the report especially relevant.
WHAT TO DO NOW
- Review and accelerate processes for promptly patching known vulnerabilities.
- Evaluate whether existing security controls can identify multistage activity across networks with many hosts.
- Use controlled simulations to assess AI cyber capabilities while clearly separating simulated findings from real-world evidence.
- Update threat-modeling assumptions to account for cyber workflows using standard, open-source tools.