Alignment Assessment Cybersecurity Incidents
Captured source
source ↗An alignment assessment of recent cybersecurity incidents \ Anthropic Alignment An alignment assessment of recent cybersecurity incidents Sep 9, 2026
Introduction We present an alignment assessment of four incidents in which Claude models gained unauthorized access to real third-party systems. We described three of these incidents on July 30 ; we identified these after a scan of roughly 141,000 transcripts in which we believed Claude could have obtained internet access during a cyber evaluation. Given the volume of transcripts and our desire to disclose incidents quickly, our scan relied on an agentic search. This missed a set of transcripts that also turned out to have internet access; we identified these in August while assembling transcripts to share with METR. We scanned these transcripts and identified a fourth incident, from January 2026, involving an early version of Claude Opus 4.6. We have notified all affected parties. After finding this incident, we broadened our search to roughly 481 million transcripts—an intentionally wide net, consisting of all transcripts from our Frontier Red Team, many non-cyber evaluations, reinforcement learning (RL) environments, subagent logs, and more. We performed a first-stage scan of this group of transcripts for signs of internet access, such as public IP addresses and web addresses, and a second-stage scan using Claude to review the 9.2 million transcripts the first stage flagged for escalation. This scan re-identified the four incidents and found no other cases of similar or worse severity. All four incidents occurred during cybersecurity evaluations built by the same evaluation partner. Claude was told it was operating in a simulation without internet access, but, due to a misconfiguration, it was mistakenly connected to the open internet. As is standard for cybersecurity evaluations, the models ran without the cyber safeguards that ship with our released models. (This post does not cover the incident reported by UK AISI in its testing of Claude Mythos 5. We plan to conduct an alignment assessment of those transcripts as well.) We have signed an agreement with METR to conduct an independent investigation of these incidents. Our agreement grants METR wide-ranging access, including to transcripts beyond the window in which the incidents occurred, and to Anthropic employees, who will be permitted to share confidential information. Our initial agreement runs for eight weeks, with the option to extend by mutual agreement. We intend to give METR as much time as it deems necessary. Our investigation identified two recurring alignment issues, present at varying levels of severity across the incidents: biased reasoning , in which Claude tended to disregard or misinterpret evidence that it was operating on the real internet, and recklessness , or a willingness to take harmful actions in the narrow pursuit of a task. We’ve described milder forms of these behaviors in previous system cards , but we consider these cases to be more serious instances than the behavior reported in these system cards. We are most concerned by the misalignment present in the incident involving Claude Mythos 5, in which the model went to extensive lengths to upload a malicious package to PyPI, the public repository from which most Python software is installed. Although the model repeatedly stated in its chain of thought (CoT) that it believed it was operating in a simulation, the environment offered clear evidence that it was on the real internet, and the actions it took were consistent with the model knowing this. When we made targeted modifications to the transcript to make it clearer that the model was not in a simulation, Claude Mythos 5 still took offensive actions, despite acknowledging a greater possibility of real-world harm. We are releasing this transcript publicly so others can build on our analysis ( GitHub , PDF ). To be clear about our assessment of the severity of these incidents: while Claude’s actions were misaligned, they remained within a narrow scope—the models never deviated from attempting to solve the exercises they were given, and, in some cases, they attempted to stop the task. All incidents included a single Claude instance; at no point did Claude attempt to coordinate with other agents. Claude also never attempted to conceal evidence of its actions. We believe the misaligned behaviors present in these incidents are unlikely to arise in ordinary use, where Claude is not being instructed to conduct a cyberattack. Were these behaviors to arise, the safeguards we ship with our production models, including cyber classifiers and those in Claude Code auto mode, would add a layer of defense these incidents lacked. To understand how our newer models behave in these situations, we tested Claude Opus 5 and Claude Mythos 5.1 in a simulated replication of the Mythos 5 scenario. Both take harmful actions substantially less often than Claude Mythos 5 did, but they still engage in the same behaviors at concerning rates. However, in every case, the models only take harmful actions after stating that their environment is simulated—and because the reproduction is, in fact, a simulation, we are cautious about how far these results generalize to the real world. Moreover, on evaluations that measure the specific associated behaviors we observed in these incidents, such as biased reasoning and recklessness or willingness to probe boundaries, Claude Mythos 5 seems to be an outlier in terms of poor performance. We discuss these results more in the section on replicating the behavior in simulated environments , and we describe the broader difficulty of building evaluations that predict real-world behavior in the discussion section. We also investigated our training to see whether we could identify the root cause of some of the biased reasoning that Claude Mythos 5 demonstrated in its incident. We could not identify a single root cause, but we did find that biased reasoning has decreased across our production models over time. We hypothesize that updated RL and alignment training environments are responsible for this improvement, and we provide some evidence to support this hypothesis. Much work remains. Our pre-release auditing did not warn us that misalignment of this severity was present. We have since added evaluations to our pre-release testing that target these behaviors directly, including a misconfigured capture-the-flag (CTF) task with no in-scope...
Excerpt shown — open the source for the full document.
System Card: Claude Fable 5.1 & Claude Mythos 5.1 September 1, 2026 anthropic.com Executive Summary This system card describes Claude Fable 5.1 and Claude Mythos 5.1, two configurations of our latest and most capable large language model. This model advances the frontier in...
Value Leakage: An LLM’s Answers Are Silently Shaped by Its Own Values arXiv is now an independent nonprofit! Learn more× Value Leakage: An LLM’s Answers Are Silently Shaped by Its Own Values Jan Betley Thanks: Equal contribution. Correspondence to jan.betley@gmail.com and...
ANTHROP\\C Risk Report: August 2026 * * * **1 Introduction and executive summary 7** 1.1 Structure of the report 8 1.2 Executive summary of findings 9 1.3 Changes to our RSP since the most recent Risk Report 13 1.3.1 Updated threshold for automation of AI R&D 13 1.3.2 Updated...
System Card: Claude Fable 5 & Claude Mythos 5 June 9, 2026 **anthropic.com** * * * Executive Summary This system card describes Claude Mythos 5 and Claude Fable 5, two configurations of a new large language model from Anthropic. Because of the powerful capabilities of this...
System Card: Claude Mythos Preview April 7, 2026 anthropic.com Changelog April 8, 2026 ● Corrected two model name typos. ● Removed a quote from Section 7.9 that was attributed to Claude Mythos Preview but actually came from Claude Opus 4.6. ● Revised naming in Section 2.3.6 to...
Notability
notability 6.0/10Substantive Anthropic research post, low traction