Anthropic Raises AI Loss-of-Control Risk to 'Low', Withholds Model 2 as OpenAI Slows Astra
Anthropic has raised its estimated probability of AI model loss-of-control in high-risk scenarios from "extremely low" to "low" in its latest risk report, citing recent cybersecurity incidents. The company also said it has no plans to release "Model 2," an internal large model that surpasses its flagship "Mythos" in capability, though the performance leap is less dramatic than the Opus 4.6-to-Mythos upgrade. Anthropic acknowledged lower confidence in this assessment as traditional evaluation methods fail to capture true capability growth. Meanwhile, competitor OpenAI has slowed the release of its next-generation Astra model over unresolved cybersecurity concerns.
In its latest risk report, Anthropic said it currently has no plans to release an internal large model designated "Model 2." The model appears to surpass the company's existing flagship model "Mythos" in capability, though the overall pace of research and development has not slowed. The report noted that while the risk of the model causing the most severe harm remains at a relatively low level, the risk indicators have shifted compared with the previous report. Citing recent cybersecurity incidents, Anthropic raised its estimated probability of model loss-of-control in high-risk scenarios from "extremely low" to "low." The company also observed an accelerating trend in the model's ability to perform automated research and development, a capability that, while conducive to technological breakthroughs, also carries the risk of misuse.
Regarding "Model 2," Anthropic explained that under its standard development process, the team internally trains and evaluates many exploratory model versions, of which "Model 2" is one. The report, released on Friday, disclosed that the unreleased model has shown "significant progress" across multiple internal tasks. It is currently used extensively alongside Mythos 5 for internal programming, agent operations, and data generation. However, in terms of the magnitude of the performance leap, "Model 2" does not replicate the dramatic jump seen when Opus 4.6 was upgraded to Mythos. The report explicitly stressed that the company has no plans to bring the model to the external market.
The decision comes against a backdrop of heightened tension across the AI industry. OpenAI, its main competitor, has also recently decided to slow the release of its next-generation Astra model, citing the inability to fully rule out potential serious cybersecurity risks. The threat signals conveyed in Anthropic's report suggest that ascertaining the upper limits of its own model capabilities and potential risks is becoming increasingly difficult. With regard to "Model 2," the report acknowledged that the team's confidence in this risk assessment is lower than before, as the most specific, task-based traditional evaluation methods can no longer accurately capture the true growth in model capabilities.
Why this event matters
The event has a measured impact on 1 industry. The strongest current signal is neutral for Artificial Intelligence, with intensity 25/100 and 70% confidence over a short term horizon.
Impact figures are analytical estimates that combine direction, intensity, confidence and event importance. They are not investment advice.