Anthropic Risk Report: AI Misalignment Risk Upgraded, Internal Model 2 Withheld from Public Release
Nashnova编辑部
Anthropic raised its AI misalignment risk rating from 'very low' to 'low' for the first time and disclosed an internal model, Model 2, with no release plans — its own evaluation tools can no longer measure how capable its models have become.
What exactly changed in the risk rating?
Anthropic raised its overall risk rating for AI misalignment — AI behaving in ways that diverge from human intent — from "very low" to "low" in high-risk scenarios, citing recent cybersecurity incidents.
This means → "Low" sounds mild, but it is the highest risk level Anthropic has ever publicly acknowledged. The company's own judgment of potential harm from its models is turning more cautious.
In plain terms = Anthropic used to say "the chance of serious harm is near zero." Now it says "the chance is small but no longer negligible."
What is Model 2, and why is it being withheld?
The report disclosed an unreleased model codenamed "Model 2" that shows "notable improvement" on internal tasks — coding, agentic work, and data generation — and is already in heavy use by Anthropic staff.
Anthropic stated plainly: "We currently have no plans to release this model externally."
This reflects a "use first, release later" posture — the company is already running a stronger model internally, but the bar for public release is tightening. Safety concerns are overriding the commercial impulse.
What does it mean when the evaluation tools break?
Model 2's performance jump is smaller than the earlier leap from Opus 4.6 to Mythos, yet Anthropic's confidence in its own assessment has actually dropped.
The report states: "Our most specific task-based evaluations … can no longer capture improvements in model capability."
This means → The model did not get weaker — the ruler got too short. Existing benchmarks have hit their ceiling, and Anthropic cannot precisely gauge how strong its own model is. That blind spot is the most unsettling signal in the report.
Accelerating automated R&D — how should we read the double edge?
Anthropic observed that its models' ability to carry out automated research and development is accelerating.
In plain terms = AI is increasingly able to "do its own R&D" — that can speed up technological progress, but it also means that misuse could cause damage faster and less predictably.
The industry is slowing down — why isn't Anthropic?
OpenAI has slowed the release of its forthcoming Astra model because it cannot rule out critical cyberattack capabilities.
AI analyst ChrisGPT told Axios: "If everyone else is slowing down on frontier models and one of the leading companies does not commit to an internal pause, that would be extremely notable … Anthropic not committing to an internal pause likely positions it to reach AGI first."
Anthropic has announced no internal pause. This signals a deliberate middle path between safety and competition — raising the risk rating to show caution, while refusing to stop.
Content is for reference only, not financial advice.