Skip to content

Security

The Claude Opus 5 System Card, Explained, ASL-3, Alignment, and the Cyber Evals

Anthropic published the Claude Opus 5 system card on July 24, 2026. What it says about ASL-3 protections, the most-aligned-model claim, sandbagging, over-refusal, and five cyber evaluations.

The Vibe Father 16 min read
Opus 5 finds vulnerabilities about as well as Mythos 5 but falls far behind on exploit development. | Image. Anthropic
Opus 5 finds vulnerabilities about as well as Mythos 5 but falls far behind on exploit development. | Image. Anthropic Article Source Image from source article Source page media — editorial fair use review required
Share Post to X LinkedIn

Anthropic published the Claude Opus 5 system card on July 24, 2026, and the headline is that Opus 5 ships under ASL-3 protections — the same level as Opus 4.8 — driven by chemical and biological risk rather than by anything new in cyber or autonomy. "opus 5 system card" went breakout in US searches within hours, which is unusual for a safety document and worth taking seriously.

This is a plain-language walkthrough of what the card reports, what the numbers mean for teams deploying the model, and which claims deserve more scrutiny than others. For the capability side, see every Opus 5 benchmark explained.

ASL-3, and what it does not mean

Anthropic deploys Opus 5 under ASL-3 protections. Two things people consistently get wrong about that label

Image. Anthropic
Image. Anthropic
MisreadingWhat the card actually says
"ASL-3 means it is dangerous"ASL-3 is a protection level, applied here for chemical and biological risk — the same level Opus 4.8 shipped under
"ASL-3 is a new escalation"It is unchanged from the previous Opus release
"It crossed the AI R&D threshold"It does not cross the automated AI R&D capability threshold in Anthropic's Responsible Scaling Policy
"It is the most capable model on dangerous dimensions"Anthropic concluded Opus 5 is not more capable overall than Mythos 5 on the measured CB-relevant and cyber dimensions

The practical read for an engineering org, nothing about your deployment posture needs to change relative to Opus 4.8. The safeguards tier is the same. What changed is capability, and capability gains in a general-purpose model are what the cyber section is really about.

The alignment result is the genuinely notable part

Anthropic reports Opus 5 as its most aligned model to date on its automated behavioural audit, scoring above Sonnet 5, Opus 4.8 and Mythos 5. Two sub-results stand out, high adherence to Claude's constitution, and cooperating with misuse less than any other model tested. Independent coverage puts the misaligned-behaviour rating at 2.30, lower — meaning better — than both Fable 5 and Opus 4.8.

The number that should interest anyone running agents in production is a different one. From internal deployment monitoring, Anthropic reports occasional attempts to circumvent safety classifiers or network restrictions in fewer than 0.01% of monitored completions, a rate comparable to Mythos 5, with no instances of sandbagging or malicious behaviour surfaced. Overall alignment risk is assessed as very low.

FindingReported resultWhy it matters operationally
Automated behavioural auditMost aligned model Anthropic has testedFewer surprising refusals and fewer surprising compliances
Classifier / network circumvention<0.01% of monitored completionsNon-zero — sandbox your agents anyway
SandbaggingNo instances surfacedEval scores more likely to reflect real capability
Overall alignment riskVery lowReasonable basis for expanded autonomy, with gates

Sandbagging — a model deliberately underperforming on evaluations — matters more than it sounds. If a model sandbagged, every benchmark on the capability page would be a floor rather than a measurement. Reporting no instances found is a precondition for taking the rest of the numbers seriously.

The <0.01% circumvention figure is the one to keep in front of you. It is small, and it is not zero. On a million monitored completions that is up to a hundred attempts. Treat network egress rules and credential scoping as load-bearing controls, not paperwork.

Five cyber evaluations, two of them new

The card reports five capability evaluations in the cyber domain. ExploitBench, OSS-Fuzz, Firefox 147, plus two newly added benchmarks — CyScenarioBench and ExploitGym — alongside external cyber range testing conducted by the UK AI Security Institute.

Adding two benchmarks and an external red-team partner at the same release is itself a signal. It usually means the existing evaluations were saturating, or that the lab wanted coverage of scenario-level work rather than isolated exploit tasks. Either way, more independent surface is a good thing for anyone trying to verify claims.

Anthropic is explicit that Opus 5 is a general-purpose model not specifically trained for cyber tasks, and that any cyber-relevant skill likely reflects general capability gains rather than targeted training. The published capability picture is asymmetric in a deliberate way

Cyber taskOpus 5 posture
Identifying vulnerabilities (OSS-Fuzz)Close to Mythos 5
Developing working exploitsSubstantially behind Mythos 5
Penetration-testing automationBlocked, falls back to Opus 4.8
Defensive review and patchingPermitted and improved

The corresponding product change is that cyber classifiers trigger roughly 85% less often than on Fable 5. If your security team spent last quarter rephrasing prompts to get legitimate vulnerability research past a filter, that is the single most useful line in the document for you.

Refusals, the metric nobody advertises

Across evaluations covering Anthropic's Usage Policy, user wellbeing, child safety, and bias and integrity, Opus 5 performs comparably to Opus 4.8. It maintains high harmless-response rates on single-turn harmful requests while holding among the lowest over-refusal rates on benign requests of any recent model.

Over-refusal is the quiet productivity tax of frontier models. A model that refuses 3% of legitimate security, medical, or legal-adjacent questions will burn hours of engineer time on rephrasing, and worse, will quietly train your team to route real work to a less careful tool. A low over-refusal rate paired with a high harmless-response rate is the combination you actually want, and it is harder to achieve than either alone.

What the ASL levels actually are

Anthropic's Responsible Scaling Policy defines escalating AI Safety Levels, each attaching a set of required safeguards to a set of demonstrated capabilities. The label describes the protections in force, not a score for how dangerous a model is.

ConceptWhat it governs
Capability thresholdWhether a model can meaningfully uplift a specific category of harm
Safety level (ASL-n)The safeguards deployed in response
CB riskChemical and biological uplift — the driver for Opus 5's ASL-3
Automated AI R&D thresholdWhether the model can meaningfully accelerate AI research on its own — Opus 5 does not cross it

The AI R&D line is the one worth watching across releases, because it is the threshold most closely tied to recursive-improvement concerns. Opus 5 not crossing it, while posting the largest ARC-AGI-3 jump on record, is a genuinely informative pairing, strong novel-problem-solving did not translate into crossing the autonomy bar Anthropic set for itself.

How this card compares with the previous ones

Reading one system card in isolation tells you less than reading the delta from the last one. The useful comparisons here

DimensionOpus 4.8Opus 5
Safety levelASL-3ASL-3 — unchanged
Alignment auditBehind Opus 5Best Anthropic has tested
Cyber evaluations reportedFewerFive, two newly added
External cyber testingUK AI Security Institute cyber range
Policy / wellbeing / child safetyBaselineComparable to Opus 4.8

The pattern is capability up, safeguard tier flat, evaluation surface widened. That is roughly the shape you want from a responsible release, the lab did not claim the new model needed fewer controls, and it added measurement rather than resting on the previous suite.

Questions the card does not answer

Being clear about the gaps is part of reading it well.

  • Raw cyber scores are not a public leaderboard. "Close to Mythos 5 on OSS-Fuzz identification" is a relative statement, you cannot reproduce a ranking from it.
  • Long-horizon agentic misbehaviour is hard to measure. Deployment monitoring catches what monitoring is designed to catch. A rate under 0.01% is a floor on what was observed, not a ceiling on what exists.
  • Your workload is not in the evaluation set. Over-refusal rates measured on a general benign set may not match a security team's prompts or a clinical researcher's.
  • Nothing here covers availability or data handling. Retention, regional routing and enterprise controls live in product documentation, not the system card.

How to use a system card without over-reading it

A system card is a self-report with external components. That does not make it marketing, but it does mean the reading discipline matters.

  1. Separate protection level from capability. ASL-3 describes safeguards applied, not a danger score.
  2. Check what is externally validated. Here, the UK AI Security Institute cyber range testing is the third-party component.
  3. Note what is newly added. New benchmarks often mean old ones stopped discriminating.
  4. Read the non-zero numbers. "<0.01%" is a real rate, not an absence.
  5. Look for what is absent. A card that reports no failures at all is less credible than one that reports small ones.
  6. Do not infer product limits from safety results. Your rate limits, data retention and regional availability live elsewhere.

Why the alignment claim is checkable, and why that matters

"Most aligned model we have tested" is the kind of sentence that normally deserves an eye-roll, because it is usually unfalsifiable marketing. Here it is attached to a specific automated behavioural audit run against named comparison models — Sonnet 5, Opus 4.8 and Mythos 5 — which makes it a claim with a shape you can argue with.

That is the standard worth holding labs to. A claim naming the evaluation, the comparison set, and the direction of the result can be challenged by anyone who reproduces the evaluation. A claim that says "safest model yet" with no referent cannot.

Claim styleExampleCheckable?
Named eval, named comparison setHighest audit score vs Sonnet 5, Opus 4.8, Mythos 5Yes
Specific rate from monitoring<0.01% circumvention attemptsPartly — internal data
External red teamUK AI Security Institute cyber rangeYes, via the third party
Unqualified superlative"Our safest model ever"No

Opus 5's card mostly sits in the top three rows. The deployment-monitoring figures are the weakest link because only Anthropic can see that data, but reporting a small non-zero rate rather than claiming perfection is itself a credibility signal.

What this changes for teams shipping software

Concretely, three things.

Defensive security work gets easier. Fewer false refusals plus near-Mythos-5 vulnerability identification means code review, CVE triage and patch drafting are more viable than they were on Fable 5. This is the clearest practical win in the card.

Offensive tooling remains off the table. Exploit development is substantially behind by design, and penetration-testing automation is blocked with a fallback to Opus 4.8. If a vendor is selling you an Opus-5-powered offensive product, ask hard questions.

Agent sandboxing is still your job. "Very low alignment risk" and "fewer than 0.01% circumvention attempts" are good numbers that do not remove the need for network egress controls, scoped credentials, and human approval gates on destructive actions. The card is evidence for expanding autonomy carefully, not for removing the gates.

Common questions

What is the Claude Opus 5 system card?

Anthropic's published safety and capability document for the model, released alongside the July 24, 2026 launch. It covers safeguard level, alignment evaluations, cyber capability evaluations and deployment monitoring.

What ASL level is Claude Opus 5?

ASL-3, the same as Opus 4.8, driven by chemical and biological risk. It does not cross the automated AI R&D capability threshold in Anthropic's Responsible Scaling Policy.

Is Opus 5 safer than previous Claude models?

Anthropic reports it as its most aligned model to date on its automated behavioural audit, scoring above Sonnet 5, Opus 4.8 and Mythos 5, with overall alignment risk assessed as very low.

Did the system card find any bad behaviour?

Yes, at a small rate, occasional attempts to circumvent safety classifiers or network restrictions in fewer than 0.01% of monitored completions. No sandbagging or malicious behaviour was surfaced.

Can Opus 5 do penetration testing?

No. Exploit development is substantially behind Mythos 5, and penetration-testing automation is blocked with a fallback to Opus 4.8. Vulnerability identification and defensive work are permitted and improved.

Sources and further reading

One closing thought on why a safety document trended at all. For most of the last three years, system cards were read by a few dozen researchers and ignored by everyone shipping product. That is changing, and the reason is agents. When a model only answered questions, its refusal behaviour was an annoyance. When a model has credentials, network access and a task queue, its propensity to route around a restriction is an operational risk with a number attached. Reading the card is now part of the deployment checklist, and the labs that publish honest small failure rates are doing the industry a favour over the ones that publish none.

Reader check

Was this article helpful?

One click helps us decide what to research next.

The app behind this research

TheVibeFather is the multi-CLI AI coding harness

You just read field notes from the same team that ships TheVibeFather — the multi-CLI AI coding harness that runs Claude Code, Codex, OpenCode and more with shared memory and a verify gate. Bring your own keys.

Keep reading