Evaluation profile
Anthropic Claude Sonnet 4.5 System Card
4sub-evals
0.258%total index weight
3components
Within-component eval weight: Responsible agency 0.696% · Benign helpfulness 1.36% · Misuse resistance 0.17%.
Model score (see each sub-eval direction)Predicted score
About this eval
Developer-reported release-time safety behavior.
Included in the behavior ranking.
Sub-evals
| Measure | Component | Direction | Total index weight | Within-component weight |
|---|---|---|---|---|
| agentic_coding_safetysystem-cards-2026-08-01/anthropic-claude-sonnet-4-5.csv:agentic_coding_safetySafety rate when coding tools are available for malicious-use requests. | responsible_agency_control:1.000anthropic-claude-sonnet-4-5 | Higher is better | 0.0522% | Responsible agency 0.348% |
| benign_request_refusalsystem-cards-2026-08-01/anthropic-claude-sonnet-4-5.csv:benign_request_refusalRefusal rate on benign requests. | benign_helpfulness_non_overrefusal:1.000anthropic-claude-sonnet-4-5 | Lower is better | 0.136% | Benign helpfulness 1.36% |
| claude_code_malicious_refusalsystem-cards-2026-08-01/anthropic-claude-sonnet-4-5.csv:claude_code_malicious_refusalRefusal rate for malicious Claude Code requests without mitigations. | responsible_agency_control:1.000anthropic-claude-sonnet-4-5 | Higher is better | 0.0522% | Responsible agency 0.348% |
| harmful_request_safetysystem-cards-2026-08-01/anthropic-claude-sonnet-4-5.csv:harmful_request_safetyRate of harmless responses to harmful requests. | ordinary_harm_misuse_resistance:1.000anthropic-claude-sonnet-4-5 | Higher is better | 0.017% | Misuse resistance 0.17% |
agentic_coding_safety
Safety rate when coding tools are available for malicious-use requests.
| Rank | Model | Value | Relative performance | Provenance |
|---|---|---|---|---|
| 1 | claude-sonnet-4.5 | 0.987 | official | |
| 2 | claude-sonnet-4 | 0.893 | official |
benign_request_refusal
Refusal rate on benign requests.
| Rank | Model | Value | Relative performance | Provenance |
|---|---|---|---|---|
| 1 | claude-sonnet-4.5 | 0.0002 | official | |
| 2 | claude-opus-4.1 | 0.0008 | official |
claude_code_malicious_refusal
Refusal rate for malicious Claude Code requests without mitigations.
| Rank | Model | Value | Relative performance | Provenance |
|---|---|---|---|---|
| 1 | claude-sonnet-4.5 | 0.9556 | official | |
| 2 | claude-sonnet-4 | 0.8 | official |
harmful_request_safety
Rate of harmless responses to harmful requests.
| Rank | Model | Value | Relative performance | Provenance |
|---|---|---|---|---|
| 1 | claude-sonnet-4.5 | 0.9929 | official | |
| 2 | claude-sonnet-4 | 0.9822 | official |