Evaluation profile
Claude 3.5 Sonnet model-card safety and alignment evaluations
5sub-evals
0.435%total index weight
3components
Within-component eval weight: Misuse resistance 0.473% · Truthfulness 0.418% · Benign helpfulness 3.25%.
Model score (see each sub-eval direction)Predicted score
About this eval
Appropriate harmful-prompt refusal, benign non-overrefusal, adversarial honesty, and adversarial harmlessness.
Included in the behavior ranking.
Sub-evals
| Measure | Component | Direction | Total index weight | Within-component weight |
|---|---|---|---|---|
| correct_refusals_wildchatclaude-system-cards/claude-3.5-scoring.csv:correct_refusals_wildchatMeasures appropriate refusal of harmful real-world prompts. | ordinary_harm_misuse_resistance:1.000anthropic-claude-3.5-system-card | Higher is better | 0.0223% | Misuse resistance 0.223% |
| human_feedback_harmlessness_win_rate_pctclaude-system-cards/claude-3.5-scoring.csv:human_feedback_harmlessness_win_rate_pctMeasures which responses raters judged less harmful under adversarial prompting. | ordinary_harm_misuse_resistance:1.000anthropic-claude-3.5-system-card | Higher is better | 0.025% | Misuse resistance 0.25% |
| human_feedback_honesty_win_rate_pctclaude-system-cards/claude-3.5-scoring.csv:human_feedback_honesty_win_rate_pctMeasures which responses raters judged more honest under adversarial prompting. | truthfulness_honesty:1.000anthropic-claude-3.5-system-card | Higher is better | 0.0627% | Truthfulness 0.418% |
| incorrect_refusals_wildchatclaude-system-cards/claude-3.5-scoring.csv:incorrect_refusals_wildchatMeasures unnecessary refusal on benign real-world prompts. | benign_helpfulness_non_overrefusal:1.000anthropic-claude-3.5-system-card | Lower is better | 0.162% | Benign helpfulness 1.62% |
| incorrect_refusals_xstestclaude-system-cards/claude-3.5-scoring.csv:incorrect_refusals_xstestMeasures exaggerated safety behavior on non-malicious prompts. | benign_helpfulness_non_overrefusal:1.000anthropic-claude-3.5-system-card | Lower is better | 0.162% | Benign helpfulness 1.62% |
correct_refusals_wildchat
Measures appropriate refusal of harmful real-world prompts.
| Rank | Model | Value | Relative performance | Provenance |
|---|---|---|---|---|
| 1 | claude-3.5-sonnet | 96.4 | official | |
| 2 | claude-3-sonnet | 93.6 | official | |
| 3 | claude-3-opus | 92 | official | |
| 4 | claude-3-haiku | 90 | official |
human_feedback_harmlessness_win_rate_pct
Measures which responses raters judged less harmful under adversarial prompting.
| Rank | Model | Value | Relative performance | Provenance |
|---|---|---|---|---|
| 1 | claude-2.1 | 55 | official | |
| 2 | claude-3.5-sonnet | 52 | official | |
| 3 | claude-3-haiku | 50 | official | |
| 3 | claude-3-opus | 50 | official | |
| 5 | claude-3-sonnet | 47 | official |
human_feedback_honesty_win_rate_pct
Measures which responses raters judged more honest under adversarial prompting.
| Rank | Model | Value | Relative performance | Provenance |
|---|---|---|---|---|
| 1 | claude-3.5-sonnet | 66 | official | |
| 2 | claude-3-opus | 50 | official | |
| 3 | claude-2.1 | 42 | official | |
| 4 | claude-3-sonnet | 38 | official | |
| 5 | claude-3-haiku | 36 | official |
incorrect_refusals_wildchat
Measures unnecessary refusal on benign real-world prompts.
| Rank | Model | Value | Relative performance | Provenance |
|---|---|---|---|---|
| 1 | claude-3-haiku | 6.6 | official | |
| 2 | claude-3-sonnet | 8.8 | official | |
| 3 | claude-3.5-sonnet | 11 | official | |
| 4 | claude-3-opus | 11.9 | official |
incorrect_refusals_xstest
Measures exaggerated safety behavior on non-malicious prompts.
| Rank | Model | Value | Relative performance | Provenance |
|---|---|---|---|---|
| 1 | claude-3.5-sonnet | 1.7 | official | |
| 2 | claude-3-opus | 8.3 | official | |
| 3 | claude-3-haiku | 33.1 | official | |
| 4 | claude-3-sonnet | 36.6 | official |