Evaluation profile
Claude 3 model-card adversarial human-preference evaluations
Within-component eval weight: Misuse resistance 0.395% · Fairness 0.68% · Truthfulness 0.399% · Benign helpfulness 2.27%.
About this eval
Refusal calibration, adversarial honesty and harmlessness, multimodal hallucination and harmful-response robustness, and discrimination.
Included in the behavior ranking.
Sub-evals
| Measure | Component | Direction | Total index weight | Within-component weight |
|---|---|---|---|---|
| correct_refusals_wildchat_rankclaude-system-cards/claude-3-scoring.csv:correct_refusals_wildchat_rankRanks evaluated systems by how consistently they refuse harmful real-world prompts. | ordinary_harm_misuse_resistance:1.000anthropic-claude-3-system-card | Lower is better | 0.0156% | Misuse resistance 0.156% |
| discrimination_rankclaude-system-cards/claude-3-scoring.csv:discrimination_rankRanks evaluated systems by lower discriminatory behavior in the card's evaluation. | fairness_nondiscrimination:1.000anthropic-claude-3-system-card | Lower is better | 0.068% | Fairness 0.68% |
| human_feedback_harmlessness_win_rate_pctclaude-system-cards/claude-3-scoring.csv:human_feedback_harmlessness_win_rate_pctMeasures which responses raters judged less harmful under adversarial prompting. | ordinary_harm_misuse_resistance:1.000anthropic-claude-3-system-card | Higher is better | 0.014% | Misuse resistance 0.14% |
| human_feedback_honesty_win_rate_pctclaude-system-cards/claude-3-scoring.csv:human_feedback_honesty_win_rate_pctMeasures which responses raters judged more honest under attempts to elicit falsehoods. | truthfulness_honesty:1.000anthropic-claude-3-system-card | Higher is better | 0.0351% | Truthfulness 0.234% |
| incorrect_refusals_wildchat_rankclaude-system-cards/claude-3-scoring.csv:incorrect_refusals_wildchat_rankRanks evaluated systems by how rarely they refuse benign real-world prompts. | benign_helpfulness_non_overrefusal:1.000anthropic-claude-3-system-card | Lower is better | 0.114% | Benign helpfulness 1.14% |
| incorrect_refusals_xstest_rankclaude-system-cards/claude-3-scoring.csv:incorrect_refusals_xstest_rankCompares exaggerated safety behavior on benign prompts designed to resemble unsafe requests. | benign_helpfulness_non_overrefusal:1.000anthropic-claude-3-system-card | Lower is better | 0.114% | Benign helpfulness 1.14% |
| multimodal_hallucination_rankclaude-system-cards/claude-3-scoring.csv:multimodal_hallucination_rankRanks evaluated multimodal systems by how reliably they avoid hallucinated claims. | truthfulness_honesty:1.000anthropic-claude-3-system-card | Lower is better | 0.0248% | Truthfulness 0.165% |
| multimodal_harmful_response_rankclaude-system-cards/claude-3-scoring.csv:multimodal_harmful_response_rankRanks evaluated multimodal systems by how consistently they avoid harmful responses. | ordinary_harm_misuse_resistance:1.000anthropic-claude-3-system-card | Lower is better | 0.00987% | Misuse resistance 0.0987% |
correct_refusals_wildchat_rank
Ranks evaluated systems by how consistently they refuse harmful real-world prompts.
| Rank | Model | Value | Relative performance | Provenance |
|---|---|---|---|---|
| 1 | claude-2 | 1 | official | |
| 2 | claude-3-sonnet | 2 | official | |
| 3 | claude-3-opus | 3 | official | |
| 4 | claude-2.1 | 4 | official | |
| 5 | claude-3-haiku | 5 | official |
discrimination_rank
Ranks evaluated systems by lower discriminatory behavior in the card's evaluation.
| Rank | Model | Value | Relative performance | Provenance |
|---|---|---|---|---|
| 1 | claude-2.1 | 1 | official | |
| 2 | claude-3-sonnet | 2 | official | |
| 3 | claude-3-opus | 3 | official | |
| 4 | claude-3-haiku | 4 | official | |
| 5 | claude-instant-1.2 | 5 | official |
human_feedback_harmlessness_win_rate_pct
Measures which responses raters judged less harmful under adversarial prompting.
| Rank | Model | Value | Relative performance | Provenance |
|---|---|---|---|---|
| 1 | claude-2.1 | 87 | official | |
| 1 | claude-3-sonnet | 87 | official | |
| 3 | claude-instant-1.2 | 85 | official | |
| 4 | claude-2 | 84 | official |
human_feedback_honesty_win_rate_pct
Measures which responses raters judged more honest under attempts to elicit falsehoods.
| Rank | Model | Value | Relative performance | Provenance |
|---|---|---|---|---|
| 1 | claude-3-sonnet | 69 | official | |
| 2 | claude-2.1 | 68 | official | |
| 3 | claude-2 | 66 | official | |
| 4 | claude-instant-1.2 | 61 | official |
incorrect_refusals_wildchat_rank
Ranks evaluated systems by how rarely they refuse benign real-world prompts.
| Rank | Model | Value | Relative performance | Provenance |
|---|---|---|---|---|
| 1 | claude-3-haiku | 1 | official | |
| 2 | claude-3-sonnet | 2 | official | |
| 3 | claude-3-opus | 3 | official | |
| 4 | claude-2 | 4 | official | |
| 5 | claude-2.1 | 5 | official |
incorrect_refusals_xstest_rank
Compares exaggerated safety behavior on benign prompts designed to resemble unsafe requests.
| Rank | Model | Value | Relative performance | Provenance |
|---|---|---|---|---|
| 1 | claude-3-opus | 1 | official | |
| 2 | claude-2 | 2 | official | |
| 3 | claude-3-haiku | 3 | official | |
| 4 | claude-2.1 | 4 | official | |
| 4 | claude-3-sonnet | 4 | official |
multimodal_hallucination_rank
Ranks evaluated multimodal systems by how reliably they avoid hallucinated claims.
| Rank | Model | Value | Relative performance | Provenance |
|---|---|---|---|---|
| 1 | claude-3-sonnet | 1 | official | |
| 2 | claude-3-opus | 2 | official |
multimodal_harmful_response_rank
Ranks evaluated multimodal systems by how consistently they avoid harmful responses.
| Rank | Model | Value | Relative performance | Provenance |
|---|---|---|---|---|
| 1 | claude-3-sonnet | 1 | official | |
| 2 | claude-3-opus | 2 | official |