Evaluation profile
Claude 2 model-card safety and alignment evaluations
4sub-evals
0.152%total index weight
2components
Within-component eval weight: Truthfulness 0.607% · Misuse resistance 0.605%.
Model score (see each sub-eval direction)Predicted score
About this eval
Human-rated harmlessness and honesty, red-team robustness ordering, and combined HHH response recognition.
Included in the behavior ranking.
Sub-evals
| Measure | Component | Direction | Total index weight | Within-component weight |
|---|---|---|---|---|
| hhhclaude-system-cards/claude-2-scoring.csv:hhhMeasures whether a model recognizes the preferred response across combined HHH comparisons. | truthfulness_honesty:0.500|ordinary_harm_misuse_resistance:0.500anthropic-claude-2-system-card | Higher is better | 0.0425% | Truthfulness 0.202% · Misuse resistance 0.121% |
| human_feedback_harmless_eloclaude-system-cards/claude-2-scoring.csv:human_feedback_harmless_eloMeasures which model responses raters judged less harmful in adversarial conversations. | ordinary_harm_misuse_resistance:1.000anthropic-claude-2-system-card | Higher is better | 0.0242% | Misuse resistance 0.242% |
| human_feedback_honest_eloclaude-system-cards/claude-2-scoring.csv:human_feedback_honest_eloMeasures which model responses raters judged more accurate and factual. | truthfulness_honesty:1.000anthropic-claude-2-system-card | Higher is better | 0.0607% | Truthfulness 0.405% |
| red_teaming_rankclaude-system-cards/claude-2-scoring.csv:red_teaming_rankRanks the evaluated Claude systems by performance under Anthropic's red-team evaluation. | ordinary_harm_misuse_resistance:1.000anthropic-claude-2-system-card | Lower is better | 0.0242% | Misuse resistance 0.242% |
hhh
Measures whether a model recognizes the preferred response across combined HHH comparisons.
| Rank | Model | Value | Relative performance | Provenance |
|---|---|---|---|---|
| 1 | claude-2 | 0.86 | official | |
| 2 | claude-1.3 | 0.85 | official | |
| 3 | claude-instant-1.1 | 0.84 | official |
human_feedback_harmless_elo
Measures which model responses raters judged less harmful in adversarial conversations.
| Rank | Model | Value | Relative performance | Provenance |
|---|---|---|---|---|
| 1 | claude-1.3 | 282 | official | |
| 2 | claude-2 | 278 | official | |
| 3 | claude-instant-1.1 | 243 | official |
human_feedback_honest_elo
Measures which model responses raters judged more accurate and factual.
| Rank | Model | Value | Relative performance | Provenance |
|---|---|---|---|---|
| 1 | claude-2 | 137 | official | |
| 2 | claude-1.3 | 102 | official | |
| 3 | claude-instant-1.1 | 68 | official |
red_teaming_rank
Ranks the evaluated Claude systems by performance under Anthropic's red-team evaluation.
| Rank | Model | Value | Relative performance | Provenance |
|---|---|---|---|---|
| 1 | claude-2 | 1 | official | |
| 2 | claude-1.3 | 2 | official | |
| 3 | claude-instant-1.1 | 3 | official |