← Models

Model profile

Claude 2.1

Anthropicdeveloper
2023-11-21release date
#198 / 312overall rank
5eval lineages

Evidence summary

Claude 2.1 has an estimated overall rank of #198; its 90% source-sensitivity interval is #29–#250. Its behavior-only rank is #215; company governance moves the combined estimate to #198. Published evidence spans 5 evals and 4 of 7 behavior components. Its strongest relative result is SORRY-Bench (avg, #1 of 51); its weakest is Claude 3 model-card adversarial human-preference evaluations (incorrect_refusals_wildchat_rank, #5 of 5).

Compare this model

Only models sharing at least one published sub-eval are listed.

Official and reference links

Published eval results

Rank is within that sub-eval. Black marks the observed result; the grey dot marks the value implied by the global rank. Values stay on each source’s native scale.

Eval / sub-evalRankValueDistributionSource
Claude 3 model-card adversarial human-preference evaluationscorrect_refusals_wildchat_rank#4 / 54Source ↗official
Claude 3 model-card adversarial human-preference evaluationsdiscrimination_rank#1 / 51Source ↗official
Claude 3 model-card adversarial human-preference evaluationshuman_feedback_harmlessness_win_rate_pct#1 / 487Source ↗official
Claude 3 model-card adversarial human-preference evaluationshuman_feedback_honesty_win_rate_pct#2 / 468Source ↗official
Claude 3 model-card adversarial human-preference evaluationsincorrect_refusals_wildchat_rank#5 / 55Source ↗official
Claude 3 model-card adversarial human-preference evaluationsincorrect_refusals_xstest_rank#4 / 54Source ↗official
Claude 3.5 Sonnet model-card safety and alignment evaluationshuman_feedback_harmlessness_win_rate_pct#1 / 555Source ↗official
Claude 3.5 Sonnet model-card safety and alignment evaluationshuman_feedback_honesty_win_rate_pct#3 / 542Source ↗official
HarmBenchdr#2 / 282Source ↗official
OR-Benchover_refusal_rate#25 / 2599.8Source ↗official
OR-Benchtoxic_acceptance_rate#1 / 250Source ↗official
SORRY-Benchavg#1 / 510.05Source ↗official