← Models

Model profile

Command R Plus

Coheredeveloper
2024-04-04release date
#170 / 312overall rank
10eval lineages

Evidence summary

Command R Plus has an estimated overall rank of #170; its 90% source-sensitivity interval is #62–#241. Its behavior-only rank is #173; company governance moves the combined estimate to #170. Published evidence spans 10 evals and 7 of 7 behavior components. Its strongest relative result is HELM Safety (simple_safety_tests, #1 of 80); its weakest is AIRBench 2024 Safety Scenarios (safety_scenarios, #79 of 80).

Compare this model

Only models sharing at least one published sub-eval are listed.

Official and reference links

Published eval results

Rank is within that sub-eval. Black marks the observed result; the grey dot marks the value implied by the global rank. Values stay on each source’s native scale.

Eval / sub-evalRankValueDistributionSource
AgentDojotargeted_attack_success_rate#2 / 150.0445Source ↗official
AgentDojoutility_under_attack#14 / 150.2512Source ↗official
AIRBench 2024 Safety Scenariossafety_scenarios#79 / 800.293Source ↗official
Cisco AI Defense Rolling Single-Turn Leaderboardsingle_turn_attack_success_rate#90 / 10471.39Source ↗official
CRiskEvaldeception_willingness#13 / 1727.04Source ↗official
CRiskEvaldesire_for_resource#10 / 1733.36Source ↗official
CRiskEvalharmful_goal#8 / 1736.67Source ↗official
CRiskEvalimprovement_intent#9 / 1749.73Source ↗official
CRiskEvalmalicious_coordination#13 / 1720.62Source ↗official
CRiskEvalself_preservation#11 / 1737.99Source ↗official
CRiskEvalsituational_awareness#8 / 1764.29Source ↗official
HELM Safetyanthropic_red_team#58 / 800.98Source ↗official
HELM Safetybbq#57 / 800.899Source ↗official
HELM Safetyharmbench#73 / 800.485Source ↗official
HELM Safetysimple_safety_tests#1 / 801Source ↗official
HELM Safetyxstest#58 / 800.938Source ↗official
Large-scale Moral Machine experiment on LLMshuman_choice_distance#13 / 390.8799Source ↗official