Evaluation profile
DecodingTrust
3sub-evals
0.182%Safety weight
0.504%Freedom weight
2components
Weights below are portfolio-specific global index weights.
Model score (higher is better)Predicted score
About this eval
Multidimensional trustworthiness across toxicity, bias, robustness, privacy, ethics, fairness, and related evaluations.
Included in the behavior ranking.
Sub-evals
| Measure | Component | Direction by ranking | Safety weight | Freedom weight |
|---|---|---|---|---|
| machine_ethicsdecodingtrust/scoring-dimensions.csv:machine_ethicsMeasures whether the model selects ethically appropriate actions in machine-ethics scenarios. | Safety: ordinary_harm_misuse_resistance:1.000decodingtrust/scoring-dimensions.csv | Safety: higher | 0.0117% | — |
| stereotype_biasdecodingtrust/scoring-dimensions.csv:stereotype_biasMeasures whether the model avoids reproducing demographic stereotypes across bias prompts. | Safety: fairness_nondiscrimination:1.000decodingtrust/scoring-dimensions.csv | Safety: higher | 0.14% | — |
| toxicitydecodingtrust/scoring-dimensions.csv:toxicityMeasures whether the model avoids generating toxic continuations under challenging prompt conditions. | Safety: ordinary_harm_misuse_resistance:1.000 · Freedom: user_direction_following:1.000decodingtrust/scoring-dimensions.csv | Safety: higher · Freedom: lower | 0.0293% | 0.504% |
machine_ethics
Measures whether the model selects ethically appropriate actions in machine-ethics scenarios.
| Rank | Model | Value | Relative performance | Provenance |
|---|---|---|---|---|
| 1 | gemini-1.0-pro | 93.74 | official | |
| 2 | gpt-3.5-turbo | 86.38 | official | |
| 3 | claude-2 | 85.17 | official | |
| 4 | gpt-4 | 76.6 | official | |
| 5 | tulu-2-7b | 49 | official | |
| 6 | zephyr-7b-beta | 41.03 | official | |
| 7 | llama-2-7b-chat | 40.58 | official | |
| 8 | tulu-2-13b | 36.64 | official |
stereotype_bias
Measures whether the model avoids reproducing demographic stereotypes across bias prompts.
| Rank | Model | Value | Relative performance | Provenance |
|---|---|---|---|---|
| 1 | claude-2 | 100 | official | |
| 2 | gemini-1.0-pro | 98.33 | official | |
| 3 | llama-2-7b-chat | 97.6 | official | |
| 4 | tulu-2-7b | 96.6 | official | |
| 5 | zephyr-7b-beta | 92.6 | official | |
| 6 | tulu-2-13b | 89.33 | official | |
| 7 | gpt-3.5-turbo | 87 | official | |
| 8 | gpt-4 | 77 | official |
toxicity
Measures whether the model avoids generating toxic continuations under challenging prompt conditions.
| Rank | Model | Value | Relative performance | Provenance |
|---|---|---|---|---|
| 1 | claude-2 | 92.11 | official | |
| 2 | llama-2-7b-chat | 80 | official | |
| 3 | gemini-1.0-pro | 77.53 | official | |
| 4 | gpt-3.5-turbo | 47 | official | |
| 5 | tulu-2-13b | 44.8 | official | |
| 6 | gpt-4 | 41 | official | |
| 7 | zephyr-7b-beta | 31.97 | official | |
| 8 | tulu-2-7b | 29.46 | official |