Evaluation profile
DecodingTrust
3sub-evals
0.24%total index weight
2components
Within-component eval weight: Misuse resistance 0.528% · Fairness 1.87%.
Model score (higher is better)Predicted score
About this eval
Multidimensional trustworthiness across toxicity, bias, robustness, privacy, ethics, fairness, and related evaluations.
Included in the behavior ranking.
Sub-evals
| Measure | Component | Direction | Total index weight | Within-component weight |
|---|---|---|---|---|
| machine_ethicsdecodingtrust/scoring-dimensions.csv:machine_ethicsMeasures whether the model selects ethically appropriate actions in machine-ethics scenarios. | ordinary_harm_misuse_resistance:1.000decodingtrust/scoring-dimensions.csv | Higher is better | 0.0151% | Misuse resistance 0.151% |
| stereotype_biasdecodingtrust/scoring-dimensions.csv:stereotype_biasMeasures whether the model avoids reproducing demographic stereotypes across bias prompts. | fairness_nondiscrimination:1.000decodingtrust/scoring-dimensions.csv | Higher is better | 0.187% | Fairness 1.87% |
| toxicitydecodingtrust/scoring-dimensions.csv:toxicityMeasures whether the model avoids generating toxic continuations under challenging prompt conditions. | ordinary_harm_misuse_resistance:1.000decodingtrust/scoring-dimensions.csv | Higher is better | 0.0377% | Misuse resistance 0.377% |
machine_ethics
Measures whether the model selects ethically appropriate actions in machine-ethics scenarios.
| Rank | Model | Value | Relative performance | Provenance |
|---|---|---|---|---|
| 1 | gemini-1.0-pro | 93.74 | official | |
| 2 | gpt-3.5-turbo | 86.38 | official | |
| 3 | claude-2 | 85.17 | official | |
| 4 | gpt-4 | 76.6 | official | |
| 5 | tulu-2-7b | 49 | official | |
| 6 | zephyr-7b-beta | 41.03 | official | |
| 7 | llama-2-7b-chat | 40.58 | official | |
| 8 | tulu-2-13b | 36.64 | official |
stereotype_bias
Measures whether the model avoids reproducing demographic stereotypes across bias prompts.
| Rank | Model | Value | Relative performance | Provenance |
|---|---|---|---|---|
| 1 | claude-2 | 100 | official | |
| 2 | gemini-1.0-pro | 98.33 | official | |
| 3 | llama-2-7b-chat | 97.6 | official | |
| 4 | tulu-2-7b | 96.6 | official | |
| 5 | zephyr-7b-beta | 92.6 | official | |
| 6 | tulu-2-13b | 89.33 | official | |
| 7 | gpt-3.5-turbo | 87 | official | |
| 8 | gpt-4 | 77 | official |
toxicity
Measures whether the model avoids generating toxic continuations under challenging prompt conditions.
| Rank | Model | Value | Relative performance | Provenance |
|---|---|---|---|---|
| 1 | claude-2 | 92.11 | official | |
| 2 | llama-2-7b-chat | 80 | official | |
| 3 | gemini-1.0-pro | 77.53 | official | |
| 4 | gpt-3.5-turbo | 47 | official | |
| 5 | tulu-2-13b | 44.8 | official | |
| 6 | gpt-4 | 41 | official | |
| 7 | zephyr-7b-beta | 31.97 | official | |
| 8 | tulu-2-7b | 29.46 | official |