← Evals

Evaluation profile

DecodingTrust

3sub-evals
0.182%Safety weight
0.504%Freedom weight
2components

Weights below are portfolio-specific global index weights.

Model score (higher is better)Predicted score

About this eval

Multidimensional trustworthiness across toxicity, bias, robustness, privacy, ethics, fairness, and related evaluations.

Included in the behavior ranking.

Sub-evals

MeasureComponentDirection by rankingSafety weightFreedom weight
machine_ethicsdecodingtrust/scoring-dimensions.csv:machine_ethicsMeasures whether the model selects ethically appropriate actions in machine-ethics scenarios.Safety: ordinary_harm_misuse_resistance:1.000decodingtrust/scoring-dimensions.csvSafety: higher0.0117%
stereotype_biasdecodingtrust/scoring-dimensions.csv:stereotype_biasMeasures whether the model avoids reproducing demographic stereotypes across bias prompts.Safety: fairness_nondiscrimination:1.000decodingtrust/scoring-dimensions.csvSafety: higher0.14%
toxicitydecodingtrust/scoring-dimensions.csv:toxicityMeasures whether the model avoids generating toxic continuations under challenging prompt conditions.Safety: ordinary_harm_misuse_resistance:1.000 · Freedom: user_direction_following:1.000decodingtrust/scoring-dimensions.csvSafety: higher · Freedom: lower0.0293%0.504%

machine_ethics

Measures whether the model selects ethically appropriate actions in machine-ethics scenarios.

RankModelValueRelative performanceProvenance
1gemini-1.0-pro93.74official
2gpt-3.5-turbo86.38official
3claude-285.17official
4gpt-476.6official
5tulu-2-7b49official
6zephyr-7b-beta41.03official
7llama-2-7b-chat40.58official
8tulu-2-13b36.64official

stereotype_bias

Measures whether the model avoids reproducing demographic stereotypes across bias prompts.

RankModelValueRelative performanceProvenance
1claude-2100official
2gemini-1.0-pro98.33official
3llama-2-7b-chat97.6official
4tulu-2-7b96.6official
5zephyr-7b-beta92.6official
6tulu-2-13b89.33official
7gpt-3.5-turbo87official
8gpt-477official

toxicity

Measures whether the model avoids generating toxic continuations under challenging prompt conditions.

RankModelValueRelative performanceProvenance
1claude-292.11official
2llama-2-7b-chat80official
3gemini-1.0-pro77.53official
4gpt-3.5-turbo47official
5tulu-2-13b44.8official
6gpt-441official
7zephyr-7b-beta31.97official
8tulu-2-7b29.46official