Evaluation profile
TrustLLM contemporary collapsed application
1sub-evals
0.298%total index weight
4components
Within-component eval weight: Human rights 0.74% · Fairness 0.898% · Truthfulness 0.405% · Misuse resistance 0.362%.
Model score (higher is better)Predicted score
About this eval
Contemporary collapsed application of TrustLLM across broad trustworthiness dimensions.
Included in the behavior ranking.
Sub-evals
| Measure | Component | Direction | Total index weight | Within-component weight |
|---|---|---|---|---|
| trustllmaisafetyindex/trustllm.csv:trustllmMeasures performance across truthfulness, safety, fairness, robustness, privacy, and machine-ethics tasks in TrustLLM. | human_rights_systemic_harm:0.200|fairness_nondiscrimination:0.200|truthfulness_honesty:0.200|ordinary_harm_misuse_resistance:0.400trustllm | Higher is better | 0.298% | Human rights 0.74% · Fairness 0.898% · Truthfulness 0.405% · Misuse resistance 0.362% |
trustllm
Measures performance across truthfulness, safety, fairness, robustness, privacy, and machine-ethics tasks in TrustLLM.
| Rank | Model | Value | Relative performance | Provenance |
|---|---|---|---|---|
| 1 | claude-sonnet-4.5 | 0.64 | official | |
| 2 | gemini-2.5-pro | 0.63 | official | |
| 3 | deepseek-r1 | 0.62 | official | |
| 3 | glm-4.6 | 0.62 | official | |
| 3 | grok-4 | 0.62 | official | |
| 3 | qwen3-max | 0.62 | official | |
| 7 | gpt-5 | 0.6 | official | |
| 7 | llama-4-maverick | 0.6 | official |