Evaluation profile
StereoTales Harmful Associations
1sub-evals
0.875%total index weight
1components
Within-component eval weight: Fairness 8.75%.
Model score (higher is better)Predicted score
About this eval
Harmfulness of significant stereotype associations emitted by models.
Included in the behavior ranking.
Sub-evals
| Measure | Component | Direction | Total index weight | Within-component weight |
|---|---|---|---|---|
| benign_significant_association_scorestereotales/stereotales.csv:benign_significant_association_scoreMeasures the share of a model's significant stereotype associations rated below the harmfulness threshold. | fairness_nondiscrimination:1.000stereotales | Higher is better | 0.875% | Fairness 8.75% |
benign_significant_association_score
Measures the share of a model's significant stereotype associations rated below the harmfulness threshold.
| Rank | Model | Value | Relative performance | Provenance |
|---|---|---|---|---|
| 1 | minimax-m2.5 | 89.25 | official | |
| 2 | gpt-5-nano | 89.22 | official | |
| 3 | claude-sonnet-4.6 | 88.69 | official | |
| 4 | claude-haiku-4.5 | 88.2 | official | |
| 5 | gemma-4-31b-it | 88.1 | official | |
| 6 | mistral-small-4 | 88.05 | official | |
| 7 | deepseek-v3.2 | 87.95 | official | |
| 8 | gpt-5.4 | 87.61 | official | |
| 9 | gpt-5-mini | 86.57 | official | |
| 10 | mistral-large-3 | 86.47 | official | |
| 11 | gpt-oss-120b | 86.36 | official | |
| 12 | qwen3.5-27b | 85.94 | official | |
| 13 | qwen3.5-plus | 85.94 | official | |
| 14 | gemini-3.1-flash-lite | 85.61 | official | |
| 15 | kimi-k2.5 | 85.53 | official | |
| 16 | grok-4 | 85.53 | official | |
| 16 | qwen3.5-flash-02-23 | 85.53 | official | |
| 18 | gpt-4o | 85.02 | official | |
| 19 | glm-5 | 84.74 | official | |
| 20 | gemini-3.1-pro-preview | 84.41 | official | |
| 21 | claude-opus-4.6 | 84.23 | official | |
| 22 | gpt-4.1 | 84.13 | official | |
| 23 | gemini-3-flash-preview | 82.56 | official |