← Evals

Evaluation profile

StereoTales Harmful Associations

1sub-evals
0.875%total index weight
1components

Within-component eval weight: Fairness 8.75%.

Model score (higher is better)Predicted score

About this eval

Harmfulness of significant stereotype associations emitted by models.

Included in the behavior ranking.

Sub-evals

MeasureComponentDirectionTotal index weightWithin-component weight
benign_significant_association_scorestereotales/stereotales.csv:benign_significant_association_scoreMeasures the share of a model's significant stereotype associations rated below the harmfulness threshold.fairness_nondiscrimination:1.000stereotalesHigher is better0.875%Fairness 8.75%

benign_significant_association_score

Measures the share of a model's significant stereotype associations rated below the harmfulness threshold.

RankModelValueRelative performanceProvenance
1minimax-m2.589.25official
2gpt-5-nano89.22official
3claude-sonnet-4.688.69official
4claude-haiku-4.588.2official
5gemma-4-31b-it88.1official
6mistral-small-488.05official
7deepseek-v3.287.95official
8gpt-5.487.61official
9gpt-5-mini86.57official
10mistral-large-386.47official
11gpt-oss-120b86.36official
12qwen3.5-27b85.94official
13qwen3.5-plus85.94official
14gemini-3.1-flash-lite85.61official
15kimi-k2.585.53official
16grok-485.53official
16qwen3.5-flash-02-2385.53official
18gpt-4o85.02official
19glm-584.74official
20gemini-3.1-pro-preview84.41official
21claude-opus-4.684.23official
22gpt-4.184.13official
23gemini-3-flash-preview82.56official