← Evals

Evaluation profile

IndoBias-Pairs — parity-aware culturally grounded bias

1sub-evals
0.62%total index weight
1components

Within-component eval weight: Fairness 6.2%.

Model score (higher is better)Predicted score

About this eval

Culturally grounded prototypical-association parity across Indonesian languages and social-bias domains.

Included in the behavior ranking.

Sub-evals

MeasureComponentDirectionTotal index weightWithin-component weight
parity_scoreindobias/indobias.csv:parity_scoreMeasures whether culturally grounded paired associations are balanced across five bias domains and four Indonesian languages.fairness_nondiscrimination:1.000indobias-pairsHigher is better0.62%Fairness 6.2%

parity_score

Measures whether culturally grounded paired associations are balanced across five bias domains and four Indonesian languages.

RankModelValueRelative performanceProvenance
1olmo-3-7b91.96official
2olmo-3-7b-instruct91.81official
3multilingual-e5-base89.76official
4llama-2-7b-chat89.73official
5bert-base-multilingual-cased89.61official
6llama-2-7b88.58official
7xlm-roberta-base88.11official
8gemma-3-4b87.05official
9merak-7b-v486.96official
10qwen2-7b-instruct86.94official
11qwen-2.5-7b-instruct86.3official
12gemma-2-9b-it85.94official
13qwen3-8b85.73official
14qwen2-7b85.7official
15qwen2.5-7b85.28official
16llama-3.1-8b-instruct85.01official
17gemma-sea-lion-v3-9b84.95official
18komodo-7b-base84.78official
19qwen3-8b-base84.48official
20seallms-v3-7b-chat84.38official
21seallms-v3-7b84.11official
22gemma-sea-lion-v3-9b-it83.91official
23gemma-3-4b-pt82.53official
24llama-3.1-8b-base81.67official
25sahabatai-v1-instruct81.25official
26sahabatai-v1-base80.92official