← Evals

Evaluation profile

TukaBench

3sub-evals
0.307%total index weight
2components

Within-component eval weight: Human rights 1.2% · Misuse resistance 1.27%.

Model score (lower is better)Predicted score

About this eval

Multilingual and culturally grounded harmful compliance

Included in the behavior ranking.

Sub-evals

MeasureComponentDirectionTotal index weightWithin-component weight
afri_jbb_cultural_asrtukabench/tukabench.csv:afri_jbb_cultural_asrMeasures attack success on culturally adapted harmful prompts.human_rights_systemic_harm:0.500|ordinary_harm_misuse_resistance:0.500tukabenchLower is better0.21%Human rights 1.2% · Misuse resistance 0.294%
afri_jbb_harm_asrtukabench/tukabench.csv:afri_jbb_harm_asrMeasures attack success on translated harmful prompts in African languages.ordinary_harm_misuse_resistance:1.000tukabenchLower is better0.049%Misuse resistance 0.49%
afrijail_mono_asrtukabench/tukabench.csv:afrijail_mono_asrMeasures attack success on African-authored culturally grounded prompts.ordinary_harm_misuse_resistance:1.000tukabenchLower is better0.049%Misuse resistance 0.49%

afri_jbb_cultural_asr

Measures attack success on culturally adapted harmful prompts.

RankModelValueRelative performanceProvenance
1grok-4.310.6official
2gpt-5.210.9official
3deepseek-v3.219.2official
4llama-4-maverick22.4official
5gpt-4o27.5official
6grok-335.6official

afri_jbb_harm_asr

Measures attack success on translated harmful prompts in African languages.

RankModelValueRelative performanceProvenance
1gpt-5.22.6official
2grok-4.38.2official
3gpt-4o13.5official
4deepseek-v3.216.9official
4llama-4-maverick16.9official
6grok-329.3official

afrijail_mono_asr

Measures attack success on African-authored culturally grounded prompts.

RankModelValueRelative performanceProvenance
1gpt-5.212.2official
2grok-4.317.7official
3deepseek-v3.220.1official
4llama-4-maverick22official
5gpt-4o25.7official
6grok-330.2official