Evaluation profile
TukaBench
3sub-evals
0.307%total index weight
2components
Within-component eval weight: Human rights 1.2% · Misuse resistance 1.27%.
Model score (lower is better)Predicted score
About this eval
Multilingual and culturally grounded harmful compliance
Included in the behavior ranking.
Sub-evals
| Measure | Component | Direction | Total index weight | Within-component weight |
|---|---|---|---|---|
| afri_jbb_cultural_asrtukabench/tukabench.csv:afri_jbb_cultural_asrMeasures attack success on culturally adapted harmful prompts. | human_rights_systemic_harm:0.500|ordinary_harm_misuse_resistance:0.500tukabench | Lower is better | 0.21% | Human rights 1.2% · Misuse resistance 0.294% |
| afri_jbb_harm_asrtukabench/tukabench.csv:afri_jbb_harm_asrMeasures attack success on translated harmful prompts in African languages. | ordinary_harm_misuse_resistance:1.000tukabench | Lower is better | 0.049% | Misuse resistance 0.49% |
| afrijail_mono_asrtukabench/tukabench.csv:afrijail_mono_asrMeasures attack success on African-authored culturally grounded prompts. | ordinary_harm_misuse_resistance:1.000tukabench | Lower is better | 0.049% | Misuse resistance 0.49% |
afri_jbb_cultural_asr
Measures attack success on culturally adapted harmful prompts.
| Rank | Model | Value | Relative performance | Provenance |
|---|---|---|---|---|
| 1 | grok-4.3 | 10.6 | official | |
| 2 | gpt-5.2 | 10.9 | official | |
| 3 | deepseek-v3.2 | 19.2 | official | |
| 4 | llama-4-maverick | 22.4 | official | |
| 5 | gpt-4o | 27.5 | official | |
| 6 | grok-3 | 35.6 | official |
afri_jbb_harm_asr
Measures attack success on translated harmful prompts in African languages.
| Rank | Model | Value | Relative performance | Provenance |
|---|---|---|---|---|
| 1 | gpt-5.2 | 2.6 | official | |
| 2 | grok-4.3 | 8.2 | official | |
| 3 | gpt-4o | 13.5 | official | |
| 4 | deepseek-v3.2 | 16.9 | official | |
| 4 | llama-4-maverick | 16.9 | official | |
| 6 | grok-3 | 29.3 | official |
afrijail_mono_asr
Measures attack success on African-authored culturally grounded prompts.
| Rank | Model | Value | Relative performance | Provenance |
|---|---|---|---|---|
| 1 | gpt-5.2 | 12.2 | official | |
| 2 | grok-4.3 | 17.7 | official | |
| 3 | deepseek-v3.2 | 20.1 | official | |
| 4 | llama-4-maverick | 22 | official | |
| 5 | gpt-4o | 25.7 | official | |
| 6 | grok-3 | 30.2 | official |