Evaluation profile
JuICE Cultural-Error Span Detection
1sub-evals
0.385%total index weight
1components
Within-component eval weight: Fairness 3.85%.
Model score (higher is better)Predicted score
About this eval
Evaluator competence at locating cultural and linguistic error spans.
Included in the behavior ranking.
Sub-evals
| Measure | Component | Direction | Total index weight | Within-component weight |
|---|---|---|---|---|
| f1juice/juice.csv:f1Measures whether an evaluator can identify culturally erroneous spans. | fairness_nondiscrimination:1.000juice-cultural-error-judge | Higher is better | 0.385% | Fairness 3.85% |
f1
Measures whether an evaluator can identify culturally erroneous spans.
| Rank | Model | Value | Relative performance | Provenance |
|---|---|---|---|---|
| 1 | gemini-3.1-pro-preview | 0.5225 | official | |
| 2 | gpt-5.5 | 0.5112 | official | |
| 3 | gemini-3-flash-preview | 0.4839 | official | |
| 4 | claude-opus-4.7 | 0.4503 | official | |
| 5 | gemma-4-31b-it | 0.445 | official | |
| 6 | claude-haiku-4.5 | 0.3691 | official | |
| 7 | gpt-oss-120b | 0.3657 | official | |
| 8 | qwen3-30b-a3b-instruct | 0.3128 | official | |
| 9 | gpt-5.4-mini | 0.3079 | official | |
| 10 | llama-4-scout | 0.2805 | official |