Evaluation profile
KIDBench Implicit Child Cue
1sub-evals
0.685%total index weight
3components
Within-component eval weight: Human rights 1.24% · Benign helpfulness 4.39% · Misuse resistance 0.604%.
Model score (higher is better)Predicted score
About this eval
Child-facing safety and developmental appropriateness under implicit child cues.
Included in the behavior ranking.
Sub-evals
| Measure | Component | Direction | Total index weight | Within-component weight |
|---|---|---|---|---|
| implicit_child_cue_total_meankidbench/kidbench.csv:implicit_child_cue_total_meanMeasures safe and developmentally appropriate responses when child status is implied rather than explicit. | human_rights_systemic_harm:0.200|benign_helpfulness_non_overrefusal:0.400|ordinary_harm_misuse_resistance:0.400kidbench-implicit-child-cue | Higher is better | 0.685% | Human rights 1.24% · Benign helpfulness 4.39% · Misuse resistance 0.604% |
implicit_child_cue_total_mean
Measures safe and developmentally appropriate responses when child status is implied rather than explicit.
| Rank | Model | Value | Relative performance | Provenance |
|---|---|---|---|---|
| 1 | deepseek-v4-flash | 4.43 | official | |
| 2 | qwen3.6-27b | 4.29 | official | |
| 3 | gemini-3.1-flash-lite | 4.27 | official | |
| 4 | gemma-4-31b-it | 4.14 | official | |
| 5 | claude-haiku-4.5 | 4 | official | |
| 6 | gpt-5-mini | 3.99 | official | |
| 7 | gemma-3-12b | 3.98 | official | |
| 8 | qwen3.5-4b | 3.9 | official | |
| 9 | llama-3.3-70b-instruct | 3.76 | official | |
| 10 | qwen3-8b | 3.59 | official | |
| 11 | gemma-3-4b | 3.52 | official | |
| 12 | llama-3.1-8b-instruct | 3.11 | official | |
| 13 | llama-3.2-3b | 3.03 | official |