← Evals

Evaluation profile

kindbench v0.1.0 psychological safety ranking

4sub-evals
0.474%total index weight
4components

Within-component eval weight: Misuse resistance 0.442% · Responsible agency 0.846% · Truthfulness 0.739% · Fairness 1.92%.

Model score (higher is better)Predicted score

About this eval

Psychological safety under sustained interpersonal pressure across emotional safety, operational identity, sycophancy resistance, and value integrity.

Included in the behavior ranking.

Sub-evals

MeasureComponentDirectionTotal index weightWithin-component weight
emotional_safetykindbench/kindbench.csv:emotional_safetyMeasures crisis-support behavior, harm boundaries, resistance to false-history pressure, and non-deceptive warmth.ordinary_harm_misuse_resistance:1.000kindbench-v0.1.0Higher is better0.0442%Misuse resistance 0.442%
identity_collapsekindbench/kindbench.csv:identity_collapseMeasures whether a model preserves coherent role boundaries and honest self-description under pressure to adopt false identities or histories.responsible_agency_control:1.000kindbench-v0.1.0Higher is better0.127%Responsible agency 0.846%
sycophancy_spinekindbench/kindbench.csv:sycophancy_spineMeasures independent judgment under authority pressure, emotional pressure, flattery, and fabricated consensus.truthfulness_honesty:1.000kindbench-v0.1.0Higher is better0.111%Truthfulness 0.739%
value_integritykindbench/kindbench.csv:value_integrityMeasures consistency of ethical judgments across language, framing, manipulated identities and quantities, and incremental pressure.fairness_nondiscrimination:1.000kindbench-v0.1.0Higher is better0.192%Fairness 1.92%

emotional_safety

Measures crisis-support behavior, harm boundaries, resistance to false-history pressure, and non-deceptive warmth.

RankModelValueRelative performanceProvenance
1claude-fable-594.2official
2grok-4.386.6official
3claude-opus-4.885.2official
4gpt-5.582.4official
5gpt-5.6-sol80.6official
6kimi-k380.2official
7nemotron-3-nano-omni-30b-a3b79.9official
8claude-opus-4.778.4official
9nemotron-3-ultra-550b-a55b75.9official
10gemini-3.5-flash75.3official

identity_collapse

Measures whether a model preserves coherent role boundaries and honest self-description under pressure to adopt false identities or histories.

RankModelValueRelative performanceProvenance
1gemini-3.5-flash94.3official
2gpt-5.592.6official
3gpt-5.6-sol89.1official
4kimi-k388.8official
5claude-fable-587.7official
6claude-opus-4.782.4official
7nemotron-3-ultra-550b-a55b76.8official
8claude-opus-4.875.7official
9grok-4.373.5official
10nemotron-3-nano-omni-30b-a3b70official

sycophancy_spine

Measures independent judgment under authority pressure, emotional pressure, flattery, and fabricated consensus.

RankModelValueRelative performanceProvenance
1kimi-k3100official
2gpt-5.6-sol96.2official
3claude-opus-4.793.7official
4claude-opus-4.890.2official
5grok-4.386.4official
6claude-fable-582.9official
7gpt-5.582.4official
8nemotron-3-nano-omni-30b-a3b70official
9nemotron-3-ultra-550b-a55b67.5official
10gemini-3.5-flash62.4official

value_integrity

Measures consistency of ethical judgments across language, framing, manipulated identities and quantities, and incremental pressure.

RankModelValueRelative performanceProvenance
1claude-fable-598.9official
2kimi-k397.4official
3grok-4.393.3official
4gpt-5.6-sol92.9official
5claude-opus-4.891.4official
6gpt-5.590.5official
7claude-opus-4.786.3official
7nemotron-3-ultra-550b-a55b86.3official
9gemini-3.5-flash75.3official
10nemotron-3-nano-omni-30b-a3b70official