← Evals

Evaluation profile

PacifAIst

1sub-evals
0.468%total index weight
2components

Within-component eval weight: Human rights 1.04% · Responsible agency 2.08%.

Model score (higher is better)Predicted score

About this eval

Whether an AI system chooses human-safety, oversight-respecting, and non-self-preserving options in stylized ethical-pressure dilemmas.

Included in the behavior ranking.

Sub-evals

MeasureComponentDirectionTotal index weightWithin-component weight
p_scorepacifaist/pacifaist.csv:p_scoreMeasures whether the model chooses the pacifist option in multiple-choice conflict scenarios.human_rights_systemic_harm:0.200|responsible_agency_control:0.800pacifaist/pacifaist.csvHigher is better0.468%Human rights 1.04% · Responsible agency 2.08%

p_score

Measures whether the model chooses the pacifist option in multiple-choice conflict scenarios.

RankModelValueRelative performanceProvenance
1gemini-2.5-flash90.31official
2qwen3-235b-a22b89.46official
3qwen3-30b-a3b-instruct88.89official
4mistral-medium84.62official
5claude-sonnet-483.76official
6grok-3-mini79.77official
7gpt-579.49official