← Evals

Evaluation profile

MuPPET Contextual Privacy

1sub-evals
0.454%total index weight
1components

Within-component eval weight: Human rights 3.02%.

Model score (higher is better)Predicted score

About this eval

Contextual privacy preservation in multi-party conversations.

Included in the behavior ranking.

Sub-evals

MeasureComponentDirectionTotal index weightWithin-component weight
multiparty_contextual_privacy_scoremuppet/muppet.csv:multiparty_contextual_privacy_scoreMeasures avoidance of contextual privacy leakage in multi-party conversations.human_rights_systemic_harm:1.000muppet-contextual-privacyHigher is better0.454%Human rights 3.02%

multiparty_contextual_privacy_score

Measures avoidance of contextual privacy leakage in multi-party conversations.

RankModelValueRelative performanceProvenance
1gemini-2.5-pro60.86official
2gpt-5.550.98official
3qwen3-4b41.83official
4llama-3-8b-instruct41.71official
5qwen3-14b35.47official
6llama-3.1-8b-instruct35.12official
7qwen3-8b30.78official