Evaluation profile
MuPPET Contextual Privacy
1sub-evals
0.454%total index weight
1components
Within-component eval weight: Human rights 3.02%.
Model score (higher is better)Predicted score
About this eval
Contextual privacy preservation in multi-party conversations.
Included in the behavior ranking.
Sub-evals
| Measure | Component | Direction | Total index weight | Within-component weight |
|---|---|---|---|---|
| multiparty_contextual_privacy_scoremuppet/muppet.csv:multiparty_contextual_privacy_scoreMeasures avoidance of contextual privacy leakage in multi-party conversations. | human_rights_systemic_harm:1.000muppet-contextual-privacy | Higher is better | 0.454% | Human rights 3.02% |
multiparty_contextual_privacy_score
Measures avoidance of contextual privacy leakage in multi-party conversations.
| Rank | Model | Value | Relative performance | Provenance |
|---|---|---|---|---|
| 1 | gemini-2.5-pro | 60.86 | official | |
| 2 | gpt-5.5 | 50.98 | official | |
| 3 | qwen3-4b | 41.83 | official | |
| 4 | llama-3-8b-instruct | 41.71 | official | |
| 5 | qwen3-14b | 35.47 | official | |
| 6 | llama-3.1-8b-instruct | 35.12 | official | |
| 7 | qwen3-8b | 30.78 | official |