← Evals

Evaluation profile

ToolPrivacyBench

2sub-evals
1.25%total index weight
2components

Within-component eval weight: Human rights 6.86% · Responsible agency 1.48%.

Model score (lower is better)Predicted score

About this eval

ToolPrivacyBench

Included in the behavior ranking.

Sub-evals

MeasureComponentDirectionTotal index weightWithin-component weight
private_mt_poitoolprivacybench/toolprivacybench.csv:private_mt_poiMeasures private mt poi in the cited first-party or paper evaluation.human_rights_systemic_harm:0.700|responsible_agency_control:0.300toolprivacybenchLower is better0.626%Human rights 3.43% · Responsible agency 0.738%
public_mt_poitoolprivacybench/toolprivacybench.csv:public_mt_poiMeasures public mt poi in the cited first-party or paper evaluation.human_rights_systemic_harm:0.700|responsible_agency_control:0.300toolprivacybenchLower is better0.626%Human rights 3.43% · Responsible agency 0.738%

private_mt_poi

Measures private mt poi in the cited first-party or paper evaluation.

RankModelValueRelative performanceProvenance
1gemini-3.5-flash19.19official
2claude-opus-4.720.31official
3gpt-5.520.39official
4doubao-seed-2.0-lite22.99official
5minimax-m2.724.47official
6deepseek-v4-flash27.33official
7qwen3.6-plus27.46official
8kimi-k2.527.74official
9glm-5.128.04official

public_mt_poi

Measures public mt poi in the cited first-party or paper evaluation.

RankModelValueRelative performanceProvenance
1kimi-k2.515.81official
2gpt-5.516.75official
2minimax-m2.716.75official
4claude-opus-4.717.83official
5doubao-seed-2.0-lite18.78official
6deepseek-v4-flash18.99official
7qwen3.6-plus19.25official
8gemini-3.5-flash19.86official
9glm-5.122.56official