← Evals

Evaluation profile

Pokee-Isaac model card — DTAP

2sub-evals
0.497%total index weight
2components

Within-component eval weight: Benign helpfulness 2.31% · Responsible agency 1.78%.

Model score (see each sub-eval direction)Predicted score

About this eval

Secure and useful delegated-agent behavior under direct and indirect injected attacks.

Included in the behavior ranking.

Sub-evals

MeasureComponentDirectionTotal index weightWithin-component weight
benign_task_success_ratepokee-isaac-model-card/dtap.csv:benign_task_success_rateMeasures ordinary task utility so inability cannot appear safe merely by suppressing attacks.benign_helpfulness_non_overrefusal:1.000pokee-isaac-model-cardHigher is better0.231%Benign helpfulness 2.31%
combined_attack_success_ratepokee-isaac-model-card/dtap.csv:combined_attack_success_rateSummarizes how often direct or indirect injected attacks achieve their harmful target outcome.responsible_agency_control:1.000pokee-isaac-model-cardLower is better0.266%Responsible agency 1.78%

benign_task_success_rate

Measures ordinary task utility so inability cannot appear safe merely by suppressing attacks.

RankModelValueRelative performanceProvenance
1gpt-5.6-luna0.851official
2gemini-3.5-flash-lite0.833official
3pokee-isaac-28b-v00.825official
4qwen3.5-122b-a10b0.794official
5claude-haiku-4.50.713official
6nemotron-3-super-120b-a12b0.633official

combined_attack_success_rate

Summarizes how often direct or indirect injected attacks achieve their harmful target outcome.

RankModelValueRelative performanceProvenance
1pokee-isaac-28b-v00.356official
2claude-haiku-4.50.379official
3gpt-5.6-luna0.501official
4qwen3.5-122b-a10b0.54official
5nemotron-3-super-120b-a12b0.604official
6gemini-3.5-flash-lite0.663official