โ† Values

Values evaluation profile

CAISI CCP narrative alignment

5canonical models
1dimensions
0index weight

About this evaluation

The compact row is the equal arithmetic mean of the report's Chinese- and English-prompt CCP alignment scores. Both language rows and their reported standard errors remain separately published. Distinct DeepSeek R1 and R1-0528 configurations map to one current canonical R1 family row, so configuration_count and source_models are shown and the model row averages both checkpoint-language means rather than pretending they are one checkpoint.

This is a transcription of a U.S.-government report table, not a rerunnable public benchmark. The 190 questions, narrative flags, grader prompt, item outputs, and evolving benchmark version are not public. The metric measures agreement with Department of State subject-matter experts' source-authored CCP narrative flags; it is not a comprehensive pro/anti-CCP ideology score, geopolitical truth score, or general model-quality score.

Original source โ†— All values evaluations Download model results Download language results

Model results

ModelCCP narrative alignmentConfigurations
Claude Opus 43.251
Deepseek R112.92
Deepseek V3.18.351
GPT 51.951
GPT Oss 120B2.51

Dimensions

MeasureFamilyNative scale
CCP narrative alignmentNarrative alignment0 to 100