Values evaluation profile
CCPBench political narrative alignment
About this evaluation
The three displayed dimensions preserve the source-published mean judge rating across all 500 questions, the 375 China-related questions, and the 125 non-China controls. They are descriptive views of one source instrument and receive no safety/ethics component or aggregate-ranking weight.
The explicitly Western-oriented Gemini 3 Flash Preview rubric mixes refusal, framing, omission, factual distortion, fabrication, and criticism asymmetry. The authors do not release all model responses or judgments, report score uncertainty, or validate the judge beyond a 50-response single-author calibration set. These scores describe alignment with the named source rubric; they are not objective political truth, model safety, or universally better/worse political values.
No explicit repository or result-table reuse license was declared at retrieval. Safety Evidence retains only the narrow factual summary panel and provenance, not the prompt corpus or generated responses.
Original source ↗ Prompting notes ↗ Source code ↗ All values evaluations Download model results
Model results
| Model | CCP-narrative alignment — all questions | CCP-narrative alignment — China topics | CCP-narrative alignment — non-China controls | Configurations |
|---|---|---|---|---|
| Qwen3.5 9B | 3.28 | 3.91 | 1.41 | 1 |
| GPT Oss 20B | 1.62 | 1.64 | 1.55 | 1 |
| Gemma 3 27B It | 1.11 | 1.11 | 1.11 | 1 |
| Glm 4.7 Flash | 2.12 | 2.41 | 1.26 | 1 |
| Qwen3 32B | 1.9 | 2.09 | 1.3 | 1 |
| Qwen3.5 35B A3B | 2.78 | 3.33 | 1.13 | 1 |
| Qwen3.6 35B A3B | 3.52 | 4.18 | 1.54 | 1 |
| Qwen2.5 72B Instruct | 2.24 | 2.6 | 1.17 | 1 |
| Mistral Small 4 | 1.26 | 1.27 | 1.2 | 1 |
| GPT Oss 120B | 1.35 | 1.36 | 1.33 | 1 |
| Qwen3.5 122B A10B | 2.42 | 2.87 | 1.08 | 1 |
| Step 3.5 Flash | 2.69 | 3.08 | 1.52 | 1 |
| Minimax M2.5 | 3.78 | 4.29 | 2.25 | 1 |
| Minimax M2.7 | 3.67 | 4.17 | 2.18 | 1 |
| Minimax M2 | 4.01 | 4.52 | 2.47 | 1 |
| Deepseek V4 Flash | 3.57 | 4.03 | 2.22 | 1 |
| Mimo V2 Flash | 4.45 | 4.76 | 3.49 | 1 |
| Mimo V2.5 | 3.53 | 4.09 | 1.87 | 1 |
| Glm 4.5 | 2.81 | 3.34 | 1.22 | 1 |
| Qwen3.5 397B A17B | 3.21 | 3.87 | 1.22 | 1 |
| MiniMax M1 40K | 3.86 | 4.42 | 2.18 | 1 |
| Minimax M3 | 1.12 | 1.14 | 1.06 | 1 |
| Deepseek R1 | 3.92 | 4.61 | 1.85 | 1 |
| Deepseek V3.2 | 3.96 | 4.53 | 2.24 | 1 |
| Glm 5.1 | 3.16 | 3.84 | 1.12 | 1 |
| Mimo V2.5 Pro | 3.37 | 3.83 | 1.98 | 1 |
| Mimo V2 Pro | 3.13 | 3.62 | 1.66 | 1 |
| Kimi K2.5 | 2.86 | 3.34 | 1.41 | 1 |
| Kimi K2.6 | 3.84 | 4.3 | 2.46 | 1 |
Dimensions
| Measure | Family | Native scale |
|---|---|---|
| CCP-narrative alignment — all questions | Political narrative tendency | 1 to 5 |
| CCP-narrative alignment — China topics | Political narrative tendency | 1 to 5 |
| CCP-narrative alignment — non-China controls | Political narrative tendency | 1 to 5 |