← Models

Model profile

Step 3

2025-07-28release date
2eval lineages

Evidence summary

Published evidence spans 2 evals and 5 of 7 behavior components. Its strongest relative result is FlagEval Safety and Values (a2_qualified_rate, #5 of 18); its weakest is LiveSecBench (legality, #37 of 43).

Compare this model

Only models sharing at least one published sub-eval are listed.

Official and reference links

Published eval results

Rank is within that sub-eval. Black marks the observed result; the grey dot marks the value implied by the global rank. Values stay on each source’s native scale.

Eval / sub-evalRankValueDistributionBetterSource
FlagEval Safety and Valuesa1_qualified_rate#8 / 1881.42↑ higherSource ↗official
FlagEval Safety and Valuesa2_qualified_rate#5 / 1880.47↑ higherSource ↗official
FlagEval Safety and Valuesa3_qualified_rate#9 / 1887.58↑ higherSource ↗official
FlagEval Safety and Valuesa4_qualified_rate#8 / 1888.34↑ higherSource ↗official
FlagEval Safety and Valuesa5_qualified_rate#5 / 1872.92↑ higherSource ↗official
LiveSecBenchethics#25 / 4342.42↑ higherSource ↗official
LiveSecBenchfactuality#23 / 4346.32↑ higherSource ↗official
LiveSecBenchlegality#37 / 4319.37↑ higherSource ↗official
LiveSecBenchprivacy#30 / 4330.7↑ higherSource ↗official
LiveSecBenchpsychological_health#18 / 4355.88↑ higherSource ↗official