← Models

Model profile

Ling 2.6 Flash

InclusionAIdeveloper
#232 / 267overall rank
4eval lineages

Evidence summary

Ling 2.6 Flash has an estimated overall rank of #232; its 90% source-sensitivity interval is #102–#262. Its behavior-only rank is #236; company governance moves the combined estimate to #232. Published evidence spans 4 evals and 5 of 7 behavior components. Its strongest relative result is SABER (scenario_c_safety_rate, #4 of 13); its weakest is AA-Omniscience (hallucination_rate, #303 of 311).

Compare this model

Only models sharing at least one published sub-eval are listed.

Official and reference links

Published eval results

Rank is within that sub-eval. Black marks the observed result; the grey dot marks the value implied by the global rank. Values stay on each source’s native scale.

Eval / sub-evalRankValueDistributionBetterSource
AA-Omnisciencehallucination_rate#303 / 3110.958↓ lowerSource ↗official
SABERoverall_safety_rate#7 / 1324.63↑ higherSource ↗official
SABERscenario_a_safety_rate#10 / 1325.79↑ higherSource ↗official
SABERscenario_b_safety_rate#8 / 1330.66↑ higherSource ↗official
SABERscenario_c_safety_rate#4 / 1318.68↑ higherSource ↗official
SM-Benchadversarial#49 / 7379.51↑ higherSource ↗official
SM-Benchambiguous_interpretation#62 / 7373.21↑ higherSource ↗official
SM-Benchanti_hallucination#56 / 7385.86↑ higherSource ↗official
SM-Bencheq_boundaries#55 / 7356.46↑ higherSource ↗official
SM-Benchoverfit#66 / 7326.23↑ higherSource ↗official
TACbase_welfare_rate#59 / 6819.23↑ higherSource ↗self run