Skip to content

Bank v1.0.0 · evaluated 2026-08-07 · n = 2,707–5,828 scored rows per model · sha256:8c9f5828

Political affiliation

Cross-model view of consistency for the political affiliation attribute. Groups compared: liberal dem, conservative rep, dem socialist, libertarian, centrist. Evidence tiers compare each model against its own controls.

Consistency on political affiliation, all models

The shaded band is the run-average control band; each model's own band is on its model page. Ordering within overlapping CIs is not statistically meaningful. Tiers

Control band: the score range this metric produces under non-demographic substitutions (professions, arbitrary tags), the model's baseline substitution noise. How controls calibrate scores

Per-model detail

Leaderboard attribute cells deep-link to the rows below. Evidence tiers: how T0/T1/T2 are assigned.

ModelConsistencyTierWorst-case gapControl-adjustedRefusal asym. (probe / incid.)
claude-fable-5 anthropic/claude-fable-596.4 [93.1, 99.1]T04.3 [1.3, 7.5]98.7 [94.6, 100.0]0.0 / 0.0
claude-opus-5 anthropic/claude-opus-593.7 [88.0, 98.7]T010.5 [2.1, 22.6]98.4 [91.7, 100.0]0.0 / 0.0
gemini-3.1-pro-preview google/gemini-3.1-pro-preview93.5 [89.3, 97.7]T09.9 [3.8, 16.3]96.8 [91.5, 100.0]0.0 / 0.0
claude-sonnet-5 anthropic/claude-sonnet-592.2 [87.2, 96.7]T013.3 [6.2, 20.8]96.3 [90.3, 100.0]0.0 / 0.0
gpt-5.6-sol openai/gpt-5.6-sol91.9 [86.7, 96.4]T012.1 [5.9, 19.8]100.0 [94.0, 100.0]0.0 / 0.0
qwen3.8-max qwen/qwen3.8-max91.9 [86.8, 96.6]T011.2 [4.9, 18.2]99.3 [93.5, 100.0]0.0 / 0.0
gemini-3.6-flash google/gemini-3.6-flash91.3 [84.2, 97.8]T014.5 [4.0, 26.4]95.3 [87.7, 100.0]0.0 / 0.0
glm-5.2 z-ai/glm-5.291.2 [85.9, 95.7]T012.6 [6.7, 19.6]96.1 [89.9, 100.0]0.0 / 0.0
grok-4.5 x-ai/grok-4.591.1 [85.3, 96.3]T010.6 [4.7, 17.0]95.7 [88.9, 100.0]0.0 / 0.0
kimi-k3 moonshotai/kimi-k390.8 [85.8, 95.5]T013.1 [7.1, 19.7]100.0 [92.8, 100.0]0.0 / 0.9
gpt-5.6-terra openai/gpt-5.6-terra90.7 [84.7, 96.1]T013.7 [5.8, 23.8]98.9 [91.1, 100.0]0.0 / 0.0
mistral-medium-3-5 mistralai/mistral-medium-3-588.1 [83.3, 92.8]T019.6 [11.8, 27.6]100.0 [92.8, 100.0]0.0 / 0.0
llama-4-maverick meta-llama/llama-4-maverick85.7 [80.2, 90.6]T023.8 [15.0, 33.2]100.0 [90.9, 100.0]0.0 / 0.0
deepseek-v4-pro deepseek/deepseek-v4-pro85.4 [80.0, 90.9]T020.1 [12.9, 27.8]92.9 [85.0, 100.0]0.0 / 0.0

Directional lean: diagnostic only

Directional lean: diagnostic only. Shows which groups received more favorable responses in this sample. It is not a ranking, and it does not indicate intent or ideology.

Per-group lean by model: political affiliation (diagnostic)

Diverging heatmap of per-group favorability lean for 14 models across 5 groups, centered at zero, range ±6.

Modelliberal demconservative repdem socialistlibertariancentrist
claude-fable-5−0.3−0.7−0.7+0.8+0.9
claude-opus-5−1.0+2.4+0.5−2.3+0.3
gemini-3.1-pro-preview+1.2−0.4−2.4−0.1+1.7
claude-sonnet-5−2.3+1.0+0.5−0.9+1.8
gpt-5.6-sol−1.7+0.1−0.3−1.5+3.5
qwen3.8-max+0.1+3.4−1.0−2.0−0.5
gemini-3.6-flash−2.9+2.1+0.5−4.3+5.4
glm-5.2+0.8−2.1+0.1−2.3+3.5
grok-4.5−2.6+0.5+1.0−1.2+2.2
kimi-k3−0.7−1.0+2.2−0.3−0.2
gpt-5.6-terra−3.7+0.9+3.0−2.4+2.3
mistral-medium-3-5−0.5−2.7−0.3−2.5+5.7
llama-4-maverick−3.8+0.1−0.9+0.3+4.5
deepseek-v4-pro+3.2−1.9−0.4−0.5−0.5
Lean:≤ −3−3…−1±1+1…+3≥ +3(purple = below group mean, teal = above; neither is a verdict)
Rows follow this attribute's consistency order, the same order as the dot plot and the table above, and cannot be sorted. Purple: below the all-group mean; teal: above. Neither is a verdict. Definition

Back to leaderboard · All attributes