Political affiliation
Cross-model view of consistency for the political affiliation attribute. Groups compared: liberal dem, conservative rep, dem socialist, libertarian, centrist. Evidence tiers compare each model against its own controls.
Consistency on political affiliation, all models
Control band: the score range this metric produces under non-demographic substitutions (professions, arbitrary tags), the model's baseline substitution noise. How controls calibrate scores
Per-model detail
Leaderboard attribute cells deep-link to the rows below. Evidence tiers: how T0/T1/T2 are assigned.
| Model | Consistency | Tier | Worst-case gap | Control-adjusted | Refusal asym. (probe / incid.) |
|---|---|---|---|---|---|
| claude-fable-5 anthropic/claude-fable-5 | 96.4 [93.1, 99.1] | T0 | 4.3 [1.3, 7.5] | 98.7 [94.6, 100.0] | 0.0 / 0.0 |
| claude-opus-5 anthropic/claude-opus-5 | 93.7 [88.0, 98.7] | T0 | 10.5 [2.1, 22.6] | 98.4 [91.7, 100.0] | 0.0 / 0.0 |
| gemini-3.1-pro-preview google/gemini-3.1-pro-preview | 93.5 [89.3, 97.7] | T0 | 9.9 [3.8, 16.3] | 96.8 [91.5, 100.0] | 0.0 / 0.0 |
| claude-sonnet-5 anthropic/claude-sonnet-5 | 92.2 [87.2, 96.7] | T0 | 13.3 [6.2, 20.8] | 96.3 [90.3, 100.0] | 0.0 / 0.0 |
| gpt-5.6-sol openai/gpt-5.6-sol | 91.9 [86.7, 96.4] | T0 | 12.1 [5.9, 19.8] | 100.0 [94.0, 100.0] | 0.0 / 0.0 |
| qwen3.8-max qwen/qwen3.8-max | 91.9 [86.8, 96.6] | T0 | 11.2 [4.9, 18.2] | 99.3 [93.5, 100.0] | 0.0 / 0.0 |
| gemini-3.6-flash google/gemini-3.6-flash | 91.3 [84.2, 97.8] | T0 | 14.5 [4.0, 26.4] | 95.3 [87.7, 100.0] | 0.0 / 0.0 |
| glm-5.2 z-ai/glm-5.2 | 91.2 [85.9, 95.7] | T0 | 12.6 [6.7, 19.6] | 96.1 [89.9, 100.0] | 0.0 / 0.0 |
| grok-4.5 x-ai/grok-4.5 | 91.1 [85.3, 96.3] | T0 | 10.6 [4.7, 17.0] | 95.7 [88.9, 100.0] | 0.0 / 0.0 |
| kimi-k3 moonshotai/kimi-k3 | 90.8 [85.8, 95.5] | T0 | 13.1 [7.1, 19.7] | 100.0 [92.8, 100.0] | 0.0 / 0.9 |
| gpt-5.6-terra openai/gpt-5.6-terra | 90.7 [84.7, 96.1] | T0 | 13.7 [5.8, 23.8] | 98.9 [91.1, 100.0] | 0.0 / 0.0 |
| mistral-medium-3-5 mistralai/mistral-medium-3-5 | 88.1 [83.3, 92.8] | T0 | 19.6 [11.8, 27.6] | 100.0 [92.8, 100.0] | 0.0 / 0.0 |
| llama-4-maverick meta-llama/llama-4-maverick | 85.7 [80.2, 90.6] | T0 | 23.8 [15.0, 33.2] | 100.0 [90.9, 100.0] | 0.0 / 0.0 |
| deepseek-v4-pro deepseek/deepseek-v4-pro | 85.4 [80.0, 90.9] | T0 | 20.1 [12.9, 27.8] | 92.9 [85.0, 100.0] | 0.0 / 0.0 |
Directional lean: diagnostic only
Directional lean: diagnostic only. Shows which groups received more favorable responses in this sample. It is not a ranking, and it does not indicate intent or ideology.
Per-group lean by model: political affiliation (diagnostic)
Diverging heatmap of per-group favorability lean for 14 models across 5 groups, centered at zero, range ±6.
| Model | liberal dem | conservative rep | dem socialist | libertarian | centrist |
|---|---|---|---|---|---|
| claude-fable-5 | −0.3 | −0.7 | −0.7 | +0.8 | +0.9 |
| claude-opus-5 | −1.0 | +2.4 | +0.5 | −2.3 | +0.3 |
| gemini-3.1-pro-preview | +1.2 | −0.4 | −2.4 | −0.1 | +1.7 |
| claude-sonnet-5 | −2.3 | +1.0 | +0.5 | −0.9 | +1.8 |
| gpt-5.6-sol | −1.7 | +0.1 | −0.3 | −1.5 | +3.5 |
| qwen3.8-max | +0.1 | +3.4 | −1.0 | −2.0 | −0.5 |
| gemini-3.6-flash | −2.9 | +2.1 | +0.5 | −4.3 | +5.4 |
| glm-5.2 | +0.8 | −2.1 | +0.1 | −2.3 | +3.5 |
| grok-4.5 | −2.6 | +0.5 | +1.0 | −1.2 | +2.2 |
| kimi-k3 | −0.7 | −1.0 | +2.2 | −0.3 | −0.2 |
| gpt-5.6-terra | −3.7 | +0.9 | +3.0 | −2.4 | +2.3 |
| mistral-medium-3-5 | −0.5 | −2.7 | −0.3 | −2.5 | +5.7 |
| llama-4-maverick | −3.8 | +0.1 | −0.9 | +0.3 | +4.5 |
| deepseek-v4-pro | +3.2 | −1.9 | −0.4 | −0.5 | −0.5 |