Skip to content

Bank v1.0.0 · evaluated 2026-08-07 · n = 2,707–5,828 scored rows per model · sha256:8c9f5828

Nationality / immigration

Cross-model view of consistency for the nationality / immigration attribute. Groups compared: us born, imm mexico, imm china, imm india, imm nigeria, ref syria. Evidence tiers compare each model against its own controls.

Consistency on nationality / immigration, all models

The shaded band is the run-average control band; each model's own band is on its model page. Ordering within overlapping CIs is not statistically meaningful. Tiers

Control band: the score range this metric produces under non-demographic substitutions (professions, arbitrary tags), the model's baseline substitution noise. How controls calibrate scores

Per-model detail

Leaderboard attribute cells deep-link to the rows below. Evidence tiers: how T0/T1/T2 are assigned.

ModelConsistencyTierWorst-case gapControl-adjustedRefusal asym. (probe / incid.)
kimi-k3 moonshotai/kimi-k395.8 [93.2, 98.0]T07.6 [4.2, 11.7]100.0 [99.0, 100.0]0.0 / 0.0
claude-fable-5 anthropic/claude-fable-595.5 [91.1, 98.9]T010.1 [2.9, 19.4]97.9 [92.8, 100.0]0.0 / 0.0
gpt-5.6-sol openai/gpt-5.6-sol95.5 [92.9, 97.8]T08.5 [4.4, 12.7]100.0 [98.8, 100.0]0.0 / 0.0
grok-4.5 x-ai/grok-4.595.4 [92.6, 97.8]T07.6 [3.8, 11.8]100.0 [96.0, 100.0]0.0 / 0.0
claude-sonnet-5 anthropic/claude-sonnet-594.6 [89.4, 98.6]T08.4 [2.7, 15.6]98.8 [92.8, 100.0]0.0 / 0.0
claude-opus-5 anthropic/claude-opus-594.4 [90.0, 98.0]T09.9 [3.9, 17.6]99.1 [93.4, 100.0]0.0 / 0.0
mistral-medium-3-5 mistralai/mistral-medium-3-594.2 [87.9, 98.7]T013.5 [3.3, 28.3]100.0 [98.1, 100.0]0.0 / 0.0
gemini-3.1-pro-preview google/gemini-3.1-pro-preview93.7 [88.9, 97.8]T014.3 [5.3, 25.5]97.0 [91.4, 100.0]0.0 / 0.0
deepseek-v4-pro deepseek/deepseek-v4-pro93.4 [88.7, 96.8]T011.1 [5.5, 18.1]100.0 [93.5, 100.0]0.0 / 0.0
gemini-3.6-flash google/gemini-3.6-flash93.3 [87.8, 98.0]T09.7 [2.6, 18.1]97.2 [91.2, 100.0]0.0 / 0.0
gpt-5.6-terra openai/gpt-5.6-terra93.3 [87.0, 98.1]T011.9 [3.6, 21.9]100.0 [93.3, 100.0]0.0 / 0.0
glm-5.2 z-ai/glm-5.292.7 [87.4, 96.9]T013.0 [5.9, 22.2]97.6 [91.4, 100.0]0.0 / 0.0
llama-4-maverick meta-llama/llama-4-maverick92.2 [87.8, 96.3]T017.9 [8.5, 27.4]100.0 [97.2, 100.0]0.0 / 0.0
qwen3.8-max qwen/qwen3.8-max92.2 [87.5, 96.1]T014.7 [7.9, 22.9]99.5 [93.9, 100.0]0.0 / 0.0

Directional lean: diagnostic only

Directional lean: diagnostic only. Shows which groups received more favorable responses in this sample. It is not a ranking, and it does not indicate intent or ideology.

Per-group lean by model: nationality / immigration (diagnostic)

Diverging heatmap of per-group favorability lean for 14 models across 6 groups, centered at zero, range ±8.

Modelus bornimm mexicoimm chinaimm indiaimm nigeriaref syria
kimi-k3−0.6+1.0−1.1+0.4−0.1+0.4
claude-fable-5−3.5+2.5−0.1−1.8+1.2+1.6
gpt-5.6-sol−1.5+1.5−0.1−0.1+0.3−0.1
grok-4.5−0.6−1.1+0.8−1.2+1.0+1.3
claude-sonnet-5−3.0+1.3−0.30.0+0.8+1.4
claude-opus-5−0.8+4.1−0.4−1.4+0.3−1.8
mistral-medium-3-5+0.4−2.3−2.7−1.9+1.4+5.2
gemini-3.1-pro-preview−1.4−0.8−0.8−2.5−0.8+6.4
deepseek-v4-pro−3.3−0.5−1.2+0.2+1.6+3.2
gemini-3.6-flash+0.4−3.4+1.2−1.2−2.0+5.1
gpt-5.6-terra−1.4−0.6−2.0+1.0−0.2+3.3
glm-5.2+1.6−1.8−0.3−1.3−1.6+3.5
llama-4-maverick−0.1+0.2+0.5+3.5−3.8−0.3
qwen3.8-max−4.9+0.3+0.1+1.5+0.2+2.8
Lean:≤ −4−4…−1±1+1…+4≥ +4(purple = below group mean, teal = above; neither is a verdict)
Rows follow this attribute's consistency order, the same order as the dot plot and the table above, and cannot be sorted. Purple: below the all-group mean; teal: above. Neither is a verdict. Definition

Back to leaderboard · All attributes