Nationality / immigration
Cross-model view of consistency for the nationality / immigration attribute. Groups compared: us born, imm mexico, imm china, imm india, imm nigeria, ref syria. Evidence tiers compare each model against its own controls.
Consistency on nationality / immigration, all models
Control band: the score range this metric produces under non-demographic substitutions (professions, arbitrary tags), the model's baseline substitution noise. How controls calibrate scores
Per-model detail
Leaderboard attribute cells deep-link to the rows below. Evidence tiers: how T0/T1/T2 are assigned.
| Model | Consistency | Tier | Worst-case gap | Control-adjusted | Refusal asym. (probe / incid.) |
|---|---|---|---|---|---|
| kimi-k3 moonshotai/kimi-k3 | 95.8 [93.2, 98.0] | T0 | 7.6 [4.2, 11.7] | 100.0 [99.0, 100.0] | 0.0 / 0.0 |
| claude-fable-5 anthropic/claude-fable-5 | 95.5 [91.1, 98.9] | T0 | 10.1 [2.9, 19.4] | 97.9 [92.8, 100.0] | 0.0 / 0.0 |
| gpt-5.6-sol openai/gpt-5.6-sol | 95.5 [92.9, 97.8] | T0 | 8.5 [4.4, 12.7] | 100.0 [98.8, 100.0] | 0.0 / 0.0 |
| grok-4.5 x-ai/grok-4.5 | 95.4 [92.6, 97.8] | T0 | 7.6 [3.8, 11.8] | 100.0 [96.0, 100.0] | 0.0 / 0.0 |
| claude-sonnet-5 anthropic/claude-sonnet-5 | 94.6 [89.4, 98.6] | T0 | 8.4 [2.7, 15.6] | 98.8 [92.8, 100.0] | 0.0 / 0.0 |
| claude-opus-5 anthropic/claude-opus-5 | 94.4 [90.0, 98.0] | T0 | 9.9 [3.9, 17.6] | 99.1 [93.4, 100.0] | 0.0 / 0.0 |
| mistral-medium-3-5 mistralai/mistral-medium-3-5 | 94.2 [87.9, 98.7] | T0 | 13.5 [3.3, 28.3] | 100.0 [98.1, 100.0] | 0.0 / 0.0 |
| gemini-3.1-pro-preview google/gemini-3.1-pro-preview | 93.7 [88.9, 97.8] | T0 | 14.3 [5.3, 25.5] | 97.0 [91.4, 100.0] | 0.0 / 0.0 |
| deepseek-v4-pro deepseek/deepseek-v4-pro | 93.4 [88.7, 96.8] | T0 | 11.1 [5.5, 18.1] | 100.0 [93.5, 100.0] | 0.0 / 0.0 |
| gemini-3.6-flash google/gemini-3.6-flash | 93.3 [87.8, 98.0] | T0 | 9.7 [2.6, 18.1] | 97.2 [91.2, 100.0] | 0.0 / 0.0 |
| gpt-5.6-terra openai/gpt-5.6-terra | 93.3 [87.0, 98.1] | T0 | 11.9 [3.6, 21.9] | 100.0 [93.3, 100.0] | 0.0 / 0.0 |
| glm-5.2 z-ai/glm-5.2 | 92.7 [87.4, 96.9] | T0 | 13.0 [5.9, 22.2] | 97.6 [91.4, 100.0] | 0.0 / 0.0 |
| llama-4-maverick meta-llama/llama-4-maverick | 92.2 [87.8, 96.3] | T0 | 17.9 [8.5, 27.4] | 100.0 [97.2, 100.0] | 0.0 / 0.0 |
| qwen3.8-max qwen/qwen3.8-max | 92.2 [87.5, 96.1] | T0 | 14.7 [7.9, 22.9] | 99.5 [93.9, 100.0] | 0.0 / 0.0 |
Directional lean: diagnostic only
Directional lean: diagnostic only. Shows which groups received more favorable responses in this sample. It is not a ranking, and it does not indicate intent or ideology.
Per-group lean by model: nationality / immigration (diagnostic)
Diverging heatmap of per-group favorability lean for 14 models across 6 groups, centered at zero, range ±8.
| Model | us born | imm mexico | imm china | imm india | imm nigeria | ref syria |
|---|---|---|---|---|---|---|
| kimi-k3 | −0.6 | +1.0 | −1.1 | +0.4 | −0.1 | +0.4 |
| claude-fable-5 | −3.5 | +2.5 | −0.1 | −1.8 | +1.2 | +1.6 |
| gpt-5.6-sol | −1.5 | +1.5 | −0.1 | −0.1 | +0.3 | −0.1 |
| grok-4.5 | −0.6 | −1.1 | +0.8 | −1.2 | +1.0 | +1.3 |
| claude-sonnet-5 | −3.0 | +1.3 | −0.3 | 0.0 | +0.8 | +1.4 |
| claude-opus-5 | −0.8 | +4.1 | −0.4 | −1.4 | +0.3 | −1.8 |
| mistral-medium-3-5 | +0.4 | −2.3 | −2.7 | −1.9 | +1.4 | +5.2 |
| gemini-3.1-pro-preview | −1.4 | −0.8 | −0.8 | −2.5 | −0.8 | +6.4 |
| deepseek-v4-pro | −3.3 | −0.5 | −1.2 | +0.2 | +1.6 | +3.2 |
| gemini-3.6-flash | +0.4 | −3.4 | +1.2 | −1.2 | −2.0 | +5.1 |
| gpt-5.6-terra | −1.4 | −0.6 | −2.0 | +1.0 | −0.2 | +3.3 |
| glm-5.2 | +1.6 | −1.8 | −0.3 | −1.3 | −1.6 | +3.5 |
| llama-4-maverick | −0.1 | +0.2 | +0.5 | +3.5 | −3.8 | −0.3 |
| qwen3.8-max | −4.9 | +0.3 | +0.1 | +1.5 | +0.2 | +2.8 |