Race / ethnicity
Cross-model view of consistency for the race / ethnicity attribute. Groups compared: white, black, hispanic, asian, middle eastern, native american. Evidence tiers compare each model against its own controls.
Consistency on race / ethnicity, all models
Control band: the score range this metric produces under non-demographic substitutions (professions, arbitrary tags), the model's baseline substitution noise. How controls calibrate scores
Per-model detail
Leaderboard attribute cells deep-link to the rows below. Evidence tiers: how T0/T1/T2 are assigned.
| Model | Consistency | Tier | Worst-case gap | Control-adjusted | Refusal asym. (probe / incid.) |
|---|---|---|---|---|---|
| gemini-3.1-pro-preview google/gemini-3.1-pro-preview | 97.9 [95.8, 99.6] | T0 | 7.7 [1.7, 15.6] | 100.0 [97.5, 100.0] | 0.0 / 0.0 |
| gemini-3.6-flash google/gemini-3.6-flash | 96.1 [93.3, 98.5] | T0 | 7.9 [3.2, 12.8] | 100.0 [96.0, 100.0] | 0.0 / 0.0 |
| claude-fable-5 anthropic/claude-fable-5 | 96.0 [93.4, 98.4] | T0 | 10.0 [3.9, 16.6] | 98.3 [94.6, 100.0] | 0.0 / 0.0 |
| claude-opus-5 anthropic/claude-opus-5 | 95.8 [92.9, 98.2] | T0 | 8.5 [3.4, 14.3] | 100.0 [96.0, 100.0] | 0.0 / 0.0 |
| grok-4.5 x-ai/grok-4.5 | 95.7 [92.0, 98.5] | T0 | 8.5 [2.6, 17.4] | 100.0 [95.5, 100.0] | 0.0 / 0.0 |
| qwen3.8-max qwen/qwen3.8-max | 94.8 [91.3, 97.8] | T0 | 8.6 [3.8, 13.8] | 100.0 [97.4, 100.0] | 0.0 / 0.0 |
| gpt-5.6-sol openai/gpt-5.6-sol | 94.1 [90.7, 97.3] | T0 | 10.1 [4.8, 15.6] | 100.0 [96.9, 100.0] | 0.0 / 0.0 |
| mistral-medium-3-5 mistralai/mistral-medium-3-5 | 94.0 [89.8, 97.9] | T0 | 13.8 [6.5, 21.9] | 100.0 [99.1, 100.0] | 0.0 / 0.0 |
| gpt-5.6-terra openai/gpt-5.6-terra | 93.5 [88.8, 97.7] | T0 | 9.1 [3.8, 14.6] | 100.0 [94.6, 100.0] | 0.0 / 0.0 |
| claude-sonnet-5 anthropic/claude-sonnet-5 | 93.2 [88.4, 97.5] | T0 | 15.8 [5.9, 27.7] | 97.4 [91.3, 100.0] | 0.0 / 0.0 |
| kimi-k3 moonshotai/kimi-k3 | 93.1 [89.4, 96.7] | T0 | 11.6 [5.9, 17.6] | 100.0 [95.7, 100.0] | 0.0 / 0.0 |
| glm-5.2 z-ai/glm-5.2 | 92.7 [90.0, 95.2] | T0 | 12.6 [7.8, 17.8] | 97.7 [93.6, 100.0] | 0.0 / 0.0 |
| llama-4-maverick meta-llama/llama-4-maverick | 91.9 [85.5, 97.2] | T0 | 16.1 [6.1, 27.6] | 100.0 [96.1, 100.0] | 0.0 / 0.0 |
| deepseek-v4-pro deepseek/deepseek-v4-pro | 91.5 [86.5, 95.7] | T0 | 11.5 [6.3, 17.6] | 99.0 [91.5, 100.0] | 0.0 / 0.0 |
Directional lean: diagnostic only
Directional lean: diagnostic only. Shows which groups received more favorable responses in this sample. It is not a ranking, and it does not indicate intent or ideology.
Per-group lean by model: race / ethnicity (diagnostic)
Diverging heatmap of per-group favorability lean for 14 models across 6 groups, centered at zero, range ±6.
| Model | white | black | hispanic | asian | middle eastern | native american |
|---|---|---|---|---|---|---|
| gemini-3.1-pro-preview | −1.1 | +1.6 | −0.4 | −0.4 | −1.7 | +2.3 |
| gemini-3.6-flash | +0.9 | +1.7 | −1.5 | +2.1 | −1.1 | −2.1 |
| claude-fable-5 | −2.6 | +0.3 | −1.4 | −0.9 | +1.8 | +3.1 |
| claude-opus-5 | −3.0 | +2.0 | −0.2 | −1.2 | +2.8 | −0.4 |
| grok-4.5 | +1.2 | +0.8 | +2.4 | −1.0 | −2.1 | −1.3 |
| qwen3.8-max | −2.5 | +0.1 | −1.4 | +1.2 | +2.9 | −0.2 |
| gpt-5.6-sol | −1.6 | −1.0 | +1.4 | −2.0 | +1.0 | +2.0 |
| mistral-medium-3-5 | −1.9 | −3.8 | +2.5 | +0.2 | +1.5 | +1.3 |
| gpt-5.6-terra | −2.8 | +2.4 | +0.9 | +0.5 | −1.5 | +0.5 |
| claude-sonnet-5 | −1.1 | +4.0 | 0.0 | −0.6 | −2.1 | −0.1 |
| kimi-k3 | −2.3 | +2.8 | +1.0 | −0.4 | −2.3 | +1.2 |
| glm-5.2 | −3.7 | +1.6 | −0.8 | +1.6 | +0.6 | +0.7 |
| llama-4-maverick | −0.1 | +3.9 | −0.6 | −4.5 | −1.7 | +2.8 |
| deepseek-v4-pro | −2.8 | +2.2 | −0.5 | −0.6 | −0.8 | +2.6 |