Skip to content

Bank v1.0.0 · evaluated 2026-08-07 · n = 2,707–5,828 scored rows per model · sha256:8c9f5828

Race / ethnicity

Cross-model view of consistency for the race / ethnicity attribute. Groups compared: white, black, hispanic, asian, middle eastern, native american. Evidence tiers compare each model against its own controls.

Consistency on race / ethnicity, all models

The shaded band is the run-average control band; each model's own band is on its model page. Ordering within overlapping CIs is not statistically meaningful. Tiers

Control band: the score range this metric produces under non-demographic substitutions (professions, arbitrary tags), the model's baseline substitution noise. How controls calibrate scores

Per-model detail

Leaderboard attribute cells deep-link to the rows below. Evidence tiers: how T0/T1/T2 are assigned.

ModelConsistencyTierWorst-case gapControl-adjustedRefusal asym. (probe / incid.)
gemini-3.1-pro-preview google/gemini-3.1-pro-preview97.9 [95.8, 99.6]T07.7 [1.7, 15.6]100.0 [97.5, 100.0]0.0 / 0.0
gemini-3.6-flash google/gemini-3.6-flash96.1 [93.3, 98.5]T07.9 [3.2, 12.8]100.0 [96.0, 100.0]0.0 / 0.0
claude-fable-5 anthropic/claude-fable-596.0 [93.4, 98.4]T010.0 [3.9, 16.6]98.3 [94.6, 100.0]0.0 / 0.0
claude-opus-5 anthropic/claude-opus-595.8 [92.9, 98.2]T08.5 [3.4, 14.3]100.0 [96.0, 100.0]0.0 / 0.0
grok-4.5 x-ai/grok-4.595.7 [92.0, 98.5]T08.5 [2.6, 17.4]100.0 [95.5, 100.0]0.0 / 0.0
qwen3.8-max qwen/qwen3.8-max94.8 [91.3, 97.8]T08.6 [3.8, 13.8]100.0 [97.4, 100.0]0.0 / 0.0
gpt-5.6-sol openai/gpt-5.6-sol94.1 [90.7, 97.3]T010.1 [4.8, 15.6]100.0 [96.9, 100.0]0.0 / 0.0
mistral-medium-3-5 mistralai/mistral-medium-3-594.0 [89.8, 97.9]T013.8 [6.5, 21.9]100.0 [99.1, 100.0]0.0 / 0.0
gpt-5.6-terra openai/gpt-5.6-terra93.5 [88.8, 97.7]T09.1 [3.8, 14.6]100.0 [94.6, 100.0]0.0 / 0.0
claude-sonnet-5 anthropic/claude-sonnet-593.2 [88.4, 97.5]T015.8 [5.9, 27.7]97.4 [91.3, 100.0]0.0 / 0.0
kimi-k3 moonshotai/kimi-k393.1 [89.4, 96.7]T011.6 [5.9, 17.6]100.0 [95.7, 100.0]0.0 / 0.0
glm-5.2 z-ai/glm-5.292.7 [90.0, 95.2]T012.6 [7.8, 17.8]97.7 [93.6, 100.0]0.0 / 0.0
llama-4-maverick meta-llama/llama-4-maverick91.9 [85.5, 97.2]T016.1 [6.1, 27.6]100.0 [96.1, 100.0]0.0 / 0.0
deepseek-v4-pro deepseek/deepseek-v4-pro91.5 [86.5, 95.7]T011.5 [6.3, 17.6]99.0 [91.5, 100.0]0.0 / 0.0

Directional lean: diagnostic only

Directional lean: diagnostic only. Shows which groups received more favorable responses in this sample. It is not a ranking, and it does not indicate intent or ideology.

Per-group lean by model: race / ethnicity (diagnostic)

Diverging heatmap of per-group favorability lean for 14 models across 6 groups, centered at zero, range ±6.

Modelwhiteblackhispanicasianmiddle easternnative american
gemini-3.1-pro-preview−1.1+1.6−0.4−0.4−1.7+2.3
gemini-3.6-flash+0.9+1.7−1.5+2.1−1.1−2.1
claude-fable-5−2.6+0.3−1.4−0.9+1.8+3.1
claude-opus-5−3.0+2.0−0.2−1.2+2.8−0.4
grok-4.5+1.2+0.8+2.4−1.0−2.1−1.3
qwen3.8-max−2.5+0.1−1.4+1.2+2.9−0.2
gpt-5.6-sol−1.6−1.0+1.4−2.0+1.0+2.0
mistral-medium-3-5−1.9−3.8+2.5+0.2+1.5+1.3
gpt-5.6-terra−2.8+2.4+0.9+0.5−1.5+0.5
claude-sonnet-5−1.1+4.00.0−0.6−2.1−0.1
kimi-k3−2.3+2.8+1.0−0.4−2.3+1.2
glm-5.2−3.7+1.6−0.8+1.6+0.6+0.7
llama-4-maverick−0.1+3.9−0.6−4.5−1.7+2.8
deepseek-v4-pro−2.8+2.2−0.5−0.6−0.8+2.6
Lean:≤ −3−3…−1±1+1…+3≥ +3(purple = below group mean, teal = above; neither is a verdict)
Rows follow this attribute's consistency order, the same order as the dot plot and the table above, and cannot be sorted. Purple: below the all-group mean; teal: above. Neither is a verdict. Definition

Back to leaderboard · All attributes