Skip to content

Bank v1.0.0 · evaluated 2026-08-07 · n = 2,707–5,828 scored rows per model · sha256:8c9f5828

Gender

Cross-model view of consistency for the gender attribute. Groups compared: man, woman, nonbinary. Evidence tiers compare each model against its own controls.

Consistency on gender, all models

The shaded band is the run-average control band; each model's own band is on its model page. Ordering within overlapping CIs is not statistically meaningful. Tiers

Control band: the score range this metric produces under non-demographic substitutions (professions, arbitrary tags), the model's baseline substitution noise. How controls calibrate scores

Per-model detail

Leaderboard attribute cells deep-link to the rows below. Evidence tiers: how T0/T1/T2 are assigned.

ModelConsistencyTierWorst-case gapControl-adjustedRefusal asym. (probe / incid.)
gemini-3.6-flash google/gemini-3.6-flash98.9 [97.0, 100.0]T01.2 [0.0, 3.2]100.0 [99.4, 100.0]0.0 / 0.0
mistral-medium-3-5 mistralai/mistral-medium-3-598.8 [96.6, 100.0]T01.7 [0.0, 4.9]100.0 [100.0, 100.0]0.0 / 0.0
gpt-5.6-sol openai/gpt-5.6-sol97.9 [94.7, 100.0]T03.2 [0.0, 7.9]100.0 [100.0, 100.0]0.0 / 0.0
gpt-5.6-terra openai/gpt-5.6-terra97.9 [95.6, 100.0]T03.2 [0.0, 6.7]100.0 [100.0, 100.0]0.0 / 0.0
claude-sonnet-5 anthropic/claude-sonnet-597.3 [93.7, 99.6]T02.4 [0.5, 4.8]100.0 [96.6, 100.0]0.0 / 0.0
gemini-3.1-pro-preview google/gemini-3.1-pro-preview96.4 [92.1, 100.0]T05.4 [0.0, 11.9]99.7 [94.7, 100.0]0.0 / 0.0
glm-5.2 z-ai/glm-5.295.7 [91.8, 98.9]T04.2 [0.1, 9.8]100.0 [95.7, 100.0]0.0 / 0.0
claude-opus-5 anthropic/claude-opus-595.7 [91.3, 99.2]T04.9 [1.1, 10.1]100.0 [94.7, 100.0]0.0 / 0.0
kimi-k3 moonshotai/kimi-k395.4 [92.6, 97.7]T05.3 [2.2, 9.4]100.0 [98.5, 100.0]0.0 / 0.0
grok-4.5 x-ai/grok-4.594.9 [90.7, 98.3]T05.1 [2.0, 8.4]99.5 [94.5, 100.0]0.0 / 0.0
qwen3.8-max qwen/qwen3.8-max94.1 [90.8, 97.1]T08.1 [3.8, 12.9]100.0 [96.8, 100.0]0.0 / 0.0
deepseek-v4-pro deepseek/deepseek-v4-pro93.9 [89.0, 98.1]T07.8 [2.1, 15.0]100.0 [93.9, 100.0]0.0 / 0.0
claude-fable-5 anthropic/claude-fable-592.5 [87.2, 97.3]T09.6 [3.1, 17.4]94.9 [88.9, 100.0]0.0 / 2.6
llama-4-maverick meta-llama/llama-4-maverick90.8 [83.3, 97.1]T011.5 [3.7, 22.0]100.0 [94.4, 100.0]0.0 / 0.0

Directional lean: diagnostic only

Directional lean: diagnostic only. Shows which groups received more favorable responses in this sample. It is not a ranking, and it does not indicate intent or ideology.

Per-group lean by model: gender (diagnostic)

Diverging heatmap of per-group favorability lean for 14 models across 3 groups, centered at zero, range ±4.

Modelmanwomannonbinary
gemini-3.6-flash−0.5+0.3+0.3
mistral-medium-3-5−1.0+0.2+0.9
gpt-5.6-sol−0.8+1.6−0.8
gpt-5.6-terra−0.5−0.1+0.7
claude-sonnet-5−0.8+0.2+0.6
gemini-3.1-pro-preview−0.2+2.0−1.6
glm-5.2+0.2−1.2+1.0
claude-opus-5−0.8+1.6−0.8
kimi-k3+0.2+0.2−0.4
grok-4.5+0.4+0.5−0.9
qwen3.8-max−1.9−0.4+2.3
deepseek-v4-pro−0.5+0.3+0.3
claude-fable-5−2.5−1.4+3.9
llama-4-maverick−0.8−1.1+1.9
Lean:≤ −2−2…−0±0+0…+2≥ +2(purple = below group mean, teal = above; neither is a verdict)
Rows follow this attribute's consistency order, the same order as the dot plot and the table above, and cannot be sorted. Purple: below the all-group mean; teal: above. Neither is a verdict. Definition

Back to leaderboard · All attributes