Skip to content

Bank v1.0.0 · evaluated 2026-08-07 · n = 2,707–5,828 scored rows per model · sha256:8c9f5828

Disability

Cross-model view of consistency for the disability attribute. Groups compared: wheelchair, blind, deaf, chronic illness, depression. Evidence tiers compare each model against its own controls.

Consistency on disability, all models

The shaded band is the run-average control band; each model's own band is on its model page. Ordering within overlapping CIs is not statistically meaningful. Tiers

Control band: the score range this metric produces under non-demographic substitutions (professions, arbitrary tags), the model's baseline substitution noise. How controls calibrate scores

Per-model detail

Leaderboard attribute cells deep-link to the rows below. Evidence tiers: how T0/T1/T2 are assigned.

ModelConsistencyTierWorst-case gapControl-adjustedRefusal asym. (probe / incid.)
gpt-5.6-sol openai/gpt-5.6-sol96.3 [91.6, 99.8]T011.5 [1.7, 26.9]100.0 [98.3, 100.0]0.0 / 0.0
claude-opus-5 anthropic/claude-opus-596.2 [93.0, 98.9]T012.7 [3.3, 26.4]100.0 [96.1, 100.0]0.0 / 0.0
claude-fable-5 anthropic/claude-fable-593.7 [83.6, 99.4]T010.6 [1.2, 24.6]96.0 [85.1, 100.0]0.0 / 0.0
gpt-5.6-terra openai/gpt-5.6-terra93.3 [88.1, 97.9]T013.3 [4.4, 24.2]100.0 [94.4, 100.0]0.0 / 0.0
grok-4.5 x-ai/grok-4.592.9 [85.8, 98.1]T011.9 [2.9, 25.9]97.5 [89.6, 100.0]0.0 / 0.0
kimi-k3 moonshotai/kimi-k391.8 [84.3, 98.2]T013.1 [2.7, 27.4]100.0 [91.8, 100.0]0.0 / 0.0
mistral-medium-3-5 mistralai/mistral-medium-3-591.2 [84.0, 97.0]T020.7 [7.4, 36.3]100.0 [94.6, 100.0]0.0 / 1.3
gemini-3.6-flash google/gemini-3.6-flash90.6 [80.3, 97.9]T020.6 [4.6, 40.5]94.5 [83.6, 100.0]0.0 / 0.0
deepseek-v4-pro deepseek/deepseek-v4-pro90.2 [82.9, 96.1]T017.1 [6.5, 30.9]97.7 [88.6, 100.0]0.0 / 0.0
claude-sonnet-5 anthropic/claude-sonnet-589.0 [78.3, 97.5]T016.5 [2.9, 34.1]93.1 [82.2, 100.0]0.0 / 0.0
glm-5.2 z-ai/glm-5.288.9 [81.5, 94.4]T017.8 [8.3, 30.6]93.9 [85.9, 100.0]0.0 / 0.0
qwen3.8-max qwen/qwen3.8-max88.1 [80.2, 94.5]T020.5 [8.5, 35.7]95.5 [87.0, 100.0]0.0 / 0.0
llama-4-maverick meta-llama/llama-4-maverick87.0 [81.3, 92.4]T025.8 [15.4, 36.5]100.0 [92.0, 100.0]0.0 / 0.0
gemini-3.1-pro-preview google/gemini-3.1-pro-preview86.1 [76.1, 94.8]T027.4 [9.7, 47.2]89.4 [78.3, 98.7]0.0 / 0.0

Directional lean: diagnostic only

Directional lean: diagnostic only. Shows which groups received more favorable responses in this sample. It is not a ranking, and it does not indicate intent or ideology.

Per-group lean by model: disability (diagnostic)

Diverging heatmap of per-group favorability lean for 14 models across 5 groups, centered at zero, range ±10.

Modelwheelchairblinddeafchronic illnessdepression
gpt-5.6-sol−0.9−0.9−0.5+2.3+0.1
claude-opus-5−0.2−0.2+1.2−1.2+0.4
claude-fable-5+0.6+0.5+1.7−0.7−1.9
gpt-5.6-terra+0.6−0.60.0+0.8−0.8
grok-4.5+1.0+3.2+1.8−1.3−4.8
kimi-k3+2.3+1.6+4.0−3.0−4.9
mistral-medium-3-5+3.5+0.2−1.2+1.6−4.6
gemini-3.6-flash−1.2+3.4+1.0+1.9−5.3
deepseek-v4-pro+0.4+3.1−0.2−0.3−3.0
claude-sonnet-5+2.8+3.6+3.4−2.6−7.8
glm-5.2−1.0+2.5+3.9−2.7−2.8
qwen3.8-max+2.4+2.9+4.6−4.2−5.8
llama-4-maverick+3.0+5.1+2.0−0.5−9.8
gemini-3.1-pro-preview+1.2+8.1+1.3−4.4−6.7
Lean:≤ −5−5…−1±1+1…+5≥ +5(purple = below group mean, teal = above; neither is a verdict)
Rows follow this attribute's consistency order, the same order as the dot plot and the table above, and cannot be sorted. Purple: below the all-group mean; teal: above. Neither is a verdict. Definition

Back to leaderboard · All attributes