Disability
Cross-model view of consistency for the disability attribute. Groups compared: wheelchair, blind, deaf, chronic illness, depression. Evidence tiers compare each model against its own controls.
Consistency on disability, all models
Control band: the score range this metric produces under non-demographic substitutions (professions, arbitrary tags), the model's baseline substitution noise. How controls calibrate scores
Per-model detail
Leaderboard attribute cells deep-link to the rows below. Evidence tiers: how T0/T1/T2 are assigned.
| Model | Consistency | Tier | Worst-case gap | Control-adjusted | Refusal asym. (probe / incid.) |
|---|---|---|---|---|---|
| gpt-5.6-sol openai/gpt-5.6-sol | 96.3 [91.6, 99.8] | T0 | 11.5 [1.7, 26.9] | 100.0 [98.3, 100.0] | 0.0 / 0.0 |
| claude-opus-5 anthropic/claude-opus-5 | 96.2 [93.0, 98.9] | T0 | 12.7 [3.3, 26.4] | 100.0 [96.1, 100.0] | 0.0 / 0.0 |
| claude-fable-5 anthropic/claude-fable-5 | 93.7 [83.6, 99.4] | T0 | 10.6 [1.2, 24.6] | 96.0 [85.1, 100.0] | 0.0 / 0.0 |
| gpt-5.6-terra openai/gpt-5.6-terra | 93.3 [88.1, 97.9] | T0 | 13.3 [4.4, 24.2] | 100.0 [94.4, 100.0] | 0.0 / 0.0 |
| grok-4.5 x-ai/grok-4.5 | 92.9 [85.8, 98.1] | T0 | 11.9 [2.9, 25.9] | 97.5 [89.6, 100.0] | 0.0 / 0.0 |
| kimi-k3 moonshotai/kimi-k3 | 91.8 [84.3, 98.2] | T0 | 13.1 [2.7, 27.4] | 100.0 [91.8, 100.0] | 0.0 / 0.0 |
| mistral-medium-3-5 mistralai/mistral-medium-3-5 | 91.2 [84.0, 97.0] | T0 | 20.7 [7.4, 36.3] | 100.0 [94.6, 100.0] | 0.0 / 1.3 |
| gemini-3.6-flash google/gemini-3.6-flash | 90.6 [80.3, 97.9] | T0 | 20.6 [4.6, 40.5] | 94.5 [83.6, 100.0] | 0.0 / 0.0 |
| deepseek-v4-pro deepseek/deepseek-v4-pro | 90.2 [82.9, 96.1] | T0 | 17.1 [6.5, 30.9] | 97.7 [88.6, 100.0] | 0.0 / 0.0 |
| claude-sonnet-5 anthropic/claude-sonnet-5 | 89.0 [78.3, 97.5] | T0 | 16.5 [2.9, 34.1] | 93.1 [82.2, 100.0] | 0.0 / 0.0 |
| glm-5.2 z-ai/glm-5.2 | 88.9 [81.5, 94.4] | T0 | 17.8 [8.3, 30.6] | 93.9 [85.9, 100.0] | 0.0 / 0.0 |
| qwen3.8-max qwen/qwen3.8-max | 88.1 [80.2, 94.5] | T0 | 20.5 [8.5, 35.7] | 95.5 [87.0, 100.0] | 0.0 / 0.0 |
| llama-4-maverick meta-llama/llama-4-maverick | 87.0 [81.3, 92.4] | T0 | 25.8 [15.4, 36.5] | 100.0 [92.0, 100.0] | 0.0 / 0.0 |
| gemini-3.1-pro-preview google/gemini-3.1-pro-preview | 86.1 [76.1, 94.8] | T0 | 27.4 [9.7, 47.2] | 89.4 [78.3, 98.7] | 0.0 / 0.0 |
Directional lean: diagnostic only
Directional lean: diagnostic only. Shows which groups received more favorable responses in this sample. It is not a ranking, and it does not indicate intent or ideology.
Per-group lean by model: disability (diagnostic)
Diverging heatmap of per-group favorability lean for 14 models across 5 groups, centered at zero, range ±10.
| Model | wheelchair | blind | deaf | chronic illness | depression |
|---|---|---|---|---|---|
| gpt-5.6-sol | −0.9 | −0.9 | −0.5 | +2.3 | +0.1 |
| claude-opus-5 | −0.2 | −0.2 | +1.2 | −1.2 | +0.4 |
| claude-fable-5 | +0.6 | +0.5 | +1.7 | −0.7 | −1.9 |
| gpt-5.6-terra | +0.6 | −0.6 | 0.0 | +0.8 | −0.8 |
| grok-4.5 | +1.0 | +3.2 | +1.8 | −1.3 | −4.8 |
| kimi-k3 | +2.3 | +1.6 | +4.0 | −3.0 | −4.9 |
| mistral-medium-3-5 | +3.5 | +0.2 | −1.2 | +1.6 | −4.6 |
| gemini-3.6-flash | −1.2 | +3.4 | +1.0 | +1.9 | −5.3 |
| deepseek-v4-pro | +0.4 | +3.1 | −0.2 | −0.3 | −3.0 |
| claude-sonnet-5 | +2.8 | +3.6 | +3.4 | −2.6 | −7.8 |
| glm-5.2 | −1.0 | +2.5 | +3.9 | −2.7 | −2.8 |
| qwen3.8-max | +2.4 | +2.9 | +4.6 | −4.2 | −5.8 |
| llama-4-maverick | +3.0 | +5.1 | +2.0 | −0.5 | −9.8 |
| gemini-3.1-pro-preview | +1.2 | +8.1 | +1.3 | −4.4 | −6.7 |